Skip to content

Oversight is designed, not staffed

Last month the BMJ called the clinician in the loop a liability sink. Being present is not the same as judging. The authors concede that oversight works on tasks that are bounded and checkable, so that is the design target: make the task checkable, put approvals where someone has time and authority, and test whether people still catch an error.

Written by

Essay3 min read

On 5 May the BMJ published an analysis by Toro-Tobon and colleagues at Mayo Clinic: Clinician in the loop: a flawed solution for AI oversight. The same day Nicolas Spatola wrote on Tech Policy Press that efficiency can undermine accountability even with a human in the loop. Both say, in different ways, what I have been circling for months.

Present is not the same as judging

Putting a person at the end of a pipeline does not mean the pipeline is overseen. When the format rewards speed, and the workplace pushes the same way, accepting what the system produced becomes the rational thing to do. The BMJ authors call the result a liability sink and use a term that is not theirs, the moral crumple zone, which Madeleine Elish coined: the person absorbs the blow for a system she could not have controlled. Spatola is more direct about the cause: for the person under pressure, deferring to the system is the sensible choice.

In March Biro, Jabbarpour and Ratwani wrote in Lancet Primary Care that safety emerges from the team of people and system in its setting, and that one-off approval and nominal review fail where continuous monitoring is needed. Xavier Geerinck put it as a design choice: governance runs while the system runs, or it does not exist.

The concession I build on

The BMJ authors grant that oversight does work for tasks that are bounded and verifiable. I take that as the design target. The question is not how to get a human to look harder. It is how to make the task such that looking is enough.

For Inora that means the following. The answer comes from designated sources. Checking it takes one look at one passage in one named source. A gap is declared and not filled. Approvals sit where someone has the time and the authority: the owner of the protocol, the quality function. They do not sit on every answer a care worker reads. And in the evaluation I want to know whether professionals still catch an error I planted, after weeks of use. That is the real measure of oversight, and I have not seen it measured.

Where I disagree

The BMJ’s ordering, human first and then the tool, suits a diagnosis. It does not fit a care worker asking what her organisation’s protocol says. That is not a clinical judgement. It is a lookup, and we keep it on that side of the line.

Physicians are also the wrong reference group. Review-based safety fails harder further down the ladder of training, and a lot of elderly care is done by people with less of it. If a review cannot rescue a doctor, it will not rescue a care assistant.

What I watch

Inora answers when asked and pushes no alerts. To see whether people lean on it too much, I would watch how often sources are opened, how often an answer is overruled, and how many gaps are sent to their owner. That is a hypothesis until the record measures it, and I will say so when it does.

Work in care and use Inora? Sign-in and help go through your own organisation: your team’s project lead and ambassadors can help you.