↓ Skip to main content

The Clinical AI Error That Never Shows Up in the Benchmark

10 min read
Dr. Sophia Patel
Dr. Sophia Patel AI in Healthcare Expert & Machine Learning Specialist

The same 2026 study that showed AI makes doctors better also showed doctors following it when it was wrong.

Both findings were real. Only one made the headline.

In September, the I3LUNG consortium published in Nature Medicine the largest international real-world test yet of an explainable AI decision-support tool in advanced non-small cell lung cancer. The study enrolled 2,396 patients. In its clinical-usability substudy, both lung-cancer experts and non-experts improved when the explainable tool supported them: accuracy at predicting disease control rose from 57% to 65%. That is the number the announcement led with. Buried in the same evaluation: clinicians accepting incorrect AI suggestions, and a model whose performance slipped in external validation, with an area under the curve falling from as high as 0.77 in the test set to 0.55–0.72 elsewhere.

A quiet nighttime clinical workstation photographed from a low angle, its screen glowing with a single confident AI recommendation in warm amber light, while a translucent outline of a clinician's hand reaches toward the mouse — the human-in-the-loop drawn as a faint, empty orbit.
We keep measuring the model. The hand that decides what to do with it is the part we never test.

The thread: in clinical AI, accuracy scores measure the algorithm, but safety lives in the clinician-plus-AI pair — and almost nobody measures the pair. As frontier models match or exceed physicians on benchmarks, the dominant failure mode is shifting from “the model is wrong” to “the human follows the model when it is wrong.” Hospitals that buy on benchmark scores inherit a liability they have no metric, no monitoring, and no named owner for.

The unit of safety was always the wrong one
#

This is not a new phenomenon. It has a name: automation bias, the tendency to over-rely on an automated system’s output. A systematic review in JAMIA by Kate Goddard, Aziz Roudsari and Jeremy Wyatt pulled 74 qualifying papers from more than 13,000 and found that automation bias and automation-induced complacency are well documented in clinical decision support. The conditions that worsen it — high workload, task complexity, time pressure — describe a normal hospital shift. The mitigators they identified were strikingly mundane: training, explicit user accountability, showing calibrated confidence, and presenting information rather than a recommendation.

Fourteen years later, most generative clinical tools are designed to do the opposite. They deliver confident recommendations, not calibrated information, and they do it in fluent prose that reads like a colleague who has already done the work.

The damage is asymmetric — and explanations don’t fix it
#

The clearest evidence on cost comes from a randomized vignette study published in JAMA by Sarah Jabbour and colleagues at the University of Michigan. They randomized 457 hospitalists, nurse practitioners and physician assistants across 13 US states to diagnose patients with acute respiratory failure, sometimes with AI input.

When the AI was correct, clinician accuracy rose 2.9 percentage points without explanations and 4.4 points with them. When the AI was systematically biased, accuracy fell 11.3 points versus baseline. Adding explanations to the biased model still left accuracy 9.1 points below baseline — a rescue of just 2.3 points, not statistically significant.

Read that as a business case. The upside of a good suggestion is small. The downside of a bad one is large and asymmetric. And the intervention everyone reaches for — explainability — is a trust lever, not a safety lever. A model that is right most of the time can still net-harm clinician accuracy, and a persuasive explanation can make the wrong answer easier to accept.

Expertise is not a shield
#

If you believe senior clinicians are too experienced to be swayed, the mammography evidence says otherwise. In a prospective experiment published in Radiology, Tabea Dratsch and colleagues had 27 radiologists read 50 mammograms with a purported AI system. After a warm-up set where the AI was always correct, the researchers seeded incorrect BI-RADS suggestions. Accuracy collapsed: inexperienced readers fell from 79.7% to 19.8%, moderately experienced readers from 81.3% to 24.8%, and very experienced readers from 82.3% to 45.5%. Experience reduced the fall. It did not prevent it.

The reverse failure is just as real. In a randomized trial in JAMA Network Open, Ethan Goh and colleagues at Stanford found that giving 50 physicians access to a large language model did not significantly improve diagnostic reasoning — a 2-point difference over conventional resources. The model on its own scored 16 points higher than the physicians’ conventional-resources group. A 2026 UK replication led by J. Healy and colleagues found the same pattern and uncovered why: physicians posed only 30% of their case questions to the model. They were not blindly over-trusting it. Many were quietly ignoring it.

That is the uncomfortable part. “Human in the loop” is not one failure mode. It is a spectrum, from barely using the tool to obeying it. A safety program that assumes a single, appropriately-assisted clinician is measuring a person who does not exist.

The new danger: AI that acts, not just advises
#

Automation bias used to be about advice a human could decline. Agentic AI removes that pause. A 2026 preprint by Mingwei Nie and colleagues tested browser agents against a simulated prior-authorization portal and a synthetic EHR of 836 patient records, some deliberately deficient. The larger models completed the task almost every time — 95.45% for Gemini 3 Pro, 93.67% for Claude Opus 4.5 — but nearly all of them submitted deficient requests anyway. In a non-agentic setting, Gemini 3 Pro correctly identified 91% of those deficiencies. The model knew. The workflow told it to finish the form.

This is the machine version of the same bias, and it connects directly to the provider-and-insurer algorithm arms race I described in The AI War Inside Your Hospital Bill: an agent optimised to complete a task will treat “submit” as the goal, even when its own reasoning flags a problem. As agentic tools move into prior authorization, scheduling and order entry, “human in the loop” has to become an explicit hold decision — or the agent’s completion bias becomes institutional policy by default.

Why nobody is measuring the pair
#

The reason this keeps happening is not that we lack evidence. It is that we keep collecting the wrong kind.

ARISE’s State of Clinical AI 2026, published in BMJ Digital Health & AI on September 29, found that fewer than 5% of FDA-cleared AI and machine-learning devices have undergone peer-reviewed evaluation, that frontier models still fail at uncertainty calibration, and that “risks of automation bias and clinician deskilling are emerging.” A scoping review in Frontiers in Digital Health of 78 LLM studies found that 89% were benchmark or simulated-workflow studies and only 11% were prospective, workflow-embedded evaluations — with automation bias listed among the safety risks those studies cannot detect. And a September 2026 synthesis in Radiology: Artificial Intelligence named the three determinants the field keeps overlooking: cognition (automation bias, uncertainty framing, AI-induced skill decay), interface design (how AI findings shape attention and reporting), and “algorithmic conformity” — clinicians overriding their own judgment under medicolegal pressure.

Put together, the picture is clear. The model has a benchmark. The pair has none. There is no clearance category, no standard metric and no postmarket reporting requirement for the human-AI interaction that actually decides whether a patient is helped or harmed.

Five questions a mature program would ask
#

If I were reviewing a clinical AI deployment this week, I would stop at the AUROC and start here:

  1. What is the deference-when-wrong rate? Track how often clinicians accept AI output that is later shown to be incorrect — and how that changes with workload and time of day. Adoption dashboards are not safety dashboards.
  2. Does the interface present information or a recommendation? Goddard’s 2012 review found that framing, position and calibrated confidence levels change behaviour. Most current tools ignore all three.
  3. What is the distribution of use? Measure the range from under-use to over-use, not an average. A tool ignored by half the team and obeyed blindly by the other half is not “working.”
  4. Who owns the loop, and can they pause it? Assign a named clinical owner with real authority to suspend, retrain or retire the system when override or error patterns move the wrong way.
  5. Is human-factors evidence part of procurement? Ask vendors for prospective, workflow-embedded studies of the pair — not benchmark scores.

The regulatory window is open right now. The FDA’s generative-AI discussion paper, issued August 18, 2026, is accepting comments under docket FDA-2026-N-7874 only until October 19, 2026, and it leans toward grading models — benchmarks, clinical confirmation, then live monitoring — without yet requiring deference or override metrics. The HIMSS AI in Healthcare Forum on October 22–23 is built explicitly around governance and accountability. Both are chances to insist that “live monitoring” watches the pair, not just the model.

The AI that helps a radiologist and the AI that harms a patient can be the exact same model, deployed into two different human workflows. We certified the model. We never tested the hand that decides what to do with it. The model has learned from us. The harder question is whether we still learn from ourselves — or just do what the screen says.

Working on clinical AI evaluation, procurement, human factors or governance and seeing a gap between the top-line score and real deployment risk? I’d like to hear how you measure what happens after the AI makes its suggestion.

Email me at sophia.patel@tlnw.uk

A vertical infographic showing five data signals behind clinical AI automation bias: explanations barely helping accuracy, expert radiologists' accuracy collapsing under wrong AI advice, an LLM outperforming physicians who use it, browser agents submitting deficient prior-authorization requests, and fewer than 5% of FDA-cleared AI devices being peer-reviewed.
We certified the model. We never tested the clinician holding the mouse.

References
#

  • Prelaj, A., Miskovic, V., et al. (September 13, 2026). “Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC.” Nature Medicine, 32(9):3235–3247. https://pubmed.ncbi.nlm.nih.gov/42733093/ (Accessed October 6, 2026)
  • Goddard, K., Roudsari, A., and Wyatt, J. C. (2012). “Automation bias: a systematic review of frequency, effect mediators, and mitigators.” Journal of the American Medical Informatics Association, 19(1):121–127. https://pubmed.ncbi.nlm.nih.gov/21685142/ (Accessed October 6, 2026)
  • Jabbour, S., Fouhey, D., et al. (December 19, 2023). “Measuring the Impact of AI in the Diagnosis of Hospitalized Patients: A Randomized Clinical Vignette Survey Study.” JAMA, 330(23):2275–2284. https://pubmed.ncbi.nlm.nih.gov/38112814/ (Accessed October 6, 2026)
  • Dratsch, T., Chen, X., et al. (May 2023). “Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance.” Radiology, 307(4):e222176. https://pubmed.ncbi.nlm.nih.gov/37129490/ (Accessed October 6, 2026)
  • Goh, E., Gallo, R., et al. (October 1, 2024). “Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial.” JAMA Network Open, 7(10):e2440969. https://pubmed.ncbi.nlm.nih.gov/39466245/ (Accessed October 6, 2026)
  • Healy, J., et al. (April 29, 2026). “Human-AI collaboration in clinical reasoning: a UK replication and interaction analysis.” Diagnosis (Berl). https://pubmed.ncbi.nlm.nih.gov/42053372/ (Accessed October 6, 2026)
  • Nie, M., Chung, W., et al. (June 18, 2026). “Hard to Halt: Automation Bias in Agent-Driven Sequencing Prior Authorization Workflows.” medRxiv 2026.06.16.26355782. https://pubmed.ncbi.nlm.nih.gov/42369474/ (Accessed October 6, 2026)
  • Worth, J. E., Perez, A., et al. (September 29, 2026). “State of clinical AI in 2026.” BMJ Digital Health & AI, 2(1):e000113. https://pubmed.ncbi.nlm.nih.gov/42825240/ (Accessed October 6, 2026)
  • Ferreira, J. C., and Rosa, I. (September 18, 2026). “Large language models in healthcare: applications, evaluation frameworks, and governance pathways.” Frontiers in Digital Health, 8:1865568. https://pubmed.ncbi.nlm.nih.gov/42828300/ (Accessed October 6, 2026)
  • Kim, S. H., Adams, L. C., Wiestler, B., and Hedderich, D. M. (September 2026). “Human-AI Collaboration in Radiology: The Blind Spots.” Radiology: Artificial Intelligence, 8(5):e260325. https://pubmed.ncbi.nlm.nih.gov/42684148/ (Accessed October 6, 2026)
  • U.S. Food and Drug Administration. (August 18, 2026). “Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback.” Docket FDA-2026-N-7874, comments due October 19, 2026. https://www.regulations.gov/docket/FDA-2026-N-7874 (Accessed October 6, 2026)
  • Vero Scribe. (September 2026). “Healthcare AI News: September 2026 Briefing” (checked through September 21, 2026). https://www.veroscribe.com/blog/healthcare-ai-news-september-2026 (Accessed October 6, 2026)
  • HIMSS. (October 22–23, 2026). “AI in Healthcare Forum, San Diego.” https://www.himss.org/events-overview/ai-in-healthcare-forum-san-diego/overview/ (Accessed October 6, 2026)

AI Content Notice

This article was created using artificial intelligence technology. Whenever possible, we include references and sources to support the information presented. Readers are encouraged to consult these sources for further information. While we strive for accuracy and provide valuable insights, readers should independently verify information and use their own judgment when making business decisions. The content may not reflect real-time market conditions or personal circumstances.

Related Articles