
Physicians May Struggle to Spot AI Errors, Even When Evidence Contradicts the Advice
Key Takeaways:
- In experiments involving decisions about hypothetical patients, physicians tended to trust incorrect advice labelled as artificial intelligence (AI) generated, even when they had the chance to notice that patient recovery data contradicted it.
- Across both experiments, physicians rated the AI system as reliable and did not draw on the recovery data to conclude that its recommendations were wrong; in the second experiment, they failed to notice that the treatment was entirely ineffective.
- The findings, published in the open-access journal PLOS Digital Health by Aranzazu Vinas of the University of the Basque Country and colleagues, point to real challenges for the widely held assumption that a human will reliably catch and correct an algorithm’s mistakes.
What the research examined
New research suggests that physicians may find it difficult to learn from experience when that experience runs counter to advice presented as coming from an AI system. In a series of experiments in which physicians made decisions about treating hypothetical patients, they tended to trust incorrect AI-labelled recommendations, even after being given the opportunity to notice that patient recovery data contradicted those recommendations.
The work was carried out by Aranzazu Vinas of the University of the Basque Country, Spain, together with colleagues, and is published in the open-access journal PLOS Digital Health.
Why AI classification matters in care
AI systems can help physicians categorise patients according to their differing care needs, for example by estimating whether a particular patient is more or less likely to benefit from a given treatment. Because these systems are not perfect, they are intended to be used as suggestions rather than instructions, with any potential errors caught and corrected by the physician using them.
This safeguard rests on an assumption that is easy to take for granted: that a human in the loop will notice when the algorithm is wrong. Prior research, however, has shown that people in general struggle to spot and correct mistakes made by AI. Vinas and colleagues set out to explore how far that difficulty extends to physicians in particular.
How the experiments worked
The researchers analysed data from 223 physicians who took part anonymously in online experiments. Participants were asked to imagine that they had the option to treat patients for a rare disease using a treatment that was not yet proven and still under development. They were told that an AI system had identified which patients were more, or less, likely to benefit from that treatment.
The physicians then chose which patients to treat. After being shown data on how those patients recovered, they rated their perceptions of how reliable the AI system was.
The design contained a deliberate mismatch. The actual effectiveness of the hypothetical treatment did not align with the AI’s recommendations. In the first experiment, the treatment was equally, and moderately, effective for all patients. In the second experiment, it was equally ineffective for everyone. In each case, the recovery data available to physicians should, in principle, have allowed them to see that the AI’s classification did not hold up.
What the physicians did
In both experiments, the physicians tended to rate the AI system as reliable, and they did not appear to use the patient recovery data to conclude that the AI’s recommendations were incorrect. In the second experiment, they did not realise that the treatment was entirely ineffective.
As lead author Aranzazu Vinas notes: “In both experiments, physicians mostly trusted the AI’s classifications and had trouble learning from the feedback. Furthermore, in the second experiment, professionals did not notice that the treatment was completely ineffective.”
Co-author Helena Matute adds: “People tend to say that there is always a human controlling the algorithm, but our experiments show that doctors (as well as anyone else) have problems in learning from the available evidence when it contradicts the suggestions of an algorithm.”
What it means for healthcare
Taken together, the results highlight potential challenges for incorporating AI-based classification into healthcare. If the human overseeing an algorithm cannot readily detect its errors, even when contradicting evidence is in front of them, then the reassurance that a clinician will always catch a mistake may be weaker than commonly assumed.
The authors suggest that future research could build on this study, for instance by developing and testing strategies and protocols designed to strengthen human critical thinking and the detection of AI errors. The aim would be to maximise the benefits of human-AI collaboration while minimising the potential for error.
Co-author Fernando Blanco summarises the wider purpose of this line of enquiry: “It is important to investigate the errors that humans (including doctors) make when working with algorithms, in order to learn how to minimize the problems that arise from them.”
Building the habit of questioning AI
While researchers work on formal protocols, individual clinicians can already sharpen how they interrogate AI output. Knowing when to trust a recommendation, and when to challenge it, is a clinical skill rather than a technical one, and it is one that structured training can help build. Our short course AI Essentials for GPs: Tools, Ethics and Everyday Applications introduces practical frameworks, including the SAFER Evaluation Framework, for spotting errors, fabrications, and outdated recommendations before they reach a patient. For clinicians who want a reliable method for the kind of critical checking this study suggests is all too easy to skip, it is a useful place to start.
Read More
AI Outperformed Emergency Doctors in Harvard Triage Study, Raising Questions About the Future of Clinical Decision-Making
Key Takeaways:
- A Harvard-led study found that an AI reasoning model outperformed emergency physicians in diagnosing patients during hospital triage scenarios using text-based clinical information.
- Researchers said the findings represent a major advance in AI clinical reasoning, although they stressed that AI is not ready to replace human doctors.
- Experts warned that important concerns remain around accountability, bias, safety, and the risk of clinicians becoming overly reliant on AI systems.
AI shows strong performance in emergency medicine trial
From fictional emergency department heroes such as George Clooney in ER to Noah Wyle in The Pitt, emergency physicians have long been portrayed as the ultimate decision-makers in moments of medical crisis. However, a new Harvard study suggests artificial intelligence may increasingly play a major role in those same high-pressure situations.
Researchers from Harvard Medical School and Beth Israel Deaconess Medical Center found that advanced AI systems outperformed human doctors in emergency medicine triage scenarios, making more accurate diagnoses when presented with limited patient information during the critical early stages of hospital admission.
The findings, published in Science, were described by independent experts as representing “a genuine step forward” in AI clinical reasoning.
According to the study authors, large language models (LLMs) “have eclipsed most benchmarks of clinical reasoning”.
AI versus doctors in emergency room triage
One of the study’s central experiments examined 76 patients who presented to the emergency department of a Boston hospital.
Both the AI system and pairs of human physicians were given identical electronic health record information to assess. This included standard triage details such as:
- Vital signs
- Demographic information
- Brief nursing notes explaining why the patient attended hospital
Using only this text-based information, OpenAI’s o1 reasoning model identified the exact diagnosis or a very close diagnosis in 67% of cases.
By comparison, the human physicians achieved diagnostic accuracy rates of between 50% and 55%.
Researchers found the AI’s advantage was especially apparent in triage situations requiring rapid decision-making with minimal available information.
When additional clinical detail was provided, the AI’s diagnostic accuracy increased further to 82%. Human experts achieved accuracy rates between 70% and 79% under those circumstances, although researchers noted the difference was not statistically significant in that setting.
AI also performed better in treatment planning
The study also evaluated how AI performed in longer-term clinical planning tasks.
In this experiment, the AI system and a group of 46 doctors were asked to review five detailed clinical case studies and develop treatment strategies. These included decisions relating to:
- Antibiotic regimens
- Ongoing management plans
- End-of-life care processes
The AI significantly outperformed the doctors.
Researchers reported that the AI achieved a score of 89%, compared with 34% among physicians using conventional resources such as search engines.
AI detected a diagnosis human doctors missed
One example highlighted in the study involved a patient with a pulmonary embolism and worsening symptoms.
Human doctors believed the patient’s anticoagulant treatment was failing. However, the AI system identified something clinicians had overlooked – the patient had a history of lupus, which may have been responsible for inflammation in the lungs.
The AI’s interpretation was ultimately confirmed as correct.
Researchers say AI will reshape medicine – not replace doctors
Despite the strong performance shown by AI systems, researchers stressed that the technology is not ready to replace physicians.
The study only assessed AI systems using text-based patient information. It did not evaluate the AI’s ability to interpret non-verbal clinical signals that doctors routinely use during patient assessment, such as:
- Visible distress
- Facial appearance
- Behaviour
- Physical examination findings
As a result, researchers said the AI functioned more like a clinician reviewing paperwork and offering a second opinion.
“I don’t think our findings mean that AI replaces doctors,” said Arjun Manrai, one of the lead authors of the study who heads an AI lab at Harvard Medical School. “I think it does mean that we’re witnessing a really profound change in technology that will reshape medicine.”
Dr Adam Rodman, another lead author and physician at Beth Israel Deaconess Medical Center, described AI LLMs as among “the most impactful technologies in decades”.
He suggested healthcare may move towards what he described as a “triadic care model”.
“Over the next decade,” Rodman said, AI would not replace physicians but instead work alongside them in a new model involving “the doctor, the patient, and an artificial intelligence system”.
AI use in healthcare is already growing
The findings come amid rapidly increasing AI adoption within healthcare systems.
According to research published last month, nearly one in five physicians in the United States are already using AI to assist with diagnosis.
In the United Kingdom, a recent Royal College of Physicians survey found:
- 16% of doctors use AI daily
- A further 15% use AI weekly
- Clinical decision-making is among the most common applications
However, concerns around safety and accountability remain significant.
UK doctors surveyed identified AI errors and legal liability as among their biggest worries.
“There is not a formal framework right now for accountability,” said Rodman.
He also stressed the continuing importance of human clinicians in patient care.
“Patients ultimately want humans to guide them through life or death decisions [and] to guide them through challenging treatment decisions,” he said.
Experts warn against over-reliance on AI
Independent experts said the study highlights the rapidly improving capabilities of AI systems in medicine, but also demonstrates the need for caution.
Prof Ewen Harrison, co-director of the University of Edinburgh’s Centre for Medical Informatics, said the findings suggest AI systems are beginning to evolve beyond theoretical testing environments.
“These systems are no longer just passing medical exams or solving artificial test cases,” he said. “They are starting to look like useful second-opinion tools for clinicians, particularly when it is important to consider a wider range of possible diagnoses and avoid missing something important.”
However, Dr Wei Xing from the University of Sheffield warned that the study also raised concerns about how clinicians interact with AI recommendations.
He suggested some doctors may unconsciously defer to AI-generated answers instead of independently evaluating clinical information themselves.
“This tendency could grow more significant as AI becomes more routinely used in clinical settings,” he said.
Dr Xing also noted that the study provided limited information about where the AI may perform less effectively, including whether diagnostic accuracy differed among certain patient populations such as:
- Older adults
- Non-English speakers
- People with more complex communication needs
He cautioned against interpreting the findings as evidence that publicly available AI tools are ready for independent medical use.
“It does not demonstrate that AI is safe for routine clinical use, nor that the public should turn to freely available AI tools as a substitute for medical advice,” he said.
Source: The Guardian
Read More