
Machine Learning Tool Helps Paediatricians Identify Children at Risk of Persistent Asthma
Key Takeaways:
- A machine learning tool that reads data already held in a child’s electronic health record helped paediatricians more accurately judge which young children are at risk of persistent asthma.
- In a pilot randomised trial using standardised clinical cases, clinicians using the tool reached an average accuracy of 83%, compared with 61% for standard assessment alone.
- The tool is designed to support clinical judgement rather than replace it, and requires no additional tests or questionnaires.
Support for a difficult clinical judgement
A machine learning tool that analyses information already captured in a child’s electronic health record (EHR) has helped paediatricians assess asthma risk more accurately in standardised clinical case scenarios, according to a pilot randomised clinical trial led by a researcher at the Regenstrief Institute. The study was published in the journal Scientific Reports.
The trial evaluated a machine learning-enabled clinical decision support tool known as the Passive Digital Marker. The tool draws on routinely collected EHR data to classify young children as having either a high or a low risk of going on to develop persistent asthma.
Why early asthma risk is hard to predict
Asthma is one of the most common long-term conditions of childhood, yet predicting which young children who have wheezing or other respiratory symptoms will later develop persistent asthma remains difficult. Some children outgrow their early symptoms, while others need ongoing treatment. That uncertainty makes early risk assessment an important, but genuinely challenging, part of paediatric care.
“This tool doesn’t replace a pediatrician’s clinical judgment,” said Arthur H. Owora, PhD, Regenstrief Institute research scientist and lead author of the study. “It helps bring together years of clinical information that’s already in the electronic health record, giving clinicians another source of information when making decisions about a child’s asthma risk.”
How the Passive Digital Marker works
Unlike many prediction tools, the Passive Digital Marker requires no extra testing and asks families to complete no additional questionnaires. Instead, it analyses information that has already been documented in the child’s EHR, including respiratory symptoms, allergies, medication history, respiratory infections and family history. It then presents clinicians with a straightforward high-risk or low-risk assessment.
This approach is intended to save clinicians’ time and reduce the burden on families, since it relies on data that has been gathered over the course of a child’s routine care rather than requiring anything new at the point of decision.
What the trial found
Paediatricians using the tool correctly predicted future asthma more often than those relying on standard assessment alone, achieving an average accuracy of 83% compared with 61%. The improvement was largely driven by better identification of children who went on to develop persistent asthma – the group that is most important to recognise early and hardest to spot.
The researchers stress that the tool is meant to support clinical decision-making, not to supplant it. Its value lies in helping clinicians quickly synthesise years of patient information into a single, easy-to-interpret risk assessment that sits alongside their own expertise. That distinction – between having an AI tool to hand and knowing how to weigh what it tells you – is becoming central to how clinicians are expected to work with these systems.
Limitations and next steps
Because the study used standardised patient cases rather than real-world clinical encounters, further research is needed to establish whether the tool improves outcomes for children in everyday paediatric practice. The pilot demonstrates promise in a controlled setting, but real-world validation is the necessary next stage before wider adoption.
CCH insight
Tools like the Passive Digital Marker are only ever as good as a clinician’s ability to judge when to lean on them and when to look again. That skill – evaluating an AI tool, recognising where it can mislead, and putting sensible governance around its use – is exactly what our short course AI Essentials for GPs: Tools, Ethics and Everyday Applications is designed to build. It’s a 3.5-hour, fully online CPD course led by Prof. Mike Bewick and Dr Dipesh Naik. [Explore the course →]
Funding and authorship
The study was supported in part by the National Institutes of Health under grant K01HL166436. In addition to Owora, it was co-authored by Bowen Jiang, M.S., and Yash Shah, M.S., of the Division of Pediatric Pulmonology, Allergy/Immunology and Sleep Medicine, Department of Pediatrics, Riley Hospital for Children, Indiana University School of Medicine.
Source: Regenstrief Institute
Read More
Physicians May Struggle to Spot AI Errors, Even When Evidence Contradicts the Advice
Key Takeaways:
- In experiments involving decisions about hypothetical patients, physicians tended to trust incorrect advice labelled as artificial intelligence (AI) generated, even when they had the chance to notice that patient recovery data contradicted it.
- Across both experiments, physicians rated the AI system as reliable and did not draw on the recovery data to conclude that its recommendations were wrong; in the second experiment, they failed to notice that the treatment was entirely ineffective.
- The findings, published in the open-access journal PLOS Digital Health by Aranzazu Vinas of the University of the Basque Country and colleagues, point to real challenges for the widely held assumption that a human will reliably catch and correct an algorithm’s mistakes.
What the research examined
New research suggests that physicians may find it difficult to learn from experience when that experience runs counter to advice presented as coming from an AI system. In a series of experiments in which physicians made decisions about treating hypothetical patients, they tended to trust incorrect AI-labelled recommendations, even after being given the opportunity to notice that patient recovery data contradicted those recommendations.
The work was carried out by Aranzazu Vinas of the University of the Basque Country, Spain, together with colleagues, and is published in the open-access journal PLOS Digital Health.
Why AI classification matters in care
AI systems can help physicians categorise patients according to their differing care needs, for example by estimating whether a particular patient is more or less likely to benefit from a given treatment. Because these systems are not perfect, they are intended to be used as suggestions rather than instructions, with any potential errors caught and corrected by the physician using them.
This safeguard rests on an assumption that is easy to take for granted: that a human in the loop will notice when the algorithm is wrong. Prior research, however, has shown that people in general struggle to spot and correct mistakes made by AI. Vinas and colleagues set out to explore how far that difficulty extends to physicians in particular.
How the experiments worked
The researchers analysed data from 223 physicians who took part anonymously in online experiments. Participants were asked to imagine that they had the option to treat patients for a rare disease using a treatment that was not yet proven and still under development. They were told that an AI system had identified which patients were more, or less, likely to benefit from that treatment.
The physicians then chose which patients to treat. After being shown data on how those patients recovered, they rated their perceptions of how reliable the AI system was.
The design contained a deliberate mismatch. The actual effectiveness of the hypothetical treatment did not align with the AI’s recommendations. In the first experiment, the treatment was equally, and moderately, effective for all patients. In the second experiment, it was equally ineffective for everyone. In each case, the recovery data available to physicians should, in principle, have allowed them to see that the AI’s classification did not hold up.
What the physicians did
In both experiments, the physicians tended to rate the AI system as reliable, and they did not appear to use the patient recovery data to conclude that the AI’s recommendations were incorrect. In the second experiment, they did not realise that the treatment was entirely ineffective.
As lead author Aranzazu Vinas notes: “In both experiments, physicians mostly trusted the AI’s classifications and had trouble learning from the feedback. Furthermore, in the second experiment, professionals did not notice that the treatment was completely ineffective.”
Co-author Helena Matute adds: “People tend to say that there is always a human controlling the algorithm, but our experiments show that doctors (as well as anyone else) have problems in learning from the available evidence when it contradicts the suggestions of an algorithm.”
What it means for healthcare
Taken together, the results highlight potential challenges for incorporating AI-based classification into healthcare. If the human overseeing an algorithm cannot readily detect its errors, even when contradicting evidence is in front of them, then the reassurance that a clinician will always catch a mistake may be weaker than commonly assumed.
The authors suggest that future research could build on this study, for instance by developing and testing strategies and protocols designed to strengthen human critical thinking and the detection of AI errors. The aim would be to maximise the benefits of human-AI collaboration while minimising the potential for error.
Co-author Fernando Blanco summarises the wider purpose of this line of enquiry: “It is important to investigate the errors that humans (including doctors) make when working with algorithms, in order to learn how to minimize the problems that arise from them.”
Building the habit of questioning AI
While researchers work on formal protocols, individual clinicians can already sharpen how they interrogate AI output. Knowing when to trust a recommendation, and when to challenge it, is a clinical skill rather than a technical one, and it is one that structured training can help build. Our short course AI Essentials for GPs: Tools, Ethics and Everyday Applications introduces practical frameworks, including the SAFER Evaluation Framework, for spotting errors, fabrications, and outdated recommendations before they reach a patient. For clinicians who want a reliable method for the kind of critical checking this study suggests is all too easy to skip, it is a useful place to start.
Read More
Autonomous AI Agent Matches and Exceeds Physicians Across Simulated Electronic Health Record Cases
Key Takeaways:
- MIRA is an autonomous AI agent that diagnoses and plans treatment inside a simulated electronic health record, rather than acting as a narrow chat tool.
- It reached 88.9% diagnostic accuracy across 574 cases, outperforming board-certified physicians (78.1%) and a mixed-seniority team (71.1%).
- Safety results were strong but preliminary, and the authors stress that MIRA is not a replacement for human clinicians.
A new kind of medical AI agent
A recent study published in the journal Nature introduced MIRA, an autonomous AI agent designed to operate within sandboxed EHR environments. Rather than acting as a single-purpose assistant, MIRA uses a suite of digital tools to simulate the full arc of a clinical workflow. It can order tests, synthesise the results, and produce diagnoses and treatment plans, all while communicating through a chat interface with a patient AI agent that is grounded in the documented history of present illness extracted from retrospective notes from genuine cases.
The system runs on a Fast Healthcare Interoperability Resources (FHIR) based architecture, which executes the agent’s tool calls and records its medical outputs. The researchers note that the example data presented in the paper were shortened and slightly modified to comply with the privacy restrictions attached to the dataset.
Unlike earlier implementations, which were predominantly task-specific chat applications, MIRA was built to independently take in patient histories, order the relevant diagnostic tests, and then use those datasets to reach diagnoses and treatment plans within a controlled simulation. Across the 574 MIMIC-IV cases, MIRA achieved 88.9% diagnostic accuracy, and in a matched 311-case physician comparison it reached 87.8% accuracy, significantly outperforming experienced human physicians under identical simulated conditions while demonstrating strong, though not perfect, safety and guideline performance.
Background: from passing exams to working a ward
Large language models (LLMs) have already proven highly capable at passing standardised medical examinations and answering complex clinical questions. Reviews of the field show, however, that translating this raw clinical knowledge into the operational workflow of a hospital has remained a major challenge.
This gap is attributed to the architectural design of traditional medical AI tools, which behave as narrow, task-specific search or text-generation utilities rather than as active partners in care. By contrast, true clinical decision-making is characterised as an intricate, multi-step process in which doctors repeatedly interview the people in their care, order blood tests or imaging, synthesise conflicting results, and update their hypotheses before arriving at a final treatment plan.
Nearly all of this clinical work takes place within EHR systems that rely on complex, standardised coding protocols. Until now, it remained unproven whether an automated system could reliably handle this end-to-end clinical action space in a realistic, EHR-style environment without committing unacceptable errors.
About the study
The study set out to address this functional gap by developing MIRA, a novel AI tool designed to autonomously ingest and access medical records, identify knowledge gaps, and order diagnostic tests to supplement the EHR record, before using the completed dataset to recommend clinical interventions.
The researchers then tested MIRA’s capabilities in a sandboxed, virtual EHR environment compliant with standard healthcare protocols, including HL7 FHIR. The sandboxed test was conducted on a curated benchmarking dataset of 574 real-world emergency department cases from the Medical Information Mart for Intensive Care (MIMIC-IV) database.
The cases included spanned eight distinct diagnoses across surgery (appendicitis), internal medicine (pneumonia), and oncology (pancreatic cancer), which MIRA navigated using 11 specialised digital tools offering more than 85,000 operational choices. The agent was permitted to request physical examinations, order targeted laboratory values, look up medical histories, and generate medication orders within the simulated EHR, rather than in live patient care.
How MIRA was compared with clinicians
MIRA’s output was compared against two distinct groups of human physicians managing exactly the same cases under identical conditions. The first group was a cohort of four board-certified physicians. The second was a mixed-seniority team consisting of four residents and two board-certified doctors.
A separate, conventional text-based AI agent was used to simulate the people under MIRA’s care, and under the care of the human physician teams. This agent was instructed to respond to questions posed by MIRA or its human counterparts solely on the basis of authentic clinical histories, while resisting adversarial attempts to trick it into prematurely leaking information. The authors noted, however, that simulated patient speech may be more structured than real emergency department conversations.
Study findings
The results revealed that MIRA performed at or above the level of experienced human doctors. It achieved 88.9% diagnostic accuracy across the full 574-case dataset and 87.8% accuracy in the matched 311-case physician comparison. By comparison, the board-certified physicians reached an average accuracy of 78.1% (p < 0.001), while the mixed-seniority medical cohort averaged 71.1% (p < 0.001).
MIRA was found to excel at identifying appendicitis and pancreatitis, achieving a perfect 100% recall for laparoscopic appendectomies. For pancreatic cancer, its diagnostic performance was equivalent to that of the board-certified physicians, while pneumonia and urinary tract infections remained more challenging.
Accuracy without simply “ordering everything”
Notably, MIRA did not achieve its superior accuracy by simply “ordering everything”. While it was observed to request a broader, more comprehensive set of individual blood parameters than the human doctors, its overall test selection remained well below the historical baselines recorded in the dataset.
The findings further demonstrated that the model successfully avoided the systematic over-ordering of high-cost radiological imaging, matching or exceeding physicians on overall resource-alignment metrics.
Safety performance
The safety evaluations were similarly encouraging, though still preliminary. An independent, blinded medical review of 56 patient-level outputs, together with a separate assessment of 468 prescriptions written by MIRA, established that the agent caused zero high-severity drug–drug interactions, zero renal dosing incompatibilities, and zero medication-allergy mismatches. Route specification was the weakest prescription field, at 97% correctness.
When making critical hospital admission decisions for pneumonia and pulmonary embolism, MIRA achieved a perfect recall score of 1.00, indicating that it never missed a single person who required inpatient care. The pulmonary embolism analysis did, however, suggest a tendency towards over-admission, reflecting a cautious disposition strategy.
Conclusions
The study introduces an integrated EHR AI agent, MIRA, that successfully translates clinical intents into structured, safe, and accurate operations, with the potential to support physicians in their work. The authors are careful to caution, however, that MIRA and similar AI agents are not replacements for expert human staff.
The model did not reach 100% perfection across all treatment choices, such as specific antibiotic selections, which highlights the ongoing need for strict human supervision and patient-level safeguards. Future iterations of the model may improve their performance by incorporating evidence from retrieval-based support, stronger governance, and prospective real-world validation before any clinical deployment.
Read More
AI Model Improves Early Detection of Serious Lung Disease in Newborns
Key Takeaways:
- Researchers at the University of Rochester have developed a time-series AI machine learning model that predicts bronchopulmonary dysplasia (BPD) in premature newborns more accurately than existing prediction tools.
- The model uses detailed electronic health record data rather than the limited datasets used in many current online BPD calculators.
- Researchers hope the technology could eventually support real-time clinical decision-making in neonatal intensive care units and help reduce the severity of lung disease in vulnerable infants.
AI and neonatal care
A research team from the University of Rochester has developed a new artificial intelligence and machine learning model designed to improve the prediction of bronchopulmonary dysplasia (BPD), a serious lung disease that affects premature newborns. Their findings were published in The Journal of Pediatrics in a study titled “Time-Series Machine Learning for Prediction of Bronchopulmonary Dysplasia.”
The project centres on the use of time-series machine learning, an approach that analyses patterns in data collected over time, allowing researchers to build more dynamic and potentially more accurate disease prediction models.
Understanding bronchopulmonary dysplasia
Bronchopulmonary dysplasia is a chronic lung condition that primarily affects babies born prematurely and with low birth weight. Because their lungs are still underdeveloped, many premature infants require oxygen therapy and mechanical ventilation shortly after birth to survive. However, this early exposure to oxygen and ventilatory support can contribute to lung injury and long-term respiratory complications.
Children who develop BPD may experience ongoing breathing difficulties and other long-term health challenges linked to impaired lung development.
“We take great effort in the neonatal intensive care unit to prevent lung damage,” said Associate Professor Andrew Dylag, MD, from the Department of Pediatrics, Neonatology. “Despite this, premature infants still develop BPD. There are BPD ‘calculators’ on the internet that can predict the severity of lung disease while the baby is still in the hospital, but they use a very limited set of data.”
According to the research team, recently updated versions of these existing calculators demonstrated lower accuracy than earlier models, prompting the group to explore a different strategy.
“We thought that using more detailed data from the University’s electronic health record would improve disease predictions and allow us to pinpoint vulnerable times when we might be able to intervene to prevent lung disease in newborns,” Dylag said.
Funding supports new collaboration
To support the project, the researchers secured a 2023 Digital Health Seedling award from the Clinical and Translational Science Institute (CTSI).
“We hoped that CTSI could help us test the hypothesis that machine learning could improve disease predictions in hospitalized premature newborns,” Dylag said. “The 2023 Digital Health Seedling award was exactly the type of funding we needed to develop new collaborations across the University community and kickstart our team’s academic interactions.”
The funding enabled the formation of a multidisciplinary research team combining expertise from neonatology, engineering, computer science, biostatistics and health informatics.
The collaboration included Jiebo Luo, PhD, Albert Arendt Hopeman Professor of Engineering in the Department of Computer Science, who connected graduate students to the project, as well as Professor Xing Qiu, PhD, from the Department of Biostatistics and Computational Biology.
“The neonatology team brought content and clinical expertise to the work, the computer science team developed and tested the models, and the biostatisticians ensured the rigor and testing of the models and algorithms,” Dylag said.
The project also expanded on an existing partnership with the University of Rochester Clinical and Translational Science Institute’s Informatics and Analytics group, particularly in relation to electronic health record research.
Building a secure AI research environment
Because the project relied on a very large dataset containing sensitive patient information, the research required substantial data security and privacy protections.
“We initially got involved by helping the study team pull clinical data from eRecord,” said Jack Chang, PhD, associate director of Research Informatics. “Recognizing the project involved a very large patient population and massive data including Protected Health Information, we identified a need for a more secure analytical workspace.”
To address these concerns, the Informatics team transitioned the data into the Secure Environment for Research Data Analytics (SERDA), a protected platform designed for high-risk clinical research and advanced analytics.
High-risk healthcare data projects typically involve extensive administrative oversight and cybersecurity requirements to ensure patient privacy and regulatory compliance.
“SERDA removes that obstacle by providing a secure, scalable, and ‘ready-to-go’ environment tailored for advanced analytics and machine learning,” Chang said. “It allows researchers to focus on their science while knowing their data is protected and compliant with all privacy regulations.”
Chang said the Informatics team worked closely with institutional partners to build the infrastructure required for the study.
“Our team—working with our ISD partners—handled the technical heavy lifting: setting up virtual machines, configuring project shares, ensuring secure access with the security team, and customizing the environment with specialized analytical software,” Chang said. “We also provided training and facilitated the numerous exports of analytical outcomes from SERDA for their publication.”
Towards real-time clinical decision support
Researchers believe the improved AI model may help clinicians identify which infants are most likely to benefit from early intervention and more precise treatment strategies.
The long-term goal is to integrate the model into clinical decision support systems capable of updating disease risk predictions in real time as patient conditions evolve.
“We want to build clinical decision support tools to identify how disease predictions change in real time,” Dylag said. “If we validate our algorithm and can present the disease prediction to the clinical team, we can test guideline implementation for how to manage or treat infants that may reduce BPD severity.”
The research team said that continued collaboration between departments, alongside infrastructure support from CTSI and ISD, will be essential as the project progresses into its next phase of development and validation.
Source: University of Rochester Medicine
Read More
AI Outperformed Emergency Doctors in Harvard Triage Study, Raising Questions About the Future of Clinical Decision-Making
Key Takeaways:
- A Harvard-led study found that an AI reasoning model outperformed emergency physicians in diagnosing patients during hospital triage scenarios using text-based clinical information.
- Researchers said the findings represent a major advance in AI clinical reasoning, although they stressed that AI is not ready to replace human doctors.
- Experts warned that important concerns remain around accountability, bias, safety, and the risk of clinicians becoming overly reliant on AI systems.
AI shows strong performance in emergency medicine trial
From fictional emergency department heroes such as George Clooney in ER to Noah Wyle in The Pitt, emergency physicians have long been portrayed as the ultimate decision-makers in moments of medical crisis. However, a new Harvard study suggests artificial intelligence may increasingly play a major role in those same high-pressure situations.
Researchers from Harvard Medical School and Beth Israel Deaconess Medical Center found that advanced AI systems outperformed human doctors in emergency medicine triage scenarios, making more accurate diagnoses when presented with limited patient information during the critical early stages of hospital admission.
The findings, published in Science, were described by independent experts as representing “a genuine step forward” in AI clinical reasoning.
According to the study authors, large language models (LLMs) “have eclipsed most benchmarks of clinical reasoning”.
AI versus doctors in emergency room triage
One of the study’s central experiments examined 76 patients who presented to the emergency department of a Boston hospital.
Both the AI system and pairs of human physicians were given identical electronic health record information to assess. This included standard triage details such as:
- Vital signs
- Demographic information
- Brief nursing notes explaining why the patient attended hospital
Using only this text-based information, OpenAI’s o1 reasoning model identified the exact diagnosis or a very close diagnosis in 67% of cases.
By comparison, the human physicians achieved diagnostic accuracy rates of between 50% and 55%.
Researchers found the AI’s advantage was especially apparent in triage situations requiring rapid decision-making with minimal available information.
When additional clinical detail was provided, the AI’s diagnostic accuracy increased further to 82%. Human experts achieved accuracy rates between 70% and 79% under those circumstances, although researchers noted the difference was not statistically significant in that setting.
AI also performed better in treatment planning
The study also evaluated how AI performed in longer-term clinical planning tasks.
In this experiment, the AI system and a group of 46 doctors were asked to review five detailed clinical case studies and develop treatment strategies. These included decisions relating to:
- Antibiotic regimens
- Ongoing management plans
- End-of-life care processes
The AI significantly outperformed the doctors.
Researchers reported that the AI achieved a score of 89%, compared with 34% among physicians using conventional resources such as search engines.
AI detected a diagnosis human doctors missed
One example highlighted in the study involved a patient with a pulmonary embolism and worsening symptoms.
Human doctors believed the patient’s anticoagulant treatment was failing. However, the AI system identified something clinicians had overlooked – the patient had a history of lupus, which may have been responsible for inflammation in the lungs.
The AI’s interpretation was ultimately confirmed as correct.
Researchers say AI will reshape medicine – not replace doctors
Despite the strong performance shown by AI systems, researchers stressed that the technology is not ready to replace physicians.
The study only assessed AI systems using text-based patient information. It did not evaluate the AI’s ability to interpret non-verbal clinical signals that doctors routinely use during patient assessment, such as:
- Visible distress
- Facial appearance
- Behaviour
- Physical examination findings
As a result, researchers said the AI functioned more like a clinician reviewing paperwork and offering a second opinion.
“I don’t think our findings mean that AI replaces doctors,” said Arjun Manrai, one of the lead authors of the study who heads an AI lab at Harvard Medical School. “I think it does mean that we’re witnessing a really profound change in technology that will reshape medicine.”
Dr Adam Rodman, another lead author and physician at Beth Israel Deaconess Medical Center, described AI LLMs as among “the most impactful technologies in decades”.
He suggested healthcare may move towards what he described as a “triadic care model”.
“Over the next decade,” Rodman said, AI would not replace physicians but instead work alongside them in a new model involving “the doctor, the patient, and an artificial intelligence system”.
AI use in healthcare is already growing
The findings come amid rapidly increasing AI adoption within healthcare systems.
According to research published last month, nearly one in five physicians in the United States are already using AI to assist with diagnosis.
In the United Kingdom, a recent Royal College of Physicians survey found:
- 16% of doctors use AI daily
- A further 15% use AI weekly
- Clinical decision-making is among the most common applications
However, concerns around safety and accountability remain significant.
UK doctors surveyed identified AI errors and legal liability as among their biggest worries.
“There is not a formal framework right now for accountability,” said Rodman.
He also stressed the continuing importance of human clinicians in patient care.
“Patients ultimately want humans to guide them through life or death decisions [and] to guide them through challenging treatment decisions,” he said.
Experts warn against over-reliance on AI
Independent experts said the study highlights the rapidly improving capabilities of AI systems in medicine, but also demonstrates the need for caution.
Prof Ewen Harrison, co-director of the University of Edinburgh’s Centre for Medical Informatics, said the findings suggest AI systems are beginning to evolve beyond theoretical testing environments.
“These systems are no longer just passing medical exams or solving artificial test cases,” he said. “They are starting to look like useful second-opinion tools for clinicians, particularly when it is important to consider a wider range of possible diagnoses and avoid missing something important.”
However, Dr Wei Xing from the University of Sheffield warned that the study also raised concerns about how clinicians interact with AI recommendations.
He suggested some doctors may unconsciously defer to AI-generated answers instead of independently evaluating clinical information themselves.
“This tendency could grow more significant as AI becomes more routinely used in clinical settings,” he said.
Dr Xing also noted that the study provided limited information about where the AI may perform less effectively, including whether diagnostic accuracy differed among certain patient populations such as:
- Older adults
- Non-English speakers
- People with more complex communication needs
He cautioned against interpreting the findings as evidence that publicly available AI tools are ready for independent medical use.
“It does not demonstrate that AI is safe for routine clinical use, nor that the public should turn to freely available AI tools as a substitute for medical advice,” he said.
Source: The Guardian
Read More
Large Language Models Show Promise in Detecting Drug Safety Signals from Clinical Notes
Key Takeaways:
- Large language models can identify immune-related adverse events in clinical notes without task-specific training, offering a potential alternative to labour-intensive manual review
- Performance remains below the threshold required for clinical decision support, with models tending to overpredict adverse events
- Despite limitations, this approach may support large-scale safety monitoring and accelerate research into cancer immunotherapies
The challenge of detecting drug safety signals
Drug safety signals are often embedded within unstructured clinical text, particularly in electronic health records. Identifying these signals has traditionally required either manual chart abstraction, which is resource-intensive, or natural language processing systems tailored to specific drugs and healthcare settings.
This challenge is particularly evident in the case of immune checkpoint inhibitors. These cancer therapies, first introduced in 2011, are associated with a broad range of immune-related adverse events. These events can affect multiple organ systems, including the colon, liver, lungs, heart, nervous system, skin, and endocrine system, making systematic detection complex and time-consuming.
Exploring large language models as a solution
Large language models are increasingly being explored as a way to streamline the identification of drug safety signals within clinical text. A multicentre study, published in eBioMedicine, evaluated whether these models could detect immune-related adverse events associated with immune checkpoint inhibitors.
The study focused on a zero-shot learning approach. In this setting, the model receives a single, detailed prompt without prior examples. The prompt used by the researchers began: “You are a clinical expert in identifying immune-related adverse events caused by immune checkpoint inhibitors …” and included a list of six immune checkpoint inhibitors alongside numerous associated adverse events.
This prompt was applied to clinical notes from multiple sources. These included records from 100 people treated at Vanderbilt Health, 70 people from the University of California, San Francisco, and 272 people enrolled in seven Roche-sponsored clinical trials.
Study design and model performance
The research team evaluated three models: GPT-3.5, GPT-4, and GPT-4o, with GPT-4o demonstrating the strongest overall performance.
To assess accuracy, the investigators used F1 scores, a metric that balances false positives and false negatives. Scores range from zero to one, with values above 90 percent considered excellent. A score of 80 percent or higher may be sufficient for use in automated clinical decision support systems.
At the patient level, GPT-4o achieved average F1 scores of 56 percent for Vanderbilt Health data, 66 percent for University of California, San Francisco data, and 62 percent for Roche clinical trial data. The models showed a consistent tendency to overpredict the presence of immune-related adverse events.
When analysing individual clinical notes, the model achieved an average F1 score of 57 percent across 667 notes from Vanderbilt Health, evaluating 17 different adverse events.
Implications for clinical practice and research
The findings suggest that large language models can play a role in identifying drug safety signals, even without task-specific training data.
“Manual patient chart abstraction for monitoring the safety and efficacy of drugs already at market requires tremendous resources and puts a drag on the pace of discovery in precision medicine. And that’s especially true with immune checkpoint inhibitors, where the adverse events are so varied. If zero-shot learning with LLMs could help with these notes, it could significantly reduce time and costs for all concerned,” said the report’s corresponding author, Cosmin Bejan, PhD, assistant professor of Biomedical Informatics at Vanderbilt Health.
However, the current level of performance falls short of what would be required for clinical decision support.
“These results show that zero-shot learning with a powerful LLM is useful for detecting these adverse events,” Bejan said. “This performance does not rise to the level required for clinical decision support, but the method could be valuable for automated irAE extraction across multiple sites, potentially speeding discovery and enhancing the safety and effectiveness of cancer immunotherapies.”
Wider research context
The study involved collaboration among multiple researchers at Vanderbilt Health, including Yaomin Xu, PhD, Eric Mukherjee, MD, PhD, Matthew Krantz, MD, Douglas Johnson, MD, MSCI, Elizabeth Phillips, MD, and Justin Balko, PhD. Funding support was provided in part by the National Institutes of Health.
Related research further highlights safety concerns associated with immune checkpoint inhibitors. In a research letter published in JAMA Oncology, Mukherjee, Phillips, and colleagues used logistic regression analysis of adverse event reports from the Food and Drug Administration. They confirmed that these therapies are independently associated with an increased risk of Stevens-Johnson syndrome and toxic epidermal necrolysis, which are severe and potentially life-threatening skin reactions. The study also found that this risk may be linked to exposure to human leukocyte antigen–restricted drugs.
Conclusion
Large language models represent a promising tool for extracting clinically meaningful insights from unstructured health data. While their current performance limits direct clinical application, their ability to operate across multiple datasets without task-specific training suggests potential for supporting large-scale pharmacovigilance efforts. As these models continue to improve, they may contribute to more efficient and comprehensive monitoring of drug safety in clinical practice.
Read More
‘Shadow AI’ on the Rise in Healthcare as Clinicians Turn to Unauthorised Tools to Improve Workflows
Key Takeaways:
- A survey of healthcare professionals found that 57% have encountered or used unauthorised artificial intelligence tools in their organisations, highlighting the growing presence of so-called “shadow AI” in healthcare settings.
- Many clinicians and administrators report using these tools to improve efficiency, analyse data, and manage administrative tasks, particularly when approved solutions or clear guidance are lacking.
- While most respondents believe AI will significantly improve healthcare within five years, concerns about patient safety, data privacy, and security risks remain widespread.
Unauthorised AI tools emerging in healthcare workplaces
A new survey suggests that artificial intelligence tools are already being used in healthcare organisations in ways that fall outside formal governance structures. According to the findings, a significant proportion of healthcare professionals have either encountered or used AI tools that have not been authorised by their employer.
The survey, conducted by Wolters Kluwer Health, gathered responses from 518 healthcare professionals, including both clinical providers and administrators. The research was carried out in December 2025 and was released publicly last week.
Overall, the findings indicate that four in ten healthcare professionals reported encountering unauthorised AI tools within their organisation, while 17% acknowledged personally using such tools.
When responses were analysed by professional role, 15% of physicians admitted to using an unauthorised AI tool, compared with 19% of administrators. In addition, one in ten respondents reported using an unauthorised AI tool in connection with direct patient care.
The report refers to the unauthorised adoption of artificial intelligence tools in professional environments as “shadow AI.”
Why healthcare staff turn to unauthorised AI
The survey findings suggest that healthcare professionals are often motivated by practical needs rather than deliberate attempts to bypass organisational policies.
According to the report:
“Clinical and administrative teams want to adhere to rules surrounding AI usage, but if the organization hasn’t provided guidance or approved solutions, they’ll experiment with generic tools to improve their workflows.”
Many respondents indicated that the absence of formal guidance or approved AI platforms has encouraged individuals to explore publicly available tools on their own.
The most frequently cited motivation for using unauthorised AI tools was the need to accelerate workflows and improve efficiency. Approximately half of respondents identified faster workflows as the primary reason for using these tools.
However, the survey also revealed differences in how clinical and administrative staff tend to use AI technologies.
Administrators were more likely to employ AI tools for operational or analytical tasks such as:
- Data analysis
- Predictive analytics
- Administrative processes
Healthcare providers, meanwhile, reported using AI for activities such as:
- Data analysis
- Patient scheduling
- Patient engagement tasks
The findings also indicate that clinicians were more likely than administrators to experiment with AI tools out of curiosity.
Governance and policy development remain uneven
The survey results highlight a notable imbalance in how different professional groups participate in the development of AI policies within healthcare organisations.
According to the report, administrators were three times more likely than clinical providers to be actively involved in developing AI governance policies.
Specifically:
- 30% of administrators reported involvement in AI policy development
- Only 9% of providers said they had participated in such efforts
This difference suggests that policy ownership around AI adoption may currently be concentrated within administrative leadership rather than clinical teams.
Administrators also reported greater familiarity with their organisation’s AI policies compared with providers, although awareness varied across both groups.
Security and privacy risks associated with “shadow AI”
The use of unauthorised AI tools raises important concerns about data security, privacy protection, and governance oversight.
The Wolters Kluwer report notes that inconsistent or unsanctioned AI usage can expose organisations to potential vulnerabilities. Without clear oversight, the integration of external AI tools may lead to data privacy violations, security breaches, or inappropriate handling of sensitive information.
To illustrate these risks, the report references a 2025 study by IBM, which found that 97% of organisations that experienced an AI-related security incident lacked adequate AI access controls.
Security incidents involving AI systems can have significant consequences, including financial losses, operational disruption, and damage to public trust.
Healthcare professionals remain optimistic about AI’s future
Despite concerns about governance and security, the survey indicates that most healthcare professionals remain broadly optimistic about the long-term role of artificial intelligence in healthcare.
Nearly 90% of respondents said they believe AI will significantly improve healthcare within the next five years. Administrators were found to be slightly more optimistic than clinical providers about the potential benefits of the technology.
At the same time, respondents recognised that AI implementation carries important risks that must be addressed.
Patient safety was identified by around half of respondents as the most significant risk associated with AI adoption.
Meanwhile, nearly half of respondents also expressed concerns about data privacy risks.
These findings suggest that healthcare professionals recognise both the transformative potential of artificial intelligence and the need for careful governance, clear guidance, and secure systems.
Addressing the rise of “shadow AI”
The report concludes that addressing the growth of shadow AI requires organisations to understand why staff are turning to unauthorised tools rather than focusing solely on restricting access.
According to the report:
“Ultimately, addressing shadow AI is not about restricting access to productivity tools. Leaders must understand why teams are using unsanctioned tools and which challenges they’re trying to solve, and then identify enterprise-level tools that can accomplish these goals safely and securely.”
As artificial intelligence becomes increasingly embedded in healthcare workflows, organisations may need to develop clearer policies, provide approved tools, and involve both clinical and administrative staff in governance decisions.
Such measures may help ensure that the benefits of AI can be realised while protecting patient safety, safeguarding sensitive data, and maintaining organisational trust.
Read More