
Large Language Models Show Promise in Detecting Drug Safety Signals from Clinical Notes
Key Takeaways:
- Large language models can identify immune-related adverse events in clinical notes without task-specific training, offering a potential alternative to labour-intensive manual review
- Performance remains below the threshold required for clinical decision support, with models tending to overpredict adverse events
- Despite limitations, this approach may support large-scale safety monitoring and accelerate research into cancer immunotherapies
The challenge of detecting drug safety signals
Drug safety signals are often embedded within unstructured clinical text, particularly in electronic health records. Identifying these signals has traditionally required either manual chart abstraction, which is resource-intensive, or natural language processing systems tailored to specific drugs and healthcare settings.
This challenge is particularly evident in the case of immune checkpoint inhibitors. These cancer therapies, first introduced in 2011, are associated with a broad range of immune-related adverse events. These events can affect multiple organ systems, including the colon, liver, lungs, heart, nervous system, skin, and endocrine system, making systematic detection complex and time-consuming.
Exploring large language models as a solution
Large language models are increasingly being explored as a way to streamline the identification of drug safety signals within clinical text. A multicentre study, published in eBioMedicine, evaluated whether these models could detect immune-related adverse events associated with immune checkpoint inhibitors.
The study focused on a zero-shot learning approach. In this setting, the model receives a single, detailed prompt without prior examples. The prompt used by the researchers began: “You are a clinical expert in identifying immune-related adverse events caused by immune checkpoint inhibitors …” and included a list of six immune checkpoint inhibitors alongside numerous associated adverse events.
This prompt was applied to clinical notes from multiple sources. These included records from 100 people treated at Vanderbilt Health, 70 people from the University of California, San Francisco, and 272 people enrolled in seven Roche-sponsored clinical trials.
Study design and model performance
The research team evaluated three models: GPT-3.5, GPT-4, and GPT-4o, with GPT-4o demonstrating the strongest overall performance.
To assess accuracy, the investigators used F1 scores, a metric that balances false positives and false negatives. Scores range from zero to one, with values above 90 percent considered excellent. A score of 80 percent or higher may be sufficient for use in automated clinical decision support systems.
At the patient level, GPT-4o achieved average F1 scores of 56 percent for Vanderbilt Health data, 66 percent for University of California, San Francisco data, and 62 percent for Roche clinical trial data. The models showed a consistent tendency to overpredict the presence of immune-related adverse events.
When analysing individual clinical notes, the model achieved an average F1 score of 57 percent across 667 notes from Vanderbilt Health, evaluating 17 different adverse events.
Implications for clinical practice and research
The findings suggest that large language models can play a role in identifying drug safety signals, even without task-specific training data.
“Manual patient chart abstraction for monitoring the safety and efficacy of drugs already at market requires tremendous resources and puts a drag on the pace of discovery in precision medicine. And that’s especially true with immune checkpoint inhibitors, where the adverse events are so varied. If zero-shot learning with LLMs could help with these notes, it could significantly reduce time and costs for all concerned,” said the report’s corresponding author, Cosmin Bejan, PhD, assistant professor of Biomedical Informatics at Vanderbilt Health.
However, the current level of performance falls short of what would be required for clinical decision support.
“These results show that zero-shot learning with a powerful LLM is useful for detecting these adverse events,” Bejan said. “This performance does not rise to the level required for clinical decision support, but the method could be valuable for automated irAE extraction across multiple sites, potentially speeding discovery and enhancing the safety and effectiveness of cancer immunotherapies.”
Wider research context
The study involved collaboration among multiple researchers at Vanderbilt Health, including Yaomin Xu, PhD, Eric Mukherjee, MD, PhD, Matthew Krantz, MD, Douglas Johnson, MD, MSCI, Elizabeth Phillips, MD, and Justin Balko, PhD. Funding support was provided in part by the National Institutes of Health.
Related research further highlights safety concerns associated with immune checkpoint inhibitors. In a research letter published in JAMA Oncology, Mukherjee, Phillips, and colleagues used logistic regression analysis of adverse event reports from the Food and Drug Administration. They confirmed that these therapies are independently associated with an increased risk of Stevens-Johnson syndrome and toxic epidermal necrolysis, which are severe and potentially life-threatening skin reactions. The study also found that this risk may be linked to exposure to human leukocyte antigen–restricted drugs.
Conclusion
Large language models represent a promising tool for extracting clinically meaningful insights from unstructured health data. While their current performance limits direct clinical application, their ability to operate across multiple datasets without task-specific training suggests potential for supporting large-scale pharmacovigilance efforts. As these models continue to improve, they may contribute to more efficient and comprehensive monitoring of drug safety in clinical practice.
Read More
AI in Healthcare: Promise, Pitfalls, and the Risk of Misguided Medical Advice
Key Takeaways:
- People using AI for health advice often struggle to interpret and communicate symptoms effectively, leading to incorrect conclusions in many cases.
- Even when AI identifies a condition correctly, it may fail to recommend appropriate urgency, particularly in time-sensitive or complex scenarios.
- Clinicians see value in AI as a supportive tool, but stress that it should complement, not replace, professional medical care.
AI becomes a common source of health information
As technology companies continue to develop platforms tailored for healthcare consultation, artificial intelligence is becoming an increasingly influential part of how people make decisions about their health. According to OpenAI, more than 40 million people use ChatGPT each day to seek health-related information.
However, emerging research suggests that while these tools offer unprecedented access to medical knowledge, they may also mislead users in certain contexts.
Challenges in how people use AI for medical queries
One of the central issues identified by researchers is not only the capability of AI systems, but how individuals interact with them. Many people lack the knowledge required to communicate symptoms accurately or comprehensively.
A recent study published in Nature Medicine attempted to replicate real-world use of AI chatbots. Participants were given medical scenarios and asked to consult AI tools. The results highlighted notable limitations:
- Participants correctly identified the condition only about one-third of the time.
- Just 43% made the correct decision regarding next steps, such as whether to seek emergency care or remain at home.
“People don’t know what they are supposed to be telling the model,” says Andrew Bean, who studies AI systems at Oxford University and was one of the authors on this study.
Bean explains that effective use of AI often depends on precise wording. “Doctors are trained to ask you questions about symptoms you might not have realised you should have mentioned,” says Bean.
Small differences in input can lead to dangerous outcomes
The study demonstrated how subtle differences in language can significantly alter the advice provided by AI systems.
In one example, two individuals described the same clinical scenario slightly differently. One described experiencing “the worst headache I’ve ever had” and was advised to go to the emergency room immediately. The other, who did not include that specific phrasing, was advised to take aspirin and remain at home.
“Turns out this was actually a life-threatening condition,” says Bean.
This highlights a critical limitation: AI systems rely heavily on the information they are given, and may not prompt for missing but clinically important details in the way a trained clinician would.
When AI gets the diagnosis right but the advice wrong
Even when AI tools successfully identify a medical condition, they may still provide inappropriate guidance regarding urgency or next steps.
In a separate study, researchers evaluated how AI systems responded to a range of medical scenarios. They found that in 52% of emergency cases, the tools “under-triaged” – treating conditions as less serious than they actually were.
In one case, the AI failed to direct a hypothetical patient experiencing diabetic ketoacidosis and impending respiratory failure – both life-threatening conditions – to seek emergency care.
“When there was a textbook medical emergency, ChatGPT got it right,” said Girish Nadkarni, a doctor and AI researcher at Mount Sinai who is an author on the study. However, he noted that performance declined in more complex situations, particularly where timing was critical. In such cases, the system often misjudged how urgently care was required.
An OpenAI spokesperson responded by stating that the study did not reflect typical real-world usage and that it evaluated an older version of ChatGPT, which the company says has since been improved to address some of these concerns.
The role of AI in supporting patient understanding
Despite these concerns, many clinicians believe that AI tools can still play a constructive role in healthcare, particularly in improving patient understanding and engagement.
“I encourage patients to use these tools,” says Robert Wachter, a doctor at UC San Francisco and author of the recently published book, A Giant Leap: How AI Is Transforming Health Care and What That Means for Our Future.
Wachter points out that barriers to accessing healthcare – including cost and availability – mean that AI can sometimes provide a useful alternative source of information. “The advice you get from the tools is substantially better than nothing and better than what you would get from your second cousin,” says Wachter.
However, he emphasises that AI should never be viewed as a substitute for professional medical care.
Enhancing, not replacing, the doctor–patient relationship
Experts suggest that AI is most valuable when used alongside traditional healthcare, rather than in place of it.
Adam Rodman, a hospitalist and researcher at Harvard Medical School, advises against using AI tools to assess emergency situations. Instead, he sees their greatest benefit in preparing for or reflecting on medical consultations.
“A good time to use a large language model is when you’re about to go see a doctor – or after you see your doctor,” says Rodman.
He explains that AI can help people better understand their condition, ask more informed questions, and make more effective use of time during appointments. This can support a more collaborative relationship between patients and clinicians.
“There are no downsides to better understanding your health,” says Rodman.
The future of AI in healthcare
Healthcare professionals broadly agree that AI is now firmly embedded within modern medicine and will continue to evolve alongside clinical practice.
“ My hope is that you might see AI as an extension of a human relationship,” says Rodman. He envisions a future in which both clinicians and patients work with AI to improve communication and navigate healthcare systems more efficiently.
However, he also raises concerns about potential unintended consequences. One particular risk is the possibility that people may receive serious or distressing diagnoses – such as cancer – directly from an AI system, rather than from a clinician.
Research suggests that when healthcare becomes more transactional or resembles a marketplace, trust in clinicians may decline.
”What I hope is that this technology can be used in a way that enhances humanity in medicine,” says Rodman “and not in a way that cuts out the doctor-patient relationship.”
Conclusion
Artificial intelligence is rapidly transforming access to health information, offering both opportunities and risks. While these tools can enhance understanding and support more informed decision-making, their limitations – particularly in how they interpret incomplete or imprecise input – mean they must be used with caution.
Ultimately, AI has the potential to strengthen healthcare delivery, but only if it is integrated in a way that supports, rather than replaces, the human relationships at the heart of medicine.
Read More