
Adapting the Machine to the Medicine: New Review Shows How to Make Clinical AI Work
Key Takeaways:
- A new review in the Journal of Medical Internet Research analysed 35 recent studies and found that off-the-shelf large language models must be adapted for medical settings before they can be relied upon for diagnosis, triage and treatment planning.
- Performance depends on the clinical task: retraining the model on task-specific data suited narrow jobs such as detecting cancer in medical images, while connecting the model to live, trusted databases worked well for reasoning through complex guidelines. Hybrid systems combining both performed best of all.
- Most of the evidence to date comes from historical medical records rather than live clinical use, so prospective, real-world testing remains essential before widespread hospital adoption.
Why raw capability is not enough
Generative artificial intelligence has arrived in health care with considerable fanfare, but a new review suggests that the decisive factor is not the sheer power of the underlying model. It is what happens after the model is built. A review study published in the Journal of Medical Internet Research by JMIR Publications concludes that adapting existing language models to the clinical environment is the key to ensuring they can safely and effectively support diagnosis, patient triage and treatment planning in real-world settings.
The research team, led by Anshum Patel, MD, and Joseph Y Cheung, MD, MS, analysed 35 recent studies to understand how different customisation methods affect AI performance. Their conclusion was that standard large language models (LLMs) are undoubtedly capable, but capability alone does not translate into clinical reliability. To be trusted in a consulting room or on a ward, a model has to be tailored specifically for the medical environment in which it will be used.
What adaptation actually looks like
The review examined the two broad routes clinicians and developers can take. The first is retraining, in which a general-purpose model is further trained on specific clinical data or guidelines so that it internalises the patterns and standards of a particular task. The second is connection, in which the model is linked directly to trusted medical databases and draws on that source material at the point of use rather than relying solely on what it absorbed during training.
Both approaches produced meaningful gains. When models were connected directly to trusted medical databases or retrained on specific clinical guidelines, accuracy greatly improved, with some systems matching the diagnostic performance of human doctors. That is a striking benchmark, and it underlines the central argument of the review: the adaptation step is not a technical afterthought but the point at which a general tool becomes a clinical one.
The right method depends on the task
The most successful approach was found to depend heavily on the specific medical task at hand, and this is arguably the review’s most practical finding for anyone evaluating these tools.
For narrow, focused tasks such as detecting cancer in medical images, retraining the AI on specific data worked best. Where the job is well defined and the data are consistent, teaching the model directly on that material produces the sharpest results.
For tasks that require reasoning through complex guidelines, linking the AI to live databases was highly effective. Clinical guidance changes, and a model that consults an authoritative, current source is far better placed than one working from a fixed snapshot of its training data.
However, the researchers determined that the best performance came from hybrid systems, which combine both methods to manage complicated workflows such as stroke triage and oncology cases. These are precisely the situations in which speed, protocol adherence and nuanced judgement all matter at once, and where people presenting with time-critical or complex conditions stand to gain the most from well-designed decision support.
“There is no single best way to adapt AI for health care. The right approach depends on the clinical task, and the next step is making sure these systems are safe, reliable, and useful in real-world patient care,” says Anshum Patel.
Implications for clinical teams
For healthcare professionals, the message is not that they need to become engineers. It is that the questions worth asking about any AI tool offered to a service are increasingly practical ones: what was this model adapted for, what source material informs its outputs, how current is that material, and has it been tested on a population and a workflow resembling ours?
That kind of informed scrutiny is becoming part of everyday clinical literacy, and it is one reason structured professional development in this area has grown in demand. The College of Contemporary Health’s AI Essentials for Primary Care short course is designed for exactly this purpose, helping clinicians understand how these systems are built, where their limitations lie, and how to appraise them responsibly within their own practice.
The evidence gap that still needs closing
While these findings are promising, the researchers note that the vast majority of studies on these AI systems are based on past medical records rather than live patient testing. Retrospective analysis can demonstrate that a system performs well against a tidy historical dataset; it cannot show how that system behaves when confronted with incomplete notes, atypical presentations, or the ordinary time pressures of a busy department.
Before these advanced tools are widely adopted in hospitals, additional prospective, real-world testing is needed to guarantee patient safety and ensure the technology works reliably across different clinical environments. Until that evidence accumulates, the sensible position is one of informed optimism – recognising the genuine potential of adapted clinical AI while insisting that it earns its place through testing in the settings where people actually receive care.
CCH insight
Artificial intelligence is moving quickly from conference agendas into everyday clinical workflows, and the professionals best placed to use it well are those who understand both its strengths and its blind spots. AI Essentials for Primary Care is a CPD-accredited short course from The College of Contemporary Health, built for clinicians who want a clear, practical grounding in how AI tools work, how to evaluate them critically, and how to apply them safely in patient-facing practice.
Explore AI Essentials for Primary Care →
Source: Journal of Medical Internet Research




