
Does AI Help You Reflect, or Just Write Faster?
By Nigel Hinchliffe Director of Education, CCH.
During a curriculum review workshop several years ago, a group of newly qualified nurses were asked to share a reflective account they were proud of. Most had written competent, structured pieces that moved neatly through Gibbs’ stages: description, feelings, evaluation, analysis, conclusion, action. When asked whether those accounts had changed anything about the way they practised, the room went quiet. They had learned to write reflection. They had not necessarily learned to reflect.
“They had learned to write reflection.
They had not necessarily learned to reflect.”
That gap, between the performance of reflection and reflection as a genuine habit of professional learning, is not a new problem. It is, however, an urgent one. Generative AI tools are now capable of producing plausible reflective writing at speed. Before the profession reaches for AI as a solution to reflective practice, it is worth being aware of what the problem actually is.
Has reflection been taught as a skill?
Healthcare professionals are required to reflect, assessed on their reflections, and held accountable for them at revalidation. The Nursing and Midwifery Council, for instance, requires five written reflective accounts as part of the revalidation process (NMC, 2019).
What is less common is being explicitly taught how to reflect: how to move from description into analysis, how to connect personal experience to evidence, how to sit with uncertainty rather than resolve it prematurely into a tidy learning point.
Gibbs’ Reflective Cycle (1988) is the dominant framework, precisely because its six stages – description, feelings, evaluation, analysis, conclusion, action planning – provide a scaffold practitioners can follow without having been taught the underlying skill. That is both its strength and its limitation. Applied well, Gibbs takes a practitioner from raw experience to transferable learning. Applied as a form to be completed, it produces the kind of account that satisfies an assessor without troubling the practitioner who wrote it.
This is the baseline problem AI enters. Not a profession of skilled reflectors who lack time, but a profession in which written reflection may well be delivered, but practitioners may not have learned how to reflect
What AI offers, and what it does not do automatically
Generative AI tools – Claude, ChatGPT, Copilot and others – are capable of functioning as structured thinking partners in a way that a blank page cannot. Used well, they can prompt practitioners through Gibbs’ stages with tailored questions, surface relevant clinical frameworks, and help translate half-formed thoughts into coherent prose.
For a nurse processing a safeguarding concern, or a GP working through a difficult prescribing decision, an AI interlocutor that asks “What assumptions were you making at that point, and what would challenge them?” can move the reflection further than a prompt to “complete your analysis section”.
But this is not automatic. A general AI model’s default behaviour is to affirm and reflect back.
Asked to help with a reflective account, it will typically smooth, summarise and endorse. Getting it to meaningfully challenge a clinician’s reasoning requires deliberate prompt design. That design is itself a competency, and one that many practitioners have yet to acquire. The AI needs to be asked to:
- Notice a self-serving account
- Probe an unexamined assumption
- Push back on a premature conclusion
There is also a legitimate concern about cognitive offloading. The effort of articulating an experience is an important aspect of reflective learning: finding the words, sitting with ambiguity, returning to an account and revising it. Moon’s theoretical work on reflection and experiential learning argues that this effortful processing is part of how meaning is constructed from experience, rather than a delivery mechanism for insight that exists independently of it (Moon, 2004).
The real risk: effortless performance
A 2024 systematic review of reflective writing as summative assessment found that power dynamics between students and markers can lead to ‘performative instead of genuine reflection’ and concluded that student voices in assessed work may not represent authentic participation ‘because students may lack agency as they are on the wrong side of a power imbalance and are motivated to pass’ (Ross, Bohlmann and Marren, 2024). The conventions of a passable reflective account, such as a growth narrative, appropriate emotional awareness, or a clear action point, are learnable independently of the reflection itself.
AI makes producing that performance effortless. A practitioner who has absorbed the conventions can generate a convincing Gibbs account of a clinical encounter they have not meaningfully examined, in minutes, with minimal cognitive effort.
AI will be used in reflective writing; it already is. The real question for educators and professional bodies is whether the assessment frameworks and supervisory practices we have built are robust enough to distinguish genuine reflection from sophisticated production. Many are not, which is a problem for education design to solve, not for the technology.
Where AI can make a real difference
None of this means AI has no role. It means the role needs to be designed rather than assumed.
For student practitioners, AI can function as a responsive scaffold – asking the questions a supervisor would ask, at the moment when the experience is still live rather than weeks later in a tutorial.
The immediacy matters: Mann, Gordon and MacLeod’s review identified proximity to the experience as one of the variables that influences the quality of reflective learning (2009). A tool that meets practitioners in that proximate moment, rather than waiting for a scheduled portfolio deadline, addresses a real structural gap.
A systematic review of how health-professions students use generative AI found these tools most often supported learning through inquiry and in practice, rather than simple content acquisition (Pham et al., 2025).
For experienced clinicians preparing for supervision or appraisal, AI offers a way to organise thinking before the formal conversation rather than during it, reducing the time spent on description and increasing what is available for deeper analysis. The preparation is the reflection; the supervision becomes its extension.
For practitioners processing emotionally charged events – errors, patient deterioration, moral distress – the non-judgemental quality of an AI interlocutor is not trivial. The psychological safety required for honest reflection is not always available in formal supervisory relationships, and a first pass engagement with an AI tool may lower the threshold for authentic disclosure in a subsequent human conversation.
Prompting as a professional skill
The common thread across these applications is that the value of AI in reflective practice depends almost entirely on how it is used. A tool set up to challenge will challenge; a tool set up to affirm will affirm. The distinction is determined by the practitioner’s prompting, and prompting is a skill.
“A tool set up to challenge will challenge.
A tool set up to affirm will affirm.”
That skill can be taught directly. Practitioners can learn to build challenge into their prompts in the same way they learn any other clinical or communication skill: through worked examples, supervised practice, and feedback. This might mean asking the AI to argue the counter-case, name the assumption they have not examined, or hold back the tidy conclusion until the uncertainty has been explored.
Embedding this in professional development does not require waiting for system-wide change. It can sit inside existing CPD structures:
- A short module on prompt design for reflective practice
- Worked examples built into portfolio guidance
- Supervisors modelling the technique in appraisal conversations.
The frameworks for reflective practice already exist; what is needed is training practitioners to use AI within them deliberately, rather than assuming the tool will do the work on its own.
The Gibbs cycle has endured for nearly four decades because its structure is sound. What it has always needed is a workforce able to use it as a habit of thinking rather than a form to complete. AI does not change that requirement. It raises the stakes on meeting it. Whether AI ends up supporting genuine reflection or accelerating its performance depends on whether that training happens now.
What this means in practice
When AI makes reflective writing easier, it does not automatically make reflective learning more likely. The outcome depends on whether practitioners learn to prompt it to challenge their thinking, rather than simply affirm it.
Colleges and providers who educate healthcare professionals can address this directly, rather than wait for the wider profession to resolve it. Teaching prompt design as part of CPD, portfolio guidance, and supervision is not a future ambition. It is available now.
How to Take This Further
If this is a skill you want to build, our short course, Getting More from ChatGPT: Prompting Skills for Healthcare Professionals, includes a prompt template for reflective practice specifically, as well as the course’s wider training in getting reliable, useful output from any AI tool you use.
Click here to learn more.
About the author
This article was written by Nigel Hinchliffe, Director of Education at the College of Contemporary Health (CCH). Nigel has extensive experience in clinical education, with a particular focus on how healthcare professionals develop and demonstrate competence throughout their careers. At CCH, he leads the development of evidence-based training programmes and regularly provides feedback on learner’s reflective submissions as part of their professional development.
• Gibbs G (1988) Learning by doing: a guide to teaching and learning methods. Oxford: Further Education Unit, Oxford Polytechnic.
• Mann K, Gordon J, MacLeod A (2009) Reflection and reflective practice in health professions education: a systematic review. Advances in Health Sciences Education. 14(4):595-621.
• Moon JA (2004) A handbook of reflective and experiential learning: theory and practice. London: RoutledgeFalmer.
• Nursing and Midwifery Council (2019) Revalidation: how to revalidate with the NMC. London: NMC. Available at: https://www.nmc.org.uk/revalidation/
• Pham TD, Karunaratne N, Exintaris B, Liu D, Lay T, Yuriev E, Lim A (2025) The impact of generative AI on health professional education: a systematic review in the context of student learning. Medical Education. 59(12):1280-1289.
• Ross M, Bohlmann J, Marren A (2024) Reflective writing as summative assessment in higher education: a systematic review. Journal of Perspectives in Applied Academic Practice. 12(1):54-67

People Like AI Mental Health Chatbots. Whether They Help Is Another Question
Key Takeaways:
- Across 21 studies in 11 countries, people using generative AI mental health chatbots reported high satisfaction and found them convenient and accessible.
- Personalisation and empathy did not reliably translate into better clinical outcomes, and engagement often faded over time.
- The evidence base remains early-stage, leaving safety, equity and crisis response unresolved.
A treatment gap that digital tools are being asked to fill
Generative artificial intelligence (GenAI) chatbots designed to support mental health are winning people over on experience, but the research needed to establish whether they are safe and clinically effective has not kept pace. That is the central finding of a review, currently in press in the journal npj Digital Medicine, which examined the user experience (UX) and intervention design of GenAI mental health chatbots.
The context for this work is a substantial and persistent shortfall in care. Around 25% of people worldwide experience a mental health problem, yet approximately 85% do not receive adequate treatment. The reasons are varied and overlapping: stigma, cost, shortages of trained professionals, geographic distance from services and structural inequities, among others. As the prevalence of mental health conditions has grown while treatment gaps have remained, attention has turned towards innovative models of delivery – and digital tools, with their scalability and convenience, have become a focus of that search.
Digital mental health interventions deliver treatment or support through a range of channels, including chatbots, websites, mobile applications and wearables. Conversational agents, more commonly known as chatbots, are applications that simulate human dialogue using machine learning and natural language processing algorithms.
From scripted responses to open-ended conversation
Traditional mental health chatbots deliver pre-scripted therapeutic content through rules-based or retrieval-based systems. Their strength is predictability, but that same design limits their capacity to personalise support or to recognise what an individual actually needs in the moment.
Chatbots built on large language models (LLMs) work differently. They can simulate core aspects of a therapeutic encounter, including personalised suggestions and empathetic reflections. That flexibility comes with a trade-off. GenAI systems may produce responses that are incorrect or inappropriate, and their open-ended conversational capacity makes intervention design both more consequential and more complex than it is for rules-based systems. When a system can say almost anything, design decisions carry considerably more weight.
How the review was carried out
The researchers set out to map the design characteristics and UX outcomes of interventions involving GenAI mental health chatbots. They began with a systematic literature search to identify studies covering the design and deployment of such tools. Reviews, editorials, media articles and commentaries were excluded.
In total, 21 studies were selected, conducted across 11 countries between 2023 and 2025. The largest numbers came from China and the United Kingdom, followed by the United States. The included work spanned a wide range of maturity, from early-stage prototype evaluations through to clinical trials, with one real-world implementation study.
Most studies recruited general or clinical adult populations, including older people living with dementia. Others involved simulated users or university students. Sample sizes ranged from as few as five participants to as many as 527. Across the body of evidence, there was substantial heterogeneity in outcome measures, and most interventions remained at an early stage of development – two features that shape how much can reasonably be concluded from the literature as it stands.
What the interventions were designed to do
Chatbot interventions most often targeted depression and anxiety, and tended to adopt shared therapeutic mechanisms, including mindfulness, emotion regulation and cognitive restructuring. Some systems were oriented towards mental well-being, stress and loneliness, emphasising general support and preventive care rather than treatment for a specific condition. Others addressed eating disorders, post-traumatic stress disorder and dementia.
The dementia-focused interventions are worth distinguishing. Rather than attempting to address the central neurological features of the condition, they targeted its related psychological dimensions – carer burnout, psychological distress and loneliness among them.
Most interventions were grounded in cognitive-behavioural therapy principles. The specific techniques drawn upon included behavioural activation, psychoeducation, Socratic dialogue, acceptance and commitment therapy, cognitive restructuring and mindfulness.
How the tools were delivered
Interventions varied in frequency, delivery modality and duration. The majority were short-term, running from two to eight weeks. Most were deployed through web-based interfaces and mobile applications, while some were delivered via messaging or social media platforms – meeting people on services they already used rather than asking them to adopt something new.
About 67% of interventions were non-embodied, text-based chatbots. The remainder used voice, avatar-based, augmented reality or other multimodal forms of interaction, with the intention of improving engagement and realism.
What people made of them
All but two of the studies evaluated at least one UX domain. The majority relied on quantitative measures, typically Likert scales, while some gathered qualitative feedback through open-ended questions and semi-structured interviews.
User satisfaction and acceptability were the most commonly reported outcomes. Across studies, participants described the interventions as convenient and accessible, with acceptability generally rated moderate-to-high and reported satisfaction high.
Half of the studies examined usability, using qualitative feedback, the System Usability Scale or Likert scales. Interface design, interaction mode and deployment platform were all observed, alongside differences in usability between studies. A clear preference emerged for free-flowing chat interfaces and customisable features over predefined options. At the same time, some interventions had an unclear scope or limited functionality, leaving people uncertain about what the chatbot could actually do for them.
Usability, engagement and the drop-off problem
Only some studies reported objective utilisation and engagement metrics, such as session frequency, interaction duration, retention over time and task completion. Where these were captured, attrition patterns frequently emerged over time in repeated-measures designs. Uptake in multi-week interventions was often high at the outset before declining – a pattern familiar across digital health more broadly, and one that matters a great deal for interventions whose therapeutic logic depends on sustained practice.
Personalisation and perceived benefit
Most chatbots featured some form of personalisation, reflecting their capacity to adapt conversations and interfaces in response to previous interactions and a person’s emotional state. The most common approach was emotion detection paired with adaptive interaction, allowing people to receive tailored responses and empathetic reflections.
Perceived impact was not consistently measured as a standalone metric. It was more usually folded into qualitative feedback or broader UX evaluations. In the intervention that produced the most granular data, the most frequently reported benefit was improved clarity and awareness.
Where empathy stops being enough
The review’s more cautionary finding is that personalisation and empathy did not consistently translate into stronger clinical outcomes or sustained use. Feeling supported and being helped are not the same thing, and the studies reviewed do not yet demonstrate a reliable link between the two.
Some people reported responses that were repetitive, generic or contextually misaligned. Others raised concerns about over-reliance on chatbots, reduced human contact, data privacy and whether these systems can respond appropriately when someone is in crisis. Inaccurate or clinically misaligned outputs were also linked to an erosion of trust and to disengagement in several studies.
For healthcare professionals, the practical question is less whether these tools have promise than how to appraise them – knowing what a given system is grounded in, where its limits sit and when a conversation needs to move to a human. That judgement is increasingly treated as a core clinical competency, and it sits at the centre of CPD training on the everyday, ethical use of AI in practice.
Design features linked to a better experience
The authors identified several design features associated with better UX outcomes, while being careful to note that these were associations rather than demonstrated causes. They included:
- Deployment on platforms people already knew and used
- Richer interaction modalities beyond plain text
- Integration into existing care pathways
- Personalisation
- Grounding in domain knowledge
- Structured delivery
- Proactive outreach
- Co-design with both experts and end users
The predominance of early-stage studies, combined with limited direct comparative analyses, prevented firm conclusions about which of these features genuinely improved user experience.
What needs to happen next
Taken together, the review suggests that GenAI chatbots have meaningful potential to deliver tailored, empathetic mental health support, and that their acceptability among the people who use them is promising. That is a real finding, and not a small one given the scale of unmet need.
Significant challenges remain, however. Standardising how UX is assessed, grounding intervention design in the needs and preferences of the people who will use these tools, and sustaining engagement beyond the first few weeks all stand out as unresolved. Addressing them, the authors argue, will require co-design with experts and users, validated UX metrics applied in long-term studies, transparent reporting standards, independent evaluation, clearer reporting of model design and training data, and stronger attention to safety, equity and the limits of crisis response.
CCH insight
Generative AI tools are arriving in patient-facing care faster than the evidence base supporting them, which puts the burden of appraisal on clinicians. CCH’s CPD-accredited short course AI Essentials for Primary Care: Tools, Ethics and Everyday Applications covers the practical and ethical judgement this requires – what these tools can and cannot do, where the risks sit, and how to use them safely in day-to-day practice.
Find out more about AI Essentials for Primary Care →

Machine Learning Tool Helps Paediatricians Identify Children at Risk of Persistent Asthma
Key Takeaways:
- A machine learning tool that reads data already held in a child’s electronic health record helped paediatricians more accurately judge which young children are at risk of persistent asthma.
- In a pilot randomised trial using standardised clinical cases, clinicians using the tool reached an average accuracy of 83%, compared with 61% for standard assessment alone.
- The tool is designed to support clinical judgement rather than replace it, and requires no additional tests or questionnaires.
Support for a difficult clinical judgement
A machine learning tool that analyses information already captured in a child’s electronic health record (EHR) has helped paediatricians assess asthma risk more accurately in standardised clinical case scenarios, according to a pilot randomised clinical trial led by a researcher at the Regenstrief Institute. The study was published in the journal Scientific Reports.
The trial evaluated a machine learning-enabled clinical decision support tool known as the Passive Digital Marker. The tool draws on routinely collected EHR data to classify young children as having either a high or a low risk of going on to develop persistent asthma.
Why early asthma risk is hard to predict
Asthma is one of the most common long-term conditions of childhood, yet predicting which young children who have wheezing or other respiratory symptoms will later develop persistent asthma remains difficult. Some children outgrow their early symptoms, while others need ongoing treatment. That uncertainty makes early risk assessment an important, but genuinely challenging, part of paediatric care.
“This tool doesn’t replace a pediatrician’s clinical judgment,” said Arthur H. Owora, PhD, Regenstrief Institute research scientist and lead author of the study. “It helps bring together years of clinical information that’s already in the electronic health record, giving clinicians another source of information when making decisions about a child’s asthma risk.”
How the Passive Digital Marker works
Unlike many prediction tools, the Passive Digital Marker requires no extra testing and asks families to complete no additional questionnaires. Instead, it analyses information that has already been documented in the child’s EHR, including respiratory symptoms, allergies, medication history, respiratory infections and family history. It then presents clinicians with a straightforward high-risk or low-risk assessment.
This approach is intended to save clinicians’ time and reduce the burden on families, since it relies on data that has been gathered over the course of a child’s routine care rather than requiring anything new at the point of decision.
What the trial found
Paediatricians using the tool correctly predicted future asthma more often than those relying on standard assessment alone, achieving an average accuracy of 83% compared with 61%. The improvement was largely driven by better identification of children who went on to develop persistent asthma – the group that is most important to recognise early and hardest to spot.
The researchers stress that the tool is meant to support clinical decision-making, not to supplant it. Its value lies in helping clinicians quickly synthesise years of patient information into a single, easy-to-interpret risk assessment that sits alongside their own expertise. That distinction – between having an AI tool to hand and knowing how to weigh what it tells you – is becoming central to how clinicians are expected to work with these systems.
Limitations and next steps
Because the study used standardised patient cases rather than real-world clinical encounters, further research is needed to establish whether the tool improves outcomes for children in everyday paediatric practice. The pilot demonstrates promise in a controlled setting, but real-world validation is the necessary next stage before wider adoption.
CCH insight
Tools like the Passive Digital Marker are only ever as good as a clinician’s ability to judge when to lean on them and when to look again. That skill – evaluating an AI tool, recognising where it can mislead, and putting sensible governance around its use – is exactly what our short course AI Essentials for GPs: Tools, Ethics and Everyday Applications is designed to build. It’s a 3.5-hour, fully online CPD course led by Prof. Mike Bewick and Dr Dipesh Naik. [Explore the course →]
Funding and authorship
The study was supported in part by the National Institutes of Health under grant K01HL166436. In addition to Owora, it was co-authored by Bowen Jiang, M.S., and Yash Shah, M.S., of the Division of Pediatric Pulmonology, Allergy/Immunology and Sleep Medicine, Department of Pediatrics, Riley Hospital for Children, Indiana University School of Medicine.
Source: Regenstrief Institute
Read More
Physicians May Struggle to Spot AI Errors, Even When Evidence Contradicts the Advice
Key Takeaways:
- In experiments involving decisions about hypothetical patients, physicians tended to trust incorrect advice labelled as artificial intelligence (AI) generated, even when they had the chance to notice that patient recovery data contradicted it.
- Across both experiments, physicians rated the AI system as reliable and did not draw on the recovery data to conclude that its recommendations were wrong; in the second experiment, they failed to notice that the treatment was entirely ineffective.
- The findings, published in the open-access journal PLOS Digital Health by Aranzazu Vinas of the University of the Basque Country and colleagues, point to real challenges for the widely held assumption that a human will reliably catch and correct an algorithm’s mistakes.
What the research examined
New research suggests that physicians may find it difficult to learn from experience when that experience runs counter to advice presented as coming from an AI system. In a series of experiments in which physicians made decisions about treating hypothetical patients, they tended to trust incorrect AI-labelled recommendations, even after being given the opportunity to notice that patient recovery data contradicted those recommendations.
The work was carried out by Aranzazu Vinas of the University of the Basque Country, Spain, together with colleagues, and is published in the open-access journal PLOS Digital Health.
Why AI classification matters in care
AI systems can help physicians categorise patients according to their differing care needs, for example by estimating whether a particular patient is more or less likely to benefit from a given treatment. Because these systems are not perfect, they are intended to be used as suggestions rather than instructions, with any potential errors caught and corrected by the physician using them.
This safeguard rests on an assumption that is easy to take for granted: that a human in the loop will notice when the algorithm is wrong. Prior research, however, has shown that people in general struggle to spot and correct mistakes made by AI. Vinas and colleagues set out to explore how far that difficulty extends to physicians in particular.
How the experiments worked
The researchers analysed data from 223 physicians who took part anonymously in online experiments. Participants were asked to imagine that they had the option to treat patients for a rare disease using a treatment that was not yet proven and still under development. They were told that an AI system had identified which patients were more, or less, likely to benefit from that treatment.
The physicians then chose which patients to treat. After being shown data on how those patients recovered, they rated their perceptions of how reliable the AI system was.
The design contained a deliberate mismatch. The actual effectiveness of the hypothetical treatment did not align with the AI’s recommendations. In the first experiment, the treatment was equally, and moderately, effective for all patients. In the second experiment, it was equally ineffective for everyone. In each case, the recovery data available to physicians should, in principle, have allowed them to see that the AI’s classification did not hold up.
What the physicians did
In both experiments, the physicians tended to rate the AI system as reliable, and they did not appear to use the patient recovery data to conclude that the AI’s recommendations were incorrect. In the second experiment, they did not realise that the treatment was entirely ineffective.
As lead author Aranzazu Vinas notes: “In both experiments, physicians mostly trusted the AI’s classifications and had trouble learning from the feedback. Furthermore, in the second experiment, professionals did not notice that the treatment was completely ineffective.”
Co-author Helena Matute adds: “People tend to say that there is always a human controlling the algorithm, but our experiments show that doctors (as well as anyone else) have problems in learning from the available evidence when it contradicts the suggestions of an algorithm.”
What it means for healthcare
Taken together, the results highlight potential challenges for incorporating AI-based classification into healthcare. If the human overseeing an algorithm cannot readily detect its errors, even when contradicting evidence is in front of them, then the reassurance that a clinician will always catch a mistake may be weaker than commonly assumed.
The authors suggest that future research could build on this study, for instance by developing and testing strategies and protocols designed to strengthen human critical thinking and the detection of AI errors. The aim would be to maximise the benefits of human-AI collaboration while minimising the potential for error.
Co-author Fernando Blanco summarises the wider purpose of this line of enquiry: “It is important to investigate the errors that humans (including doctors) make when working with algorithms, in order to learn how to minimize the problems that arise from them.”
Building the habit of questioning AI
While researchers work on formal protocols, individual clinicians can already sharpen how they interrogate AI output. Knowing when to trust a recommendation, and when to challenge it, is a clinical skill rather than a technical one, and it is one that structured training can help build. Our short course AI Essentials for GPs: Tools, Ethics and Everyday Applications introduces practical frameworks, including the SAFER Evaluation Framework, for spotting errors, fabrications, and outdated recommendations before they reach a patient. For clinicians who want a reliable method for the kind of critical checking this study suggests is all too easy to skip, it is a useful place to start.
Read More
Autonomous AI Agent Matches and Exceeds Physicians Across Simulated Electronic Health Record Cases
Key Takeaways:
- MIRA is an autonomous AI agent that diagnoses and plans treatment inside a simulated electronic health record, rather than acting as a narrow chat tool.
- It reached 88.9% diagnostic accuracy across 574 cases, outperforming board-certified physicians (78.1%) and a mixed-seniority team (71.1%).
- Safety results were strong but preliminary, and the authors stress that MIRA is not a replacement for human clinicians.
A new kind of medical AI agent
A recent study published in the journal Nature introduced MIRA, an autonomous AI agent designed to operate within sandboxed EHR environments. Rather than acting as a single-purpose assistant, MIRA uses a suite of digital tools to simulate the full arc of a clinical workflow. It can order tests, synthesise the results, and produce diagnoses and treatment plans, all while communicating through a chat interface with a patient AI agent that is grounded in the documented history of present illness extracted from retrospective notes from genuine cases.
The system runs on a Fast Healthcare Interoperability Resources (FHIR) based architecture, which executes the agent’s tool calls and records its medical outputs. The researchers note that the example data presented in the paper were shortened and slightly modified to comply with the privacy restrictions attached to the dataset.
Unlike earlier implementations, which were predominantly task-specific chat applications, MIRA was built to independently take in patient histories, order the relevant diagnostic tests, and then use those datasets to reach diagnoses and treatment plans within a controlled simulation. Across the 574 MIMIC-IV cases, MIRA achieved 88.9% diagnostic accuracy, and in a matched 311-case physician comparison it reached 87.8% accuracy, significantly outperforming experienced human physicians under identical simulated conditions while demonstrating strong, though not perfect, safety and guideline performance.
Background: from passing exams to working a ward
Large language models (LLMs) have already proven highly capable at passing standardised medical examinations and answering complex clinical questions. Reviews of the field show, however, that translating this raw clinical knowledge into the operational workflow of a hospital has remained a major challenge.
This gap is attributed to the architectural design of traditional medical AI tools, which behave as narrow, task-specific search or text-generation utilities rather than as active partners in care. By contrast, true clinical decision-making is characterised as an intricate, multi-step process in which doctors repeatedly interview the people in their care, order blood tests or imaging, synthesise conflicting results, and update their hypotheses before arriving at a final treatment plan.
Nearly all of this clinical work takes place within EHR systems that rely on complex, standardised coding protocols. Until now, it remained unproven whether an automated system could reliably handle this end-to-end clinical action space in a realistic, EHR-style environment without committing unacceptable errors.
About the study
The study set out to address this functional gap by developing MIRA, a novel AI tool designed to autonomously ingest and access medical records, identify knowledge gaps, and order diagnostic tests to supplement the EHR record, before using the completed dataset to recommend clinical interventions.
The researchers then tested MIRA’s capabilities in a sandboxed, virtual EHR environment compliant with standard healthcare protocols, including HL7 FHIR. The sandboxed test was conducted on a curated benchmarking dataset of 574 real-world emergency department cases from the Medical Information Mart for Intensive Care (MIMIC-IV) database.
The cases included spanned eight distinct diagnoses across surgery (appendicitis), internal medicine (pneumonia), and oncology (pancreatic cancer), which MIRA navigated using 11 specialised digital tools offering more than 85,000 operational choices. The agent was permitted to request physical examinations, order targeted laboratory values, look up medical histories, and generate medication orders within the simulated EHR, rather than in live patient care.
How MIRA was compared with clinicians
MIRA’s output was compared against two distinct groups of human physicians managing exactly the same cases under identical conditions. The first group was a cohort of four board-certified physicians. The second was a mixed-seniority team consisting of four residents and two board-certified doctors.
A separate, conventional text-based AI agent was used to simulate the people under MIRA’s care, and under the care of the human physician teams. This agent was instructed to respond to questions posed by MIRA or its human counterparts solely on the basis of authentic clinical histories, while resisting adversarial attempts to trick it into prematurely leaking information. The authors noted, however, that simulated patient speech may be more structured than real emergency department conversations.
Study findings
The results revealed that MIRA performed at or above the level of experienced human doctors. It achieved 88.9% diagnostic accuracy across the full 574-case dataset and 87.8% accuracy in the matched 311-case physician comparison. By comparison, the board-certified physicians reached an average accuracy of 78.1% (p < 0.001), while the mixed-seniority medical cohort averaged 71.1% (p < 0.001).
MIRA was found to excel at identifying appendicitis and pancreatitis, achieving a perfect 100% recall for laparoscopic appendectomies. For pancreatic cancer, its diagnostic performance was equivalent to that of the board-certified physicians, while pneumonia and urinary tract infections remained more challenging.
Accuracy without simply “ordering everything”
Notably, MIRA did not achieve its superior accuracy by simply “ordering everything”. While it was observed to request a broader, more comprehensive set of individual blood parameters than the human doctors, its overall test selection remained well below the historical baselines recorded in the dataset.
The findings further demonstrated that the model successfully avoided the systematic over-ordering of high-cost radiological imaging, matching or exceeding physicians on overall resource-alignment metrics.
Safety performance
The safety evaluations were similarly encouraging, though still preliminary. An independent, blinded medical review of 56 patient-level outputs, together with a separate assessment of 468 prescriptions written by MIRA, established that the agent caused zero high-severity drug–drug interactions, zero renal dosing incompatibilities, and zero medication-allergy mismatches. Route specification was the weakest prescription field, at 97% correctness.
When making critical hospital admission decisions for pneumonia and pulmonary embolism, MIRA achieved a perfect recall score of 1.00, indicating that it never missed a single person who required inpatient care. The pulmonary embolism analysis did, however, suggest a tendency towards over-admission, reflecting a cautious disposition strategy.
Conclusions
The study introduces an integrated EHR AI agent, MIRA, that successfully translates clinical intents into structured, safe, and accurate operations, with the potential to support physicians in their work. The authors are careful to caution, however, that MIRA and similar AI agents are not replacements for expert human staff.
The model did not reach 100% perfection across all treatment choices, such as specific antibiotic selections, which highlights the ongoing need for strict human supervision and patient-level safeguards. Future iterations of the model may improve their performance by incorporating evidence from retrieval-based support, stronger governance, and prospective real-world validation before any clinical deployment.
Read More
AI Model Improves Early Detection of Serious Lung Disease in Newborns
Key Takeaways:
- Researchers at the University of Rochester have developed a time-series AI machine learning model that predicts bronchopulmonary dysplasia (BPD) in premature newborns more accurately than existing prediction tools.
- The model uses detailed electronic health record data rather than the limited datasets used in many current online BPD calculators.
- Researchers hope the technology could eventually support real-time clinical decision-making in neonatal intensive care units and help reduce the severity of lung disease in vulnerable infants.
AI and neonatal care
A research team from the University of Rochester has developed a new artificial intelligence and machine learning model designed to improve the prediction of bronchopulmonary dysplasia (BPD), a serious lung disease that affects premature newborns. Their findings were published in The Journal of Pediatrics in a study titled “Time-Series Machine Learning for Prediction of Bronchopulmonary Dysplasia.”
The project centres on the use of time-series machine learning, an approach that analyses patterns in data collected over time, allowing researchers to build more dynamic and potentially more accurate disease prediction models.
Understanding bronchopulmonary dysplasia
Bronchopulmonary dysplasia is a chronic lung condition that primarily affects babies born prematurely and with low birth weight. Because their lungs are still underdeveloped, many premature infants require oxygen therapy and mechanical ventilation shortly after birth to survive. However, this early exposure to oxygen and ventilatory support can contribute to lung injury and long-term respiratory complications.
Children who develop BPD may experience ongoing breathing difficulties and other long-term health challenges linked to impaired lung development.
“We take great effort in the neonatal intensive care unit to prevent lung damage,” said Associate Professor Andrew Dylag, MD, from the Department of Pediatrics, Neonatology. “Despite this, premature infants still develop BPD. There are BPD ‘calculators’ on the internet that can predict the severity of lung disease while the baby is still in the hospital, but they use a very limited set of data.”
According to the research team, recently updated versions of these existing calculators demonstrated lower accuracy than earlier models, prompting the group to explore a different strategy.
“We thought that using more detailed data from the University’s electronic health record would improve disease predictions and allow us to pinpoint vulnerable times when we might be able to intervene to prevent lung disease in newborns,” Dylag said.
Funding supports new collaboration
To support the project, the researchers secured a 2023 Digital Health Seedling award from the Clinical and Translational Science Institute (CTSI).
“We hoped that CTSI could help us test the hypothesis that machine learning could improve disease predictions in hospitalized premature newborns,” Dylag said. “The 2023 Digital Health Seedling award was exactly the type of funding we needed to develop new collaborations across the University community and kickstart our team’s academic interactions.”
The funding enabled the formation of a multidisciplinary research team combining expertise from neonatology, engineering, computer science, biostatistics and health informatics.
The collaboration included Jiebo Luo, PhD, Albert Arendt Hopeman Professor of Engineering in the Department of Computer Science, who connected graduate students to the project, as well as Professor Xing Qiu, PhD, from the Department of Biostatistics and Computational Biology.
“The neonatology team brought content and clinical expertise to the work, the computer science team developed and tested the models, and the biostatisticians ensured the rigor and testing of the models and algorithms,” Dylag said.
The project also expanded on an existing partnership with the University of Rochester Clinical and Translational Science Institute’s Informatics and Analytics group, particularly in relation to electronic health record research.
Building a secure AI research environment
Because the project relied on a very large dataset containing sensitive patient information, the research required substantial data security and privacy protections.
“We initially got involved by helping the study team pull clinical data from eRecord,” said Jack Chang, PhD, associate director of Research Informatics. “Recognizing the project involved a very large patient population and massive data including Protected Health Information, we identified a need for a more secure analytical workspace.”
To address these concerns, the Informatics team transitioned the data into the Secure Environment for Research Data Analytics (SERDA), a protected platform designed for high-risk clinical research and advanced analytics.
High-risk healthcare data projects typically involve extensive administrative oversight and cybersecurity requirements to ensure patient privacy and regulatory compliance.
“SERDA removes that obstacle by providing a secure, scalable, and ‘ready-to-go’ environment tailored for advanced analytics and machine learning,” Chang said. “It allows researchers to focus on their science while knowing their data is protected and compliant with all privacy regulations.”
Chang said the Informatics team worked closely with institutional partners to build the infrastructure required for the study.
“Our team—working with our ISD partners—handled the technical heavy lifting: setting up virtual machines, configuring project shares, ensuring secure access with the security team, and customizing the environment with specialized analytical software,” Chang said. “We also provided training and facilitated the numerous exports of analytical outcomes from SERDA for their publication.”
Towards real-time clinical decision support
Researchers believe the improved AI model may help clinicians identify which infants are most likely to benefit from early intervention and more precise treatment strategies.
The long-term goal is to integrate the model into clinical decision support systems capable of updating disease risk predictions in real time as patient conditions evolve.
“We want to build clinical decision support tools to identify how disease predictions change in real time,” Dylag said. “If we validate our algorithm and can present the disease prediction to the clinical team, we can test guideline implementation for how to manage or treat infants that may reduce BPD severity.”
The research team said that continued collaboration between departments, alongside infrastructure support from CTSI and ISD, will be essential as the project progresses into its next phase of development and validation.
Source: University of Rochester Medicine
Read More
AI Outperformed Emergency Doctors in Harvard Triage Study, Raising Questions About the Future of Clinical Decision-Making
Key Takeaways:
- A Harvard-led study found that an AI reasoning model outperformed emergency physicians in diagnosing patients during hospital triage scenarios using text-based clinical information.
- Researchers said the findings represent a major advance in AI clinical reasoning, although they stressed that AI is not ready to replace human doctors.
- Experts warned that important concerns remain around accountability, bias, safety, and the risk of clinicians becoming overly reliant on AI systems.
AI shows strong performance in emergency medicine trial
From fictional emergency department heroes such as George Clooney in ER to Noah Wyle in The Pitt, emergency physicians have long been portrayed as the ultimate decision-makers in moments of medical crisis. However, a new Harvard study suggests artificial intelligence may increasingly play a major role in those same high-pressure situations.
Researchers from Harvard Medical School and Beth Israel Deaconess Medical Center found that advanced AI systems outperformed human doctors in emergency medicine triage scenarios, making more accurate diagnoses when presented with limited patient information during the critical early stages of hospital admission.
The findings, published in Science, were described by independent experts as representing “a genuine step forward” in AI clinical reasoning.
According to the study authors, large language models (LLMs) “have eclipsed most benchmarks of clinical reasoning”.
AI versus doctors in emergency room triage
One of the study’s central experiments examined 76 patients who presented to the emergency department of a Boston hospital.
Both the AI system and pairs of human physicians were given identical electronic health record information to assess. This included standard triage details such as:
- Vital signs
- Demographic information
- Brief nursing notes explaining why the patient attended hospital
Using only this text-based information, OpenAI’s o1 reasoning model identified the exact diagnosis or a very close diagnosis in 67% of cases.
By comparison, the human physicians achieved diagnostic accuracy rates of between 50% and 55%.
Researchers found the AI’s advantage was especially apparent in triage situations requiring rapid decision-making with minimal available information.
When additional clinical detail was provided, the AI’s diagnostic accuracy increased further to 82%. Human experts achieved accuracy rates between 70% and 79% under those circumstances, although researchers noted the difference was not statistically significant in that setting.
AI also performed better in treatment planning
The study also evaluated how AI performed in longer-term clinical planning tasks.
In this experiment, the AI system and a group of 46 doctors were asked to review five detailed clinical case studies and develop treatment strategies. These included decisions relating to:
- Antibiotic regimens
- Ongoing management plans
- End-of-life care processes
The AI significantly outperformed the doctors.
Researchers reported that the AI achieved a score of 89%, compared with 34% among physicians using conventional resources such as search engines.
AI detected a diagnosis human doctors missed
One example highlighted in the study involved a patient with a pulmonary embolism and worsening symptoms.
Human doctors believed the patient’s anticoagulant treatment was failing. However, the AI system identified something clinicians had overlooked – the patient had a history of lupus, which may have been responsible for inflammation in the lungs.
The AI’s interpretation was ultimately confirmed as correct.
Researchers say AI will reshape medicine – not replace doctors
Despite the strong performance shown by AI systems, researchers stressed that the technology is not ready to replace physicians.
The study only assessed AI systems using text-based patient information. It did not evaluate the AI’s ability to interpret non-verbal clinical signals that doctors routinely use during patient assessment, such as:
- Visible distress
- Facial appearance
- Behaviour
- Physical examination findings
As a result, researchers said the AI functioned more like a clinician reviewing paperwork and offering a second opinion.
“I don’t think our findings mean that AI replaces doctors,” said Arjun Manrai, one of the lead authors of the study who heads an AI lab at Harvard Medical School. “I think it does mean that we’re witnessing a really profound change in technology that will reshape medicine.”
Dr Adam Rodman, another lead author and physician at Beth Israel Deaconess Medical Center, described AI LLMs as among “the most impactful technologies in decades”.
He suggested healthcare may move towards what he described as a “triadic care model”.
“Over the next decade,” Rodman said, AI would not replace physicians but instead work alongside them in a new model involving “the doctor, the patient, and an artificial intelligence system”.
AI use in healthcare is already growing
The findings come amid rapidly increasing AI adoption within healthcare systems.
According to research published last month, nearly one in five physicians in the United States are already using AI to assist with diagnosis.
In the United Kingdom, a recent Royal College of Physicians survey found:
- 16% of doctors use AI daily
- A further 15% use AI weekly
- Clinical decision-making is among the most common applications
However, concerns around safety and accountability remain significant.
UK doctors surveyed identified AI errors and legal liability as among their biggest worries.
“There is not a formal framework right now for accountability,” said Rodman.
He also stressed the continuing importance of human clinicians in patient care.
“Patients ultimately want humans to guide them through life or death decisions [and] to guide them through challenging treatment decisions,” he said.
Experts warn against over-reliance on AI
Independent experts said the study highlights the rapidly improving capabilities of AI systems in medicine, but also demonstrates the need for caution.
Prof Ewen Harrison, co-director of the University of Edinburgh’s Centre for Medical Informatics, said the findings suggest AI systems are beginning to evolve beyond theoretical testing environments.
“These systems are no longer just passing medical exams or solving artificial test cases,” he said. “They are starting to look like useful second-opinion tools for clinicians, particularly when it is important to consider a wider range of possible diagnoses and avoid missing something important.”
However, Dr Wei Xing from the University of Sheffield warned that the study also raised concerns about how clinicians interact with AI recommendations.
He suggested some doctors may unconsciously defer to AI-generated answers instead of independently evaluating clinical information themselves.
“This tendency could grow more significant as AI becomes more routinely used in clinical settings,” he said.
Dr Xing also noted that the study provided limited information about where the AI may perform less effectively, including whether diagnostic accuracy differed among certain patient populations such as:
- Older adults
- Non-English speakers
- People with more complex communication needs
He cautioned against interpreting the findings as evidence that publicly available AI tools are ready for independent medical use.
“It does not demonstrate that AI is safe for routine clinical use, nor that the public should turn to freely available AI tools as a substitute for medical advice,” he said.
Source: The Guardian
Read More
Large Language Models Show Promise in Detecting Drug Safety Signals from Clinical Notes
Key Takeaways:
- Large language models can identify immune-related adverse events in clinical notes without task-specific training, offering a potential alternative to labour-intensive manual review
- Performance remains below the threshold required for clinical decision support, with models tending to overpredict adverse events
- Despite limitations, this approach may support large-scale safety monitoring and accelerate research into cancer immunotherapies
The challenge of detecting drug safety signals
Drug safety signals are often embedded within unstructured clinical text, particularly in electronic health records. Identifying these signals has traditionally required either manual chart abstraction, which is resource-intensive, or natural language processing systems tailored to specific drugs and healthcare settings.
This challenge is particularly evident in the case of immune checkpoint inhibitors. These cancer therapies, first introduced in 2011, are associated with a broad range of immune-related adverse events. These events can affect multiple organ systems, including the colon, liver, lungs, heart, nervous system, skin, and endocrine system, making systematic detection complex and time-consuming.
Exploring large language models as a solution
Large language models are increasingly being explored as a way to streamline the identification of drug safety signals within clinical text. A multicentre study, published in eBioMedicine, evaluated whether these models could detect immune-related adverse events associated with immune checkpoint inhibitors.
The study focused on a zero-shot learning approach. In this setting, the model receives a single, detailed prompt without prior examples. The prompt used by the researchers began: “You are a clinical expert in identifying immune-related adverse events caused by immune checkpoint inhibitors …” and included a list of six immune checkpoint inhibitors alongside numerous associated adverse events.
This prompt was applied to clinical notes from multiple sources. These included records from 100 people treated at Vanderbilt Health, 70 people from the University of California, San Francisco, and 272 people enrolled in seven Roche-sponsored clinical trials.
Study design and model performance
The research team evaluated three models: GPT-3.5, GPT-4, and GPT-4o, with GPT-4o demonstrating the strongest overall performance.
To assess accuracy, the investigators used F1 scores, a metric that balances false positives and false negatives. Scores range from zero to one, with values above 90 percent considered excellent. A score of 80 percent or higher may be sufficient for use in automated clinical decision support systems.
At the patient level, GPT-4o achieved average F1 scores of 56 percent for Vanderbilt Health data, 66 percent for University of California, San Francisco data, and 62 percent for Roche clinical trial data. The models showed a consistent tendency to overpredict the presence of immune-related adverse events.
When analysing individual clinical notes, the model achieved an average F1 score of 57 percent across 667 notes from Vanderbilt Health, evaluating 17 different adverse events.
Implications for clinical practice and research
The findings suggest that large language models can play a role in identifying drug safety signals, even without task-specific training data.
“Manual patient chart abstraction for monitoring the safety and efficacy of drugs already at market requires tremendous resources and puts a drag on the pace of discovery in precision medicine. And that’s especially true with immune checkpoint inhibitors, where the adverse events are so varied. If zero-shot learning with LLMs could help with these notes, it could significantly reduce time and costs for all concerned,” said the report’s corresponding author, Cosmin Bejan, PhD, assistant professor of Biomedical Informatics at Vanderbilt Health.
However, the current level of performance falls short of what would be required for clinical decision support.
“These results show that zero-shot learning with a powerful LLM is useful for detecting these adverse events,” Bejan said. “This performance does not rise to the level required for clinical decision support, but the method could be valuable for automated irAE extraction across multiple sites, potentially speeding discovery and enhancing the safety and effectiveness of cancer immunotherapies.”
Wider research context
The study involved collaboration among multiple researchers at Vanderbilt Health, including Yaomin Xu, PhD, Eric Mukherjee, MD, PhD, Matthew Krantz, MD, Douglas Johnson, MD, MSCI, Elizabeth Phillips, MD, and Justin Balko, PhD. Funding support was provided in part by the National Institutes of Health.
Related research further highlights safety concerns associated with immune checkpoint inhibitors. In a research letter published in JAMA Oncology, Mukherjee, Phillips, and colleagues used logistic regression analysis of adverse event reports from the Food and Drug Administration. They confirmed that these therapies are independently associated with an increased risk of Stevens-Johnson syndrome and toxic epidermal necrolysis, which are severe and potentially life-threatening skin reactions. The study also found that this risk may be linked to exposure to human leukocyte antigen–restricted drugs.
Conclusion
Large language models represent a promising tool for extracting clinically meaningful insights from unstructured health data. While their current performance limits direct clinical application, their ability to operate across multiple datasets without task-specific training suggests potential for supporting large-scale pharmacovigilance efforts. As these models continue to improve, they may contribute to more efficient and comprehensive monitoring of drug safety in clinical practice.
Read More
‘Shadow AI’ on the Rise in Healthcare as Clinicians Turn to Unauthorised Tools to Improve Workflows
Key Takeaways:
- A survey of healthcare professionals found that 57% have encountered or used unauthorised artificial intelligence tools in their organisations, highlighting the growing presence of so-called “shadow AI” in healthcare settings.
- Many clinicians and administrators report using these tools to improve efficiency, analyse data, and manage administrative tasks, particularly when approved solutions or clear guidance are lacking.
- While most respondents believe AI will significantly improve healthcare within five years, concerns about patient safety, data privacy, and security risks remain widespread.
Unauthorised AI tools emerging in healthcare workplaces
A new survey suggests that artificial intelligence tools are already being used in healthcare organisations in ways that fall outside formal governance structures. According to the findings, a significant proportion of healthcare professionals have either encountered or used AI tools that have not been authorised by their employer.
The survey, conducted by Wolters Kluwer Health, gathered responses from 518 healthcare professionals, including both clinical providers and administrators. The research was carried out in December 2025 and was released publicly last week.
Overall, the findings indicate that four in ten healthcare professionals reported encountering unauthorised AI tools within their organisation, while 17% acknowledged personally using such tools.
When responses were analysed by professional role, 15% of physicians admitted to using an unauthorised AI tool, compared with 19% of administrators. In addition, one in ten respondents reported using an unauthorised AI tool in connection with direct patient care.
The report refers to the unauthorised adoption of artificial intelligence tools in professional environments as “shadow AI.”
Why healthcare staff turn to unauthorised AI
The survey findings suggest that healthcare professionals are often motivated by practical needs rather than deliberate attempts to bypass organisational policies.
According to the report:
“Clinical and administrative teams want to adhere to rules surrounding AI usage, but if the organization hasn’t provided guidance or approved solutions, they’ll experiment with generic tools to improve their workflows.”
Many respondents indicated that the absence of formal guidance or approved AI platforms has encouraged individuals to explore publicly available tools on their own.
The most frequently cited motivation for using unauthorised AI tools was the need to accelerate workflows and improve efficiency. Approximately half of respondents identified faster workflows as the primary reason for using these tools.
However, the survey also revealed differences in how clinical and administrative staff tend to use AI technologies.
Administrators were more likely to employ AI tools for operational or analytical tasks such as:
- Data analysis
- Predictive analytics
- Administrative processes
Healthcare providers, meanwhile, reported using AI for activities such as:
- Data analysis
- Patient scheduling
- Patient engagement tasks
The findings also indicate that clinicians were more likely than administrators to experiment with AI tools out of curiosity.
Governance and policy development remain uneven
The survey results highlight a notable imbalance in how different professional groups participate in the development of AI policies within healthcare organisations.
According to the report, administrators were three times more likely than clinical providers to be actively involved in developing AI governance policies.
Specifically:
- 30% of administrators reported involvement in AI policy development
- Only 9% of providers said they had participated in such efforts
This difference suggests that policy ownership around AI adoption may currently be concentrated within administrative leadership rather than clinical teams.
Administrators also reported greater familiarity with their organisation’s AI policies compared with providers, although awareness varied across both groups.
Security and privacy risks associated with “shadow AI”
The use of unauthorised AI tools raises important concerns about data security, privacy protection, and governance oversight.
The Wolters Kluwer report notes that inconsistent or unsanctioned AI usage can expose organisations to potential vulnerabilities. Without clear oversight, the integration of external AI tools may lead to data privacy violations, security breaches, or inappropriate handling of sensitive information.
To illustrate these risks, the report references a 2025 study by IBM, which found that 97% of organisations that experienced an AI-related security incident lacked adequate AI access controls.
Security incidents involving AI systems can have significant consequences, including financial losses, operational disruption, and damage to public trust.
Healthcare professionals remain optimistic about AI’s future
Despite concerns about governance and security, the survey indicates that most healthcare professionals remain broadly optimistic about the long-term role of artificial intelligence in healthcare.
Nearly 90% of respondents said they believe AI will significantly improve healthcare within the next five years. Administrators were found to be slightly more optimistic than clinical providers about the potential benefits of the technology.
At the same time, respondents recognised that AI implementation carries important risks that must be addressed.
Patient safety was identified by around half of respondents as the most significant risk associated with AI adoption.
Meanwhile, nearly half of respondents also expressed concerns about data privacy risks.
These findings suggest that healthcare professionals recognise both the transformative potential of artificial intelligence and the need for careful governance, clear guidance, and secure systems.
Addressing the rise of “shadow AI”
The report concludes that addressing the growth of shadow AI requires organisations to understand why staff are turning to unauthorised tools rather than focusing solely on restricting access.
According to the report:
“Ultimately, addressing shadow AI is not about restricting access to productivity tools. Leaders must understand why teams are using unsanctioned tools and which challenges they’re trying to solve, and then identify enterprise-level tools that can accomplish these goals safely and securely.”
As artificial intelligence becomes increasingly embedded in healthcare workflows, organisations may need to develop clearer policies, provide approved tools, and involve both clinical and administrative staff in governance decisions.
Such measures may help ensure that the benefits of AI can be realised while protecting patient safety, safeguarding sensitive data, and maintaining organisational trust.
Read More