
People Like AI Mental Health Chatbots. Whether They Help Is Another Question
Key Takeaways:
- Across 21 studies in 11 countries, people using generative AI mental health chatbots reported high satisfaction and found them convenient and accessible.
- Personalisation and empathy did not reliably translate into better clinical outcomes, and engagement often faded over time.
- The evidence base remains early-stage, leaving safety, equity and crisis response unresolved.
A treatment gap that digital tools are being asked to fill
Generative artificial intelligence (GenAI) chatbots designed to support mental health are winning people over on experience, but the research needed to establish whether they are safe and clinically effective has not kept pace. That is the central finding of a review, currently in press in the journal npj Digital Medicine, which examined the user experience (UX) and intervention design of GenAI mental health chatbots.
The context for this work is a substantial and persistent shortfall in care. Around 25% of people worldwide experience a mental health problem, yet approximately 85% do not receive adequate treatment. The reasons are varied and overlapping: stigma, cost, shortages of trained professionals, geographic distance from services and structural inequities, among others. As the prevalence of mental health conditions has grown while treatment gaps have remained, attention has turned towards innovative models of delivery – and digital tools, with their scalability and convenience, have become a focus of that search.
Digital mental health interventions deliver treatment or support through a range of channels, including chatbots, websites, mobile applications and wearables. Conversational agents, more commonly known as chatbots, are applications that simulate human dialogue using machine learning and natural language processing algorithms.
From scripted responses to open-ended conversation
Traditional mental health chatbots deliver pre-scripted therapeutic content through rules-based or retrieval-based systems. Their strength is predictability, but that same design limits their capacity to personalise support or to recognise what an individual actually needs in the moment.
Chatbots built on large language models (LLMs) work differently. They can simulate core aspects of a therapeutic encounter, including personalised suggestions and empathetic reflections. That flexibility comes with a trade-off. GenAI systems may produce responses that are incorrect or inappropriate, and their open-ended conversational capacity makes intervention design both more consequential and more complex than it is for rules-based systems. When a system can say almost anything, design decisions carry considerably more weight.
How the review was carried out
The researchers set out to map the design characteristics and UX outcomes of interventions involving GenAI mental health chatbots. They began with a systematic literature search to identify studies covering the design and deployment of such tools. Reviews, editorials, media articles and commentaries were excluded.
In total, 21 studies were selected, conducted across 11 countries between 2023 and 2025. The largest numbers came from China and the United Kingdom, followed by the United States. The included work spanned a wide range of maturity, from early-stage prototype evaluations through to clinical trials, with one real-world implementation study.
Most studies recruited general or clinical adult populations, including older people living with dementia. Others involved simulated users or university students. Sample sizes ranged from as few as five participants to as many as 527. Across the body of evidence, there was substantial heterogeneity in outcome measures, and most interventions remained at an early stage of development – two features that shape how much can reasonably be concluded from the literature as it stands.
What the interventions were designed to do
Chatbot interventions most often targeted depression and anxiety, and tended to adopt shared therapeutic mechanisms, including mindfulness, emotion regulation and cognitive restructuring. Some systems were oriented towards mental well-being, stress and loneliness, emphasising general support and preventive care rather than treatment for a specific condition. Others addressed eating disorders, post-traumatic stress disorder and dementia.
The dementia-focused interventions are worth distinguishing. Rather than attempting to address the central neurological features of the condition, they targeted its related psychological dimensions – carer burnout, psychological distress and loneliness among them.
Most interventions were grounded in cognitive-behavioural therapy principles. The specific techniques drawn upon included behavioural activation, psychoeducation, Socratic dialogue, acceptance and commitment therapy, cognitive restructuring and mindfulness.
How the tools were delivered
Interventions varied in frequency, delivery modality and duration. The majority were short-term, running from two to eight weeks. Most were deployed through web-based interfaces and mobile applications, while some were delivered via messaging or social media platforms – meeting people on services they already used rather than asking them to adopt something new.
About 67% of interventions were non-embodied, text-based chatbots. The remainder used voice, avatar-based, augmented reality or other multimodal forms of interaction, with the intention of improving engagement and realism.
What people made of them
All but two of the studies evaluated at least one UX domain. The majority relied on quantitative measures, typically Likert scales, while some gathered qualitative feedback through open-ended questions and semi-structured interviews.
User satisfaction and acceptability were the most commonly reported outcomes. Across studies, participants described the interventions as convenient and accessible, with acceptability generally rated moderate-to-high and reported satisfaction high.
Half of the studies examined usability, using qualitative feedback, the System Usability Scale or Likert scales. Interface design, interaction mode and deployment platform were all observed, alongside differences in usability between studies. A clear preference emerged for free-flowing chat interfaces and customisable features over predefined options. At the same time, some interventions had an unclear scope or limited functionality, leaving people uncertain about what the chatbot could actually do for them.
Usability, engagement and the drop-off problem
Only some studies reported objective utilisation and engagement metrics, such as session frequency, interaction duration, retention over time and task completion. Where these were captured, attrition patterns frequently emerged over time in repeated-measures designs. Uptake in multi-week interventions was often high at the outset before declining – a pattern familiar across digital health more broadly, and one that matters a great deal for interventions whose therapeutic logic depends on sustained practice.
Personalisation and perceived benefit
Most chatbots featured some form of personalisation, reflecting their capacity to adapt conversations and interfaces in response to previous interactions and a person’s emotional state. The most common approach was emotion detection paired with adaptive interaction, allowing people to receive tailored responses and empathetic reflections.
Perceived impact was not consistently measured as a standalone metric. It was more usually folded into qualitative feedback or broader UX evaluations. In the intervention that produced the most granular data, the most frequently reported benefit was improved clarity and awareness.
Where empathy stops being enough
The review’s more cautionary finding is that personalisation and empathy did not consistently translate into stronger clinical outcomes or sustained use. Feeling supported and being helped are not the same thing, and the studies reviewed do not yet demonstrate a reliable link between the two.
Some people reported responses that were repetitive, generic or contextually misaligned. Others raised concerns about over-reliance on chatbots, reduced human contact, data privacy and whether these systems can respond appropriately when someone is in crisis. Inaccurate or clinically misaligned outputs were also linked to an erosion of trust and to disengagement in several studies.
For healthcare professionals, the practical question is less whether these tools have promise than how to appraise them – knowing what a given system is grounded in, where its limits sit and when a conversation needs to move to a human. That judgement is increasingly treated as a core clinical competency, and it sits at the centre of CPD training on the everyday, ethical use of AI in practice.
Design features linked to a better experience
The authors identified several design features associated with better UX outcomes, while being careful to note that these were associations rather than demonstrated causes. They included:
- Deployment on platforms people already knew and used
- Richer interaction modalities beyond plain text
- Integration into existing care pathways
- Personalisation
- Grounding in domain knowledge
- Structured delivery
- Proactive outreach
- Co-design with both experts and end users
The predominance of early-stage studies, combined with limited direct comparative analyses, prevented firm conclusions about which of these features genuinely improved user experience.
What needs to happen next
Taken together, the review suggests that GenAI chatbots have meaningful potential to deliver tailored, empathetic mental health support, and that their acceptability among the people who use them is promising. That is a real finding, and not a small one given the scale of unmet need.
Significant challenges remain, however. Standardising how UX is assessed, grounding intervention design in the needs and preferences of the people who will use these tools, and sustaining engagement beyond the first few weeks all stand out as unresolved. Addressing them, the authors argue, will require co-design with experts and users, validated UX metrics applied in long-term studies, transparent reporting standards, independent evaluation, clearer reporting of model design and training data, and stronger attention to safety, equity and the limits of crisis response.
CCH insight
Generative AI tools are arriving in patient-facing care faster than the evidence base supporting them, which puts the burden of appraisal on clinicians. CCH’s CPD-accredited short course AI Essentials for Primary Care: Tools, Ethics and Everyday Applications covers the practical and ethical judgement this requires – what these tools can and cannot do, where the risks sit, and how to use them safely in day-to-day practice.
Find out more about AI Essentials for Primary Care →

Autonomous AI Agent Matches and Exceeds Physicians Across Simulated Electronic Health Record Cases
Key Takeaways:
- MIRA is an autonomous AI agent that diagnoses and plans treatment inside a simulated electronic health record, rather than acting as a narrow chat tool.
- It reached 88.9% diagnostic accuracy across 574 cases, outperforming board-certified physicians (78.1%) and a mixed-seniority team (71.1%).
- Safety results were strong but preliminary, and the authors stress that MIRA is not a replacement for human clinicians.
A new kind of medical AI agent
A recent study published in the journal Nature introduced MIRA, an autonomous AI agent designed to operate within sandboxed EHR environments. Rather than acting as a single-purpose assistant, MIRA uses a suite of digital tools to simulate the full arc of a clinical workflow. It can order tests, synthesise the results, and produce diagnoses and treatment plans, all while communicating through a chat interface with a patient AI agent that is grounded in the documented history of present illness extracted from retrospective notes from genuine cases.
The system runs on a Fast Healthcare Interoperability Resources (FHIR) based architecture, which executes the agent’s tool calls and records its medical outputs. The researchers note that the example data presented in the paper were shortened and slightly modified to comply with the privacy restrictions attached to the dataset.
Unlike earlier implementations, which were predominantly task-specific chat applications, MIRA was built to independently take in patient histories, order the relevant diagnostic tests, and then use those datasets to reach diagnoses and treatment plans within a controlled simulation. Across the 574 MIMIC-IV cases, MIRA achieved 88.9% diagnostic accuracy, and in a matched 311-case physician comparison it reached 87.8% accuracy, significantly outperforming experienced human physicians under identical simulated conditions while demonstrating strong, though not perfect, safety and guideline performance.
Background: from passing exams to working a ward
Large language models (LLMs) have already proven highly capable at passing standardised medical examinations and answering complex clinical questions. Reviews of the field show, however, that translating this raw clinical knowledge into the operational workflow of a hospital has remained a major challenge.
This gap is attributed to the architectural design of traditional medical AI tools, which behave as narrow, task-specific search or text-generation utilities rather than as active partners in care. By contrast, true clinical decision-making is characterised as an intricate, multi-step process in which doctors repeatedly interview the people in their care, order blood tests or imaging, synthesise conflicting results, and update their hypotheses before arriving at a final treatment plan.
Nearly all of this clinical work takes place within EHR systems that rely on complex, standardised coding protocols. Until now, it remained unproven whether an automated system could reliably handle this end-to-end clinical action space in a realistic, EHR-style environment without committing unacceptable errors.
About the study
The study set out to address this functional gap by developing MIRA, a novel AI tool designed to autonomously ingest and access medical records, identify knowledge gaps, and order diagnostic tests to supplement the EHR record, before using the completed dataset to recommend clinical interventions.
The researchers then tested MIRA’s capabilities in a sandboxed, virtual EHR environment compliant with standard healthcare protocols, including HL7 FHIR. The sandboxed test was conducted on a curated benchmarking dataset of 574 real-world emergency department cases from the Medical Information Mart for Intensive Care (MIMIC-IV) database.
The cases included spanned eight distinct diagnoses across surgery (appendicitis), internal medicine (pneumonia), and oncology (pancreatic cancer), which MIRA navigated using 11 specialised digital tools offering more than 85,000 operational choices. The agent was permitted to request physical examinations, order targeted laboratory values, look up medical histories, and generate medication orders within the simulated EHR, rather than in live patient care.
How MIRA was compared with clinicians
MIRA’s output was compared against two distinct groups of human physicians managing exactly the same cases under identical conditions. The first group was a cohort of four board-certified physicians. The second was a mixed-seniority team consisting of four residents and two board-certified doctors.
A separate, conventional text-based AI agent was used to simulate the people under MIRA’s care, and under the care of the human physician teams. This agent was instructed to respond to questions posed by MIRA or its human counterparts solely on the basis of authentic clinical histories, while resisting adversarial attempts to trick it into prematurely leaking information. The authors noted, however, that simulated patient speech may be more structured than real emergency department conversations.
Study findings
The results revealed that MIRA performed at or above the level of experienced human doctors. It achieved 88.9% diagnostic accuracy across the full 574-case dataset and 87.8% accuracy in the matched 311-case physician comparison. By comparison, the board-certified physicians reached an average accuracy of 78.1% (p < 0.001), while the mixed-seniority medical cohort averaged 71.1% (p < 0.001).
MIRA was found to excel at identifying appendicitis and pancreatitis, achieving a perfect 100% recall for laparoscopic appendectomies. For pancreatic cancer, its diagnostic performance was equivalent to that of the board-certified physicians, while pneumonia and urinary tract infections remained more challenging.
Accuracy without simply “ordering everything”
Notably, MIRA did not achieve its superior accuracy by simply “ordering everything”. While it was observed to request a broader, more comprehensive set of individual blood parameters than the human doctors, its overall test selection remained well below the historical baselines recorded in the dataset.
The findings further demonstrated that the model successfully avoided the systematic over-ordering of high-cost radiological imaging, matching or exceeding physicians on overall resource-alignment metrics.
Safety performance
The safety evaluations were similarly encouraging, though still preliminary. An independent, blinded medical review of 56 patient-level outputs, together with a separate assessment of 468 prescriptions written by MIRA, established that the agent caused zero high-severity drug–drug interactions, zero renal dosing incompatibilities, and zero medication-allergy mismatches. Route specification was the weakest prescription field, at 97% correctness.
When making critical hospital admission decisions for pneumonia and pulmonary embolism, MIRA achieved a perfect recall score of 1.00, indicating that it never missed a single person who required inpatient care. The pulmonary embolism analysis did, however, suggest a tendency towards over-admission, reflecting a cautious disposition strategy.
Conclusions
The study introduces an integrated EHR AI agent, MIRA, that successfully translates clinical intents into structured, safe, and accurate operations, with the potential to support physicians in their work. The authors are careful to caution, however, that MIRA and similar AI agents are not replacements for expert human staff.
The model did not reach 100% perfection across all treatment choices, such as specific antibiotic selections, which highlights the ongoing need for strict human supervision and patient-level safeguards. Future iterations of the model may improve their performance by incorporating evidence from retrieval-based support, stronger governance, and prospective real-world validation before any clinical deployment.
Read More