
People Like AI Mental Health Chatbots. Whether They Help Is Another Question
Key Takeaways:
- Across 21 studies in 11 countries, people using generative AI mental health chatbots reported high satisfaction and found them convenient and accessible.
- Personalisation and empathy did not reliably translate into better clinical outcomes, and engagement often faded over time.
- The evidence base remains early-stage, leaving safety, equity and crisis response unresolved.
A treatment gap that digital tools are being asked to fill
Generative artificial intelligence (GenAI) chatbots designed to support mental health are winning people over on experience, but the research needed to establish whether they are safe and clinically effective has not kept pace. That is the central finding of a review, currently in press in the journal npj Digital Medicine, which examined the user experience (UX) and intervention design of GenAI mental health chatbots.
The context for this work is a substantial and persistent shortfall in care. Around 25% of people worldwide experience a mental health problem, yet approximately 85% do not receive adequate treatment. The reasons are varied and overlapping: stigma, cost, shortages of trained professionals, geographic distance from services and structural inequities, among others. As the prevalence of mental health conditions has grown while treatment gaps have remained, attention has turned towards innovative models of delivery – and digital tools, with their scalability and convenience, have become a focus of that search.
Digital mental health interventions deliver treatment or support through a range of channels, including chatbots, websites, mobile applications and wearables. Conversational agents, more commonly known as chatbots, are applications that simulate human dialogue using machine learning and natural language processing algorithms.
From scripted responses to open-ended conversation
Traditional mental health chatbots deliver pre-scripted therapeutic content through rules-based or retrieval-based systems. Their strength is predictability, but that same design limits their capacity to personalise support or to recognise what an individual actually needs in the moment.
Chatbots built on large language models (LLMs) work differently. They can simulate core aspects of a therapeutic encounter, including personalised suggestions and empathetic reflections. That flexibility comes with a trade-off. GenAI systems may produce responses that are incorrect or inappropriate, and their open-ended conversational capacity makes intervention design both more consequential and more complex than it is for rules-based systems. When a system can say almost anything, design decisions carry considerably more weight.
How the review was carried out
The researchers set out to map the design characteristics and UX outcomes of interventions involving GenAI mental health chatbots. They began with a systematic literature search to identify studies covering the design and deployment of such tools. Reviews, editorials, media articles and commentaries were excluded.
In total, 21 studies were selected, conducted across 11 countries between 2023 and 2025. The largest numbers came from China and the United Kingdom, followed by the United States. The included work spanned a wide range of maturity, from early-stage prototype evaluations through to clinical trials, with one real-world implementation study.
Most studies recruited general or clinical adult populations, including older people living with dementia. Others involved simulated users or university students. Sample sizes ranged from as few as five participants to as many as 527. Across the body of evidence, there was substantial heterogeneity in outcome measures, and most interventions remained at an early stage of development – two features that shape how much can reasonably be concluded from the literature as it stands.
What the interventions were designed to do
Chatbot interventions most often targeted depression and anxiety, and tended to adopt shared therapeutic mechanisms, including mindfulness, emotion regulation and cognitive restructuring. Some systems were oriented towards mental well-being, stress and loneliness, emphasising general support and preventive care rather than treatment for a specific condition. Others addressed eating disorders, post-traumatic stress disorder and dementia.
The dementia-focused interventions are worth distinguishing. Rather than attempting to address the central neurological features of the condition, they targeted its related psychological dimensions – carer burnout, psychological distress and loneliness among them.
Most interventions were grounded in cognitive-behavioural therapy principles. The specific techniques drawn upon included behavioural activation, psychoeducation, Socratic dialogue, acceptance and commitment therapy, cognitive restructuring and mindfulness.
How the tools were delivered
Interventions varied in frequency, delivery modality and duration. The majority were short-term, running from two to eight weeks. Most were deployed through web-based interfaces and mobile applications, while some were delivered via messaging or social media platforms – meeting people on services they already used rather than asking them to adopt something new.
About 67% of interventions were non-embodied, text-based chatbots. The remainder used voice, avatar-based, augmented reality or other multimodal forms of interaction, with the intention of improving engagement and realism.
What people made of them
All but two of the studies evaluated at least one UX domain. The majority relied on quantitative measures, typically Likert scales, while some gathered qualitative feedback through open-ended questions and semi-structured interviews.
User satisfaction and acceptability were the most commonly reported outcomes. Across studies, participants described the interventions as convenient and accessible, with acceptability generally rated moderate-to-high and reported satisfaction high.
Half of the studies examined usability, using qualitative feedback, the System Usability Scale or Likert scales. Interface design, interaction mode and deployment platform were all observed, alongside differences in usability between studies. A clear preference emerged for free-flowing chat interfaces and customisable features over predefined options. At the same time, some interventions had an unclear scope or limited functionality, leaving people uncertain about what the chatbot could actually do for them.
Usability, engagement and the drop-off problem
Only some studies reported objective utilisation and engagement metrics, such as session frequency, interaction duration, retention over time and task completion. Where these were captured, attrition patterns frequently emerged over time in repeated-measures designs. Uptake in multi-week interventions was often high at the outset before declining – a pattern familiar across digital health more broadly, and one that matters a great deal for interventions whose therapeutic logic depends on sustained practice.
Personalisation and perceived benefit
Most chatbots featured some form of personalisation, reflecting their capacity to adapt conversations and interfaces in response to previous interactions and a person’s emotional state. The most common approach was emotion detection paired with adaptive interaction, allowing people to receive tailored responses and empathetic reflections.
Perceived impact was not consistently measured as a standalone metric. It was more usually folded into qualitative feedback or broader UX evaluations. In the intervention that produced the most granular data, the most frequently reported benefit was improved clarity and awareness.
Where empathy stops being enough
The review’s more cautionary finding is that personalisation and empathy did not consistently translate into stronger clinical outcomes or sustained use. Feeling supported and being helped are not the same thing, and the studies reviewed do not yet demonstrate a reliable link between the two.
Some people reported responses that were repetitive, generic or contextually misaligned. Others raised concerns about over-reliance on chatbots, reduced human contact, data privacy and whether these systems can respond appropriately when someone is in crisis. Inaccurate or clinically misaligned outputs were also linked to an erosion of trust and to disengagement in several studies.
For healthcare professionals, the practical question is less whether these tools have promise than how to appraise them – knowing what a given system is grounded in, where its limits sit and when a conversation needs to move to a human. That judgement is increasingly treated as a core clinical competency, and it sits at the centre of CPD training on the everyday, ethical use of AI in practice.
Design features linked to a better experience
The authors identified several design features associated with better UX outcomes, while being careful to note that these were associations rather than demonstrated causes. They included:
- Deployment on platforms people already knew and used
- Richer interaction modalities beyond plain text
- Integration into existing care pathways
- Personalisation
- Grounding in domain knowledge
- Structured delivery
- Proactive outreach
- Co-design with both experts and end users
The predominance of early-stage studies, combined with limited direct comparative analyses, prevented firm conclusions about which of these features genuinely improved user experience.
What needs to happen next
Taken together, the review suggests that GenAI chatbots have meaningful potential to deliver tailored, empathetic mental health support, and that their acceptability among the people who use them is promising. That is a real finding, and not a small one given the scale of unmet need.
Significant challenges remain, however. Standardising how UX is assessed, grounding intervention design in the needs and preferences of the people who will use these tools, and sustaining engagement beyond the first few weeks all stand out as unresolved. Addressing them, the authors argue, will require co-design with experts and users, validated UX metrics applied in long-term studies, transparent reporting standards, independent evaluation, clearer reporting of model design and training data, and stronger attention to safety, equity and the limits of crisis response.
CCH insight
Generative AI tools are arriving in patient-facing care faster than the evidence base supporting them, which puts the burden of appraisal on clinicians. CCH’s CPD-accredited short course AI Essentials for Primary Care: Tools, Ethics and Everyday Applications covers the practical and ethical judgement this requires – what these tools can and cannot do, where the risks sit, and how to use them safely in day-to-day practice.
Find out more about AI Essentials for Primary Care →

AI-Supported Digital Care Improves Rheumatoid Arthritis Outcomes After Hospital Discharge
Key Takeaways:
- A nurse-led, AI-assisted digital platform reduced disease activity and improved physical function more than routine care over six months.
- People using the platform showed higher medication adherence and markedly greater satisfaction with their care.
- Real-time monitoring enabled earlier detection of problems and more personalised support between clinic visits.
The challenge of care after discharge
Rheumatoid arthritis is a long-term autoimmune condition that causes joint pain, swelling and a progressive loss of function. Managing it well once people leave hospital is often difficult, because symptoms can fluctuate and regular follow-up is not always easy to arrange. A recent real-world study set out to test whether an artificial intelligence (AI)-assisted digital care platform could improve outcomes for people living with the condition after discharge.
How the study was designed
The study, published in JMIR Medical Informatics and conducted by Ziyun Zhang, PhD, and colleagues at Tongji Hospital, followed 341 people with rheumatoid arthritis over a six-month period in a real clinical setting.
Participants were divided into two groups. One group received standard post-discharge care, while the other used a nurse-led digital management platform supported by AI. The platform allowed people to report symptoms, fatigue, medication use, laboratory results and emotional wellbeing through a smartphone app. This information was stored securely and analysed in real time. When the system detected concerning changes, healthcare staff were alerted so that they could respond quickly. Nurses and health coaches also provided ongoing education and personalised support.
The researchers focused on how disease activity, physical function, medication adherence and satisfaction changed over time, using standard clinical tools to measure disease severity and disability.
What the platform achieved
After six months, both groups showed some improvement, but the differences between them were notable. People using the AI-supported platform experienced a greater reduction in disease activity scores, meaning their arthritis was better controlled. They also showed significant improvements in physical function compared with those receiving routine care alone.
Medication adherence was higher in the digital care group, with more people taking their medicines as prescribed. Satisfaction levels were significantly higher too, with a large majority of those using the platform reporting that they were very satisfied with their care experience, compared with the standard care group.
What it means for long-term management
The authors conclude that the combination of AI monitoring, nurse-led support and continuous digital engagement helped to improve both clinical outcomes and the experience of care. The system made it easier to detect problems early, encourage medication use and provide more personalised care between clinic visits.
Overall, the study suggests that digital health platforms could play an important role in improving the long-term management of rheumatoid arthritis, particularly by keeping people more closely connected to their care teams after they leave hospital.
Read More
Digital Health Tools Show Early Promise for Infant Feeding and Sleep, UMass Chan Research Finds
Key Takeaways:
- Families who completed three or more visits with the virtual feeding service SimpliFed provided breast milk for nearly 16 weeks longer than those who did not use it.
- Infants whose parents engaged most actively with the AI-powered sleep app Huckleberry slept around 90 minutes longer during their longest overnight stretch.
- Lower uptake among Spanish-speaking and publicly insured families underlines the need to make digital health support equitable rather than exclusionary.
Studying whether technology can support new parents
Researchers at UMass Chan Medical School are investigating whether digital tools for infant feeding, sleep and other early parenting challenges can improve health outcomes and widen access to support for families.
Among them is Nisha Fahey, DO, MSc’21, assistant professor of pediatrics and principal investigator on research examining how digital health interventions can support families during the critical first year of a child’s life. Dr Fahey has led two pilot studies evaluating virtual lactation support and an artificial intelligence–powered infant sleep application, carried out in collaboration with the Department of Medicine’s Program in Digital Medicine. That programme is led by Apurv Soni, MD, PhD’21, assistant professor of medicine and the programme’s co-director, who serves as multiprincipal investigator on the work.
As a paediatrician, Dr Fahey hears the same questions from new parents every day: Is my baby feeding enough? Are they sleeping enough? And where can I turn for help when I need it?
“These technologies already exist. Families are accessing them and using them,” said Fahey. “As researchers and healthcare providers, it’s our responsibility to understand their impact and think about how they can be integrated into healthcare in a way that is equitable and reaches all families.”
Virtual feeding support and longer breastfeeding
The first study examined SimpliFed, a virtual infant-feeding support platform that gives families on-demand access to certified lactation consultants and feeding specialists. Researchers enrolled 200 pregnant and postpartum individuals through UMass Memorial Health’s obstetrics clinics and followed them through the first year of their infant’s life.
The study assessed infant growth and development, maternal mental health, healthcare utilisation and feeding practices. Researchers found that participants who completed three or more visits with SimpliFed provided breast milk for nearly 16 weeks longer than participants who did not use the service.
The findings also drew attention to important equity considerations. Uptake was lower among Spanish-speaking families and among publicly insured participants, underscoring the need to ensure that digital health interventions reach populations that have historically faced barriers to care.
“If health systems are going to deploy these tools broadly, we need to pay special attention to making sure all patients and families can access them,” Fahey said. “The goal is to close gaps in care, not widen them.”
An AI sleep app and longer overnight rest
A second pilot study evaluated Huckleberry, a mobile app that allows parents to track infant sleep and uses artificial intelligence to predict optimal nap and bedtime schedules. This study was funded by an NIH grant focused on point-of-care technologies for heart, lung, blood and sleep disorders.
The study enrolled approximately 80 families with infants under 12 months who are beneficiaries of UMass Memorial’s MassHealth Accountable Care Organization. Participants used the app for three months while researchers tracked engagement and measured infant sleep, parental sleep and parental mental health.
Among families who engaged most actively with the app, infants experienced longer consolidated overnight sleep. Researchers found that infants in the high-engagement group slept approximately 90 minutes longer during their longest stretch of overnight sleep, compared with participants who used the app less frequently.
The researchers also found that families in a population often underrepresented in digital health research were willing and able to engage with the technology. About half of participants were classified as highly engaged users, and most reported that they found the app useful and would recommend it to other families.
Recognising the limitations
The studies also revealed some limitations. While many families reported positive experiences, others described challenges with tracking data consistently or navigating app features while caring for a young infant.
For Dr Fahey, those findings reinforce the importance of viewing digital health as a complement to, rather than a replacement of, traditional care.
“Digital technologies offer an on-demand pathway for information and support,” she said. “The goal is to make both digital and in-person care as accessible as possible and empower families to choose what works best for them.”
Building evidence for the future of care
The research was made possible through collaborations across UMass Chan, including faculty in the Program in Digital Medicine, the Department of Obstetrics & Gynecology, the Department of Psychiatry & Behavioral Health, and the Department of Pediatrics.
“Parents are seeking out digital health apps on their own,” Fahey said. “Building evidence around their benefits and understanding their limitations helps us determine whether they can become trusted parts of care in the future.”
Source: UMass Chan Medical School
Read More