
People Like AI Mental Health Chatbots. Whether They Help Is Another Question
Key Takeaways:
- Across 21 studies in 11 countries, people using generative AI mental health chatbots reported high satisfaction and found them convenient and accessible.
- Personalisation and empathy did not reliably translate into better clinical outcomes, and engagement often faded over time.
- The evidence base remains early-stage, leaving safety, equity and crisis response unresolved.
A treatment gap that digital tools are being asked to fill
Generative artificial intelligence (GenAI) chatbots designed to support mental health are winning people over on experience, but the research needed to establish whether they are safe and clinically effective has not kept pace. That is the central finding of a review, currently in press in the journal npj Digital Medicine, which examined the user experience (UX) and intervention design of GenAI mental health chatbots.
The context for this work is a substantial and persistent shortfall in care. Around 25% of people worldwide experience a mental health problem, yet approximately 85% do not receive adequate treatment. The reasons are varied and overlapping: stigma, cost, shortages of trained professionals, geographic distance from services and structural inequities, among others. As the prevalence of mental health conditions has grown while treatment gaps have remained, attention has turned towards innovative models of delivery – and digital tools, with their scalability and convenience, have become a focus of that search.
Digital mental health interventions deliver treatment or support through a range of channels, including chatbots, websites, mobile applications and wearables. Conversational agents, more commonly known as chatbots, are applications that simulate human dialogue using machine learning and natural language processing algorithms.
From scripted responses to open-ended conversation
Traditional mental health chatbots deliver pre-scripted therapeutic content through rules-based or retrieval-based systems. Their strength is predictability, but that same design limits their capacity to personalise support or to recognise what an individual actually needs in the moment.
Chatbots built on large language models (LLMs) work differently. They can simulate core aspects of a therapeutic encounter, including personalised suggestions and empathetic reflections. That flexibility comes with a trade-off. GenAI systems may produce responses that are incorrect or inappropriate, and their open-ended conversational capacity makes intervention design both more consequential and more complex than it is for rules-based systems. When a system can say almost anything, design decisions carry considerably more weight.
How the review was carried out
The researchers set out to map the design characteristics and UX outcomes of interventions involving GenAI mental health chatbots. They began with a systematic literature search to identify studies covering the design and deployment of such tools. Reviews, editorials, media articles and commentaries were excluded.
In total, 21 studies were selected, conducted across 11 countries between 2023 and 2025. The largest numbers came from China and the United Kingdom, followed by the United States. The included work spanned a wide range of maturity, from early-stage prototype evaluations through to clinical trials, with one real-world implementation study.
Most studies recruited general or clinical adult populations, including older people living with dementia. Others involved simulated users or university students. Sample sizes ranged from as few as five participants to as many as 527. Across the body of evidence, there was substantial heterogeneity in outcome measures, and most interventions remained at an early stage of development – two features that shape how much can reasonably be concluded from the literature as it stands.
What the interventions were designed to do
Chatbot interventions most often targeted depression and anxiety, and tended to adopt shared therapeutic mechanisms, including mindfulness, emotion regulation and cognitive restructuring. Some systems were oriented towards mental well-being, stress and loneliness, emphasising general support and preventive care rather than treatment for a specific condition. Others addressed eating disorders, post-traumatic stress disorder and dementia.
The dementia-focused interventions are worth distinguishing. Rather than attempting to address the central neurological features of the condition, they targeted its related psychological dimensions – carer burnout, psychological distress and loneliness among them.
Most interventions were grounded in cognitive-behavioural therapy principles. The specific techniques drawn upon included behavioural activation, psychoeducation, Socratic dialogue, acceptance and commitment therapy, cognitive restructuring and mindfulness.
How the tools were delivered
Interventions varied in frequency, delivery modality and duration. The majority were short-term, running from two to eight weeks. Most were deployed through web-based interfaces and mobile applications, while some were delivered via messaging or social media platforms – meeting people on services they already used rather than asking them to adopt something new.
About 67% of interventions were non-embodied, text-based chatbots. The remainder used voice, avatar-based, augmented reality or other multimodal forms of interaction, with the intention of improving engagement and realism.
What people made of them
All but two of the studies evaluated at least one UX domain. The majority relied on quantitative measures, typically Likert scales, while some gathered qualitative feedback through open-ended questions and semi-structured interviews.
User satisfaction and acceptability were the most commonly reported outcomes. Across studies, participants described the interventions as convenient and accessible, with acceptability generally rated moderate-to-high and reported satisfaction high.
Half of the studies examined usability, using qualitative feedback, the System Usability Scale or Likert scales. Interface design, interaction mode and deployment platform were all observed, alongside differences in usability between studies. A clear preference emerged for free-flowing chat interfaces and customisable features over predefined options. At the same time, some interventions had an unclear scope or limited functionality, leaving people uncertain about what the chatbot could actually do for them.
Usability, engagement and the drop-off problem
Only some studies reported objective utilisation and engagement metrics, such as session frequency, interaction duration, retention over time and task completion. Where these were captured, attrition patterns frequently emerged over time in repeated-measures designs. Uptake in multi-week interventions was often high at the outset before declining – a pattern familiar across digital health more broadly, and one that matters a great deal for interventions whose therapeutic logic depends on sustained practice.
Personalisation and perceived benefit
Most chatbots featured some form of personalisation, reflecting their capacity to adapt conversations and interfaces in response to previous interactions and a person’s emotional state. The most common approach was emotion detection paired with adaptive interaction, allowing people to receive tailored responses and empathetic reflections.
Perceived impact was not consistently measured as a standalone metric. It was more usually folded into qualitative feedback or broader UX evaluations. In the intervention that produced the most granular data, the most frequently reported benefit was improved clarity and awareness.
Where empathy stops being enough
The review’s more cautionary finding is that personalisation and empathy did not consistently translate into stronger clinical outcomes or sustained use. Feeling supported and being helped are not the same thing, and the studies reviewed do not yet demonstrate a reliable link between the two.
Some people reported responses that were repetitive, generic or contextually misaligned. Others raised concerns about over-reliance on chatbots, reduced human contact, data privacy and whether these systems can respond appropriately when someone is in crisis. Inaccurate or clinically misaligned outputs were also linked to an erosion of trust and to disengagement in several studies.
For healthcare professionals, the practical question is less whether these tools have promise than how to appraise them – knowing what a given system is grounded in, where its limits sit and when a conversation needs to move to a human. That judgement is increasingly treated as a core clinical competency, and it sits at the centre of CPD training on the everyday, ethical use of AI in practice.
Design features linked to a better experience
The authors identified several design features associated with better UX outcomes, while being careful to note that these were associations rather than demonstrated causes. They included:
- Deployment on platforms people already knew and used
- Richer interaction modalities beyond plain text
- Integration into existing care pathways
- Personalisation
- Grounding in domain knowledge
- Structured delivery
- Proactive outreach
- Co-design with both experts and end users
The predominance of early-stage studies, combined with limited direct comparative analyses, prevented firm conclusions about which of these features genuinely improved user experience.
What needs to happen next
Taken together, the review suggests that GenAI chatbots have meaningful potential to deliver tailored, empathetic mental health support, and that their acceptability among the people who use them is promising. That is a real finding, and not a small one given the scale of unmet need.
Significant challenges remain, however. Standardising how UX is assessed, grounding intervention design in the needs and preferences of the people who will use these tools, and sustaining engagement beyond the first few weeks all stand out as unresolved. Addressing them, the authors argue, will require co-design with experts and users, validated UX metrics applied in long-term studies, transparent reporting standards, independent evaluation, clearer reporting of model design and training data, and stronger attention to safety, equity and the limits of crisis response.
CCH insight
Generative AI tools are arriving in patient-facing care faster than the evidence base supporting them, which puts the burden of appraisal on clinicians. CCH’s CPD-accredited short course AI Essentials for Primary Care: Tools, Ethics and Everyday Applications covers the practical and ethical judgement this requires – what these tools can and cannot do, where the risks sit, and how to use them safely in day-to-day practice.
Find out more about AI Essentials for Primary Care →

AI Diet Recommendations for Adolescents Show Significant Nutritional Gaps, Study Finds
Key Takeaways:
- AI-generated diet plans consistently underestimated energy and key macronutrients required by adolescents
- Macronutrient balance was frequently misaligned with clinical guidelines, with lower carbohydrates and higher fat and protein levels
- Researchers caution that AI tools should not replace dietitians for adolescent nutrition without professional oversight
Growing demand for accessible nutrition support
Artificial intelligence is increasingly being used to support dietary planning, particularly in areas where access to qualified professionals is limited. However, a new study published in Frontiers in Nutrition raises important concerns about the reliability of these tools when applied to adolescents living with overweight or obesity.
Globally, adolescent overweight and obesity are rising at pace, affecting an estimated 390 million young people in 2022. In many regions, this now represents the most common form of malnutrition. Excess body weight in adolescence is associated with a range of adverse health outcomes, including type 2 diabetes, dyslipidaemia, hypertension, and sleep apnoea. It also increases the likelihood of obesity in adulthood and is linked to reduced quality of life.
Alongside physical health risks, adolescents may experience body image concerns and engage in harmful weight control behaviours such as self-induced vomiting or misuse of laxatives.
Dietary modification remains central to improving outcomes. Dietitians play a key role in delivering tailored, evidence-based nutrition plans aligned with established guidelines. However, limited access and workforce pressures can restrict the availability of personalised support.
AI tools, including chatbots and large language models, are increasingly being explored as a way to bridge this gap. While they can provide general dietary guidance, concerns remain about their accuracy, safety, and ability to replicate the individualised care provided by trained professionals.
Study design – comparing AI models with dietitian plans
To better understand the role of AI in adolescent nutrition, researchers conducted a direct comparison between AI-generated diet plans and those created by a dietitian.
Five AI systems were evaluated: ChatGPT-4o, Gemini 2.5 Pro, Claude 4.1, Bing Chat-5GPT, and Perplexity. Across two sessions, these models generated a total of 60 diet plans. Each plan covered three days and was based on four standardised adolescent profiles, including boys and girls living with overweight or obesity.
These AI-generated plans were compared with dietitian-designed one-day plans developed in line with established nutritional recommendations. The reference plans followed a macronutrient distribution of:
- 45–50 % carbohydrates
- 30–35 % fat
- 15–20 % protein
The researchers then analysed energy intake, macronutrient composition, micronutrient content, safety, and feasibility.
Consistent underestimation of energy and macronutrients
The findings revealed a clear and consistent pattern across all AI models. Diet plans generated by AI underestimated both total energy intake and key macronutrients when compared with dietitian-designed plans.
On average:
- Energy intake was lower by 695 kcal
- Protein intake was reduced by 20 g
- Fat intake was reduced by 16 g
- Carbohydrate intake was reduced by 115 g
Given the high energy demands of adolescence, such deficits could have meaningful clinical implications, particularly for growth, development, and overall health.
Macronutrient imbalance – a shift away from guidelines
Beyond total intake, the balance of macronutrients was also significantly altered in AI-generated plans.
Some AI models recommended:
- Protein intake up to 23.7 %
- Fat intake up to 44.5 %
Both values exceeded recommended levels. In contrast, carbohydrate intake accounted for no more than 36.3 %, falling below guideline recommendations.
Dietitian-designed plans, by comparison, remained closely aligned with clinical standards:
- Carbohydrates: 44 %–46 %
- Protein: 18 %–20 %
- Fat: 36 %–37 %
The authors noted:
“This pattern illustrates a systematic shift across all AI models to lower CHO, higher protein, and higher lipid meal structures, indicating that the macronutrient balance, not just the amount of gram-based nutrients, is significantly disrupted in AI-generated plans.”
Researchers suggest that AI models may be influenced by popular dietary trends, such as low-carbohydrate or ketogenic approaches, rather than evidence-based adolescent nutrition guidelines. This shift may pose risks during a critical period of physical and cognitive development.
Micronutrient variability raises additional concerns
In addition to macronutrient discrepancies, the study identified significant variability in micronutrient composition across AI-generated plans.
No model consistently matched the dietitian-designed reference diet across all nutrients. This inconsistency raises concerns about potential micronutrient deficiencies, which could further compromise adolescent health.
The findings suggest that AI tools currently lack the technical precision required to accurately estimate both macro- and micronutrient needs in personalised dietary plans for adolescents.
Strengths and limitations of the study
The study offers several notable strengths. It evaluated multiple AI models, allowing for robust comparison across systems. The use of three-day diet plans enabled identification of consistent patterns rather than isolated outputs. Dietitian-designed plans provided a credible clinical benchmark, and the inclusion of both macro- and micronutrient analysis allowed for a comprehensive assessment of dietary quality.
However, there are limitations to consider. The findings are specific to the models tested, which are rapidly evolving. Standardised adolescent profiles may not fully capture real-world complexity, limiting personalisation. The use of simulated scenarios rather than real-life behaviours may reduce ecological validity. Additionally, prompts were standardised and delivered in a single language, which may limit generalisability across populations.
Implications for clinical practice and AI use
The study highlights important risks associated with the unsupervised use of AI for adolescent dietary planning.
As the authors conclude:
“AI models have exhibited clinically significant deviations in diet plans for adolescents at both macro and micro levels.”
These deviations include consistently lower energy and carbohydrate recommendations compared with dietitian-designed plans.
Until these limitations are addressed, AI-generated diet plans should be used with caution. They may serve as a supplementary tool under professional supervision, but they are not currently a safe or reliable substitute for qualified dietary guidance in adolescents.
Read More
AI in Healthcare: Promise, Pitfalls, and the Risk of Misguided Medical Advice
Key Takeaways:
- People using AI for health advice often struggle to interpret and communicate symptoms effectively, leading to incorrect conclusions in many cases.
- Even when AI identifies a condition correctly, it may fail to recommend appropriate urgency, particularly in time-sensitive or complex scenarios.
- Clinicians see value in AI as a supportive tool, but stress that it should complement, not replace, professional medical care.
AI becomes a common source of health information
As technology companies continue to develop platforms tailored for healthcare consultation, artificial intelligence is becoming an increasingly influential part of how people make decisions about their health. According to OpenAI, more than 40 million people use ChatGPT each day to seek health-related information.
However, emerging research suggests that while these tools offer unprecedented access to medical knowledge, they may also mislead users in certain contexts.
Challenges in how people use AI for medical queries
One of the central issues identified by researchers is not only the capability of AI systems, but how individuals interact with them. Many people lack the knowledge required to communicate symptoms accurately or comprehensively.
A recent study published in Nature Medicine attempted to replicate real-world use of AI chatbots. Participants were given medical scenarios and asked to consult AI tools. The results highlighted notable limitations:
- Participants correctly identified the condition only about one-third of the time.
- Just 43% made the correct decision regarding next steps, such as whether to seek emergency care or remain at home.
“People don’t know what they are supposed to be telling the model,” says Andrew Bean, who studies AI systems at Oxford University and was one of the authors on this study.
Bean explains that effective use of AI often depends on precise wording. “Doctors are trained to ask you questions about symptoms you might not have realised you should have mentioned,” says Bean.
Small differences in input can lead to dangerous outcomes
The study demonstrated how subtle differences in language can significantly alter the advice provided by AI systems.
In one example, two individuals described the same clinical scenario slightly differently. One described experiencing “the worst headache I’ve ever had” and was advised to go to the emergency room immediately. The other, who did not include that specific phrasing, was advised to take aspirin and remain at home.
“Turns out this was actually a life-threatening condition,” says Bean.
This highlights a critical limitation: AI systems rely heavily on the information they are given, and may not prompt for missing but clinically important details in the way a trained clinician would.
When AI gets the diagnosis right but the advice wrong
Even when AI tools successfully identify a medical condition, they may still provide inappropriate guidance regarding urgency or next steps.
In a separate study, researchers evaluated how AI systems responded to a range of medical scenarios. They found that in 52% of emergency cases, the tools “under-triaged” – treating conditions as less serious than they actually were.
In one case, the AI failed to direct a hypothetical patient experiencing diabetic ketoacidosis and impending respiratory failure – both life-threatening conditions – to seek emergency care.
“When there was a textbook medical emergency, ChatGPT got it right,” said Girish Nadkarni, a doctor and AI researcher at Mount Sinai who is an author on the study. However, he noted that performance declined in more complex situations, particularly where timing was critical. In such cases, the system often misjudged how urgently care was required.
An OpenAI spokesperson responded by stating that the study did not reflect typical real-world usage and that it evaluated an older version of ChatGPT, which the company says has since been improved to address some of these concerns.
The role of AI in supporting patient understanding
Despite these concerns, many clinicians believe that AI tools can still play a constructive role in healthcare, particularly in improving patient understanding and engagement.
“I encourage patients to use these tools,” says Robert Wachter, a doctor at UC San Francisco and author of the recently published book, A Giant Leap: How AI Is Transforming Health Care and What That Means for Our Future.
Wachter points out that barriers to accessing healthcare – including cost and availability – mean that AI can sometimes provide a useful alternative source of information. “The advice you get from the tools is substantially better than nothing and better than what you would get from your second cousin,” says Wachter.
However, he emphasises that AI should never be viewed as a substitute for professional medical care.
Enhancing, not replacing, the doctor–patient relationship
Experts suggest that AI is most valuable when used alongside traditional healthcare, rather than in place of it.
Adam Rodman, a hospitalist and researcher at Harvard Medical School, advises against using AI tools to assess emergency situations. Instead, he sees their greatest benefit in preparing for or reflecting on medical consultations.
“A good time to use a large language model is when you’re about to go see a doctor – or after you see your doctor,” says Rodman.
He explains that AI can help people better understand their condition, ask more informed questions, and make more effective use of time during appointments. This can support a more collaborative relationship between patients and clinicians.
“There are no downsides to better understanding your health,” says Rodman.
The future of AI in healthcare
Healthcare professionals broadly agree that AI is now firmly embedded within modern medicine and will continue to evolve alongside clinical practice.
“ My hope is that you might see AI as an extension of a human relationship,” says Rodman. He envisions a future in which both clinicians and patients work with AI to improve communication and navigate healthcare systems more efficiently.
However, he also raises concerns about potential unintended consequences. One particular risk is the possibility that people may receive serious or distressing diagnoses – such as cancer – directly from an AI system, rather than from a clinician.
Research suggests that when healthcare becomes more transactional or resembles a marketplace, trust in clinicians may decline.
”What I hope is that this technology can be used in a way that enhances humanity in medicine,” says Rodman “and not in a way that cuts out the doctor-patient relationship.”
Conclusion
Artificial intelligence is rapidly transforming access to health information, offering both opportunities and risks. While these tools can enhance understanding and support more informed decision-making, their limitations – particularly in how they interpret incomplete or imprecise input – mean they must be used with caution.
Ultimately, AI has the potential to strengthen healthcare delivery, but only if it is integrated in a way that supports, rather than replaces, the human relationships at the heart of medicine.
Read More
‘Shadow AI’ on the Rise in Healthcare as Clinicians Turn to Unauthorised Tools to Improve Workflows
Key Takeaways:
- A survey of healthcare professionals found that 57% have encountered or used unauthorised artificial intelligence tools in their organisations, highlighting the growing presence of so-called “shadow AI” in healthcare settings.
- Many clinicians and administrators report using these tools to improve efficiency, analyse data, and manage administrative tasks, particularly when approved solutions or clear guidance are lacking.
- While most respondents believe AI will significantly improve healthcare within five years, concerns about patient safety, data privacy, and security risks remain widespread.
Unauthorised AI tools emerging in healthcare workplaces
A new survey suggests that artificial intelligence tools are already being used in healthcare organisations in ways that fall outside formal governance structures. According to the findings, a significant proportion of healthcare professionals have either encountered or used AI tools that have not been authorised by their employer.
The survey, conducted by Wolters Kluwer Health, gathered responses from 518 healthcare professionals, including both clinical providers and administrators. The research was carried out in December 2025 and was released publicly last week.
Overall, the findings indicate that four in ten healthcare professionals reported encountering unauthorised AI tools within their organisation, while 17% acknowledged personally using such tools.
When responses were analysed by professional role, 15% of physicians admitted to using an unauthorised AI tool, compared with 19% of administrators. In addition, one in ten respondents reported using an unauthorised AI tool in connection with direct patient care.
The report refers to the unauthorised adoption of artificial intelligence tools in professional environments as “shadow AI.”
Why healthcare staff turn to unauthorised AI
The survey findings suggest that healthcare professionals are often motivated by practical needs rather than deliberate attempts to bypass organisational policies.
According to the report:
“Clinical and administrative teams want to adhere to rules surrounding AI usage, but if the organization hasn’t provided guidance or approved solutions, they’ll experiment with generic tools to improve their workflows.”
Many respondents indicated that the absence of formal guidance or approved AI platforms has encouraged individuals to explore publicly available tools on their own.
The most frequently cited motivation for using unauthorised AI tools was the need to accelerate workflows and improve efficiency. Approximately half of respondents identified faster workflows as the primary reason for using these tools.
However, the survey also revealed differences in how clinical and administrative staff tend to use AI technologies.
Administrators were more likely to employ AI tools for operational or analytical tasks such as:
- Data analysis
- Predictive analytics
- Administrative processes
Healthcare providers, meanwhile, reported using AI for activities such as:
- Data analysis
- Patient scheduling
- Patient engagement tasks
The findings also indicate that clinicians were more likely than administrators to experiment with AI tools out of curiosity.
Governance and policy development remain uneven
The survey results highlight a notable imbalance in how different professional groups participate in the development of AI policies within healthcare organisations.
According to the report, administrators were three times more likely than clinical providers to be actively involved in developing AI governance policies.
Specifically:
- 30% of administrators reported involvement in AI policy development
- Only 9% of providers said they had participated in such efforts
This difference suggests that policy ownership around AI adoption may currently be concentrated within administrative leadership rather than clinical teams.
Administrators also reported greater familiarity with their organisation’s AI policies compared with providers, although awareness varied across both groups.
Security and privacy risks associated with “shadow AI”
The use of unauthorised AI tools raises important concerns about data security, privacy protection, and governance oversight.
The Wolters Kluwer report notes that inconsistent or unsanctioned AI usage can expose organisations to potential vulnerabilities. Without clear oversight, the integration of external AI tools may lead to data privacy violations, security breaches, or inappropriate handling of sensitive information.
To illustrate these risks, the report references a 2025 study by IBM, which found that 97% of organisations that experienced an AI-related security incident lacked adequate AI access controls.
Security incidents involving AI systems can have significant consequences, including financial losses, operational disruption, and damage to public trust.
Healthcare professionals remain optimistic about AI’s future
Despite concerns about governance and security, the survey indicates that most healthcare professionals remain broadly optimistic about the long-term role of artificial intelligence in healthcare.
Nearly 90% of respondents said they believe AI will significantly improve healthcare within the next five years. Administrators were found to be slightly more optimistic than clinical providers about the potential benefits of the technology.
At the same time, respondents recognised that AI implementation carries important risks that must be addressed.
Patient safety was identified by around half of respondents as the most significant risk associated with AI adoption.
Meanwhile, nearly half of respondents also expressed concerns about data privacy risks.
These findings suggest that healthcare professionals recognise both the transformative potential of artificial intelligence and the need for careful governance, clear guidance, and secure systems.
Addressing the rise of “shadow AI”
The report concludes that addressing the growth of shadow AI requires organisations to understand why staff are turning to unauthorised tools rather than focusing solely on restricting access.
According to the report:
“Ultimately, addressing shadow AI is not about restricting access to productivity tools. Leaders must understand why teams are using unsanctioned tools and which challenges they’re trying to solve, and then identify enterprise-level tools that can accomplish these goals safely and securely.”
As artificial intelligence becomes increasingly embedded in healthcare workflows, organisations may need to develop clearer policies, provide approved tools, and involve both clinical and administrative staff in governance decisions.
Such measures may help ensure that the benefits of AI can be realised while protecting patient safety, safeguarding sensitive data, and maintaining organisational trust.
Read More
Study Highlights Benefits and Limits of Generative AI in Weight Management
Key Takeaways:
- A short field experiment suggests that generative AI can support modest reductions in weight and body mass index through personalised dietary feedback.
- Private use of AI tools appears more effective than public sharing, with public analysis associated with higher dropout rates.
- People with lower levels of nutritional knowledge benefited most, indicating potential for AI to help reduce health inequalities, although it does not replicate the value of human community support.
Introduction
Nearly three-quarters of adults in the United States are living with overweight or obesity, and prevalence continues to rise globally. As a result, demand for high-cost interventions such as bariatric surgery and glucagon-like peptide-1 medications has increased, placing significant financial pressure on health care systems.
A new working paper suggests that generative artificial intelligence may offer a low-cost way to support people with weight loss by helping them make more informed dietary choices. However, the research also indicates that AI tools do not replicate the benefits of community-based programmes where people can share experiences and openly discuss the physical and psychological challenges associated with obesity.
The study was conducted by Catherine Tucker, Professor of Marketing at MIT Sloan School of Management, and Linyi Li of Singapore Management University. They followed 416 adult participants of varying ages over a three-week period in late 2024.
Study design and intervention
The researchers partnered with an Asia-based Fortune 500 company that runs an online weight loss boot camp combining guidance on healthy eating and physical activity. The programme included a group chat function using WeChat, enabling participants to interact, share experiences and support one another.
Participants were divided into three groups to assess the impact of a generative AI tool designed to analyse meals. The tool evaluated the nutritional content of food based on photographs and provided real-time, personalised suggestions such as adding more vegetables or choosing leaner protein sources.
The three groups were structured as follows:
- Group 1 – control group: Participants received general healthy-diet tips and access to the group chat but did not use the AI food-analysis tool.
- Group 2 – private analysis group: Participants sent photos of their meals privately to an administrator and received personalised AI-generated nutrition reports.
- Group 3 – public analysis group: Participants shared meal photos within the group chat, where both the images and the AI-generated nutrition reports were visible to all group members.
Finding 1 – Generative AI supported weight loss
Compared with the control group, both groups that used the AI food-analysis tool showed higher engagement with the programme, greater weight loss and larger reductions in body mass index.
On average, participants in Group 1 lost 0.966 kg over the three-week period. Those in Group 2 lost 1.426 kg, while participants in Group 3 lost 1.358 kg.
Although the absolute numbers were modest, Tucker emphasised their significance given the short duration of the intervention.
“Weight loss is such a big challenge. If it were easy for us all to lose weight, we’d just lose weight,” Tucker said. “The fact that a digital tool such as AI can have any effect is wonderful because interventions such as surgery or injectables are expensive. This is evidence of the cost efficacy of a very small intervention in terms of changing behavior.”
According to Tucker, the results highlight the value of generative AI in personalising individual experiences by offering tailored feedback, practical knowledge and guidance on day-to-day dietary decisions.
Finding 2 – Public analysis reduced participation
The way in which the AI tool was used had a clear impact on engagement. Participants with private access to the food-analysis tool were significantly more likely to remain in the programme for the full three weeks.
In contrast, Group 3, where meal photos and AI feedback were shared publicly, had the highest dropout rate. Tucker suggested that some participants may have felt discouraged by seeing highly engaged or high-performing peers, leading to disengagement.
“Dropout is the big enemy of weight loss,” Tucker said. “A likely explanation [for dropouts in Group 3] is that staying in the group introduced pressure [when] consistently reporting less-favorable statistics compared to others.”
The findings suggest that making AI-generated feedback public may alienate some individuals and reduce sustained participation. Community-based programmes such as Weight Watchers have historically succeeded by fostering mutual support during both successful and challenging periods.
As Tucker noted,
“There’s a set of people there to support you through good or bad weeks. I think what we are demonstrating is that if you make it too easy to post success stories, then you lose some of that [shared] vulnerability within the community.”
Finding 3 – Potential to reduce health inequalities
The researchers also found that the greatest benefits from the AI tool were seen among participants with lower levels of education and less prior nutritional knowledge. These individuals often struggle to interpret standard weight loss advice and appeared to gain particular value from detailed, personalised recommendations generated by the AI system.
The authors suggest that this capability could help reduce health inequalities by improving access to understandable, tailored dietary guidance for people who may otherwise be disadvantaged by traditional educational approaches.
Implications for the use of AI in health behaviour change
Although the study focused specifically on weight loss, the authors argue that the findings have broader relevance for how people interact with AI systems. Generative AI appears well suited to supporting individual behaviour change through personalisation, prompts and reminders. However, it does not replicate the social connection and emotional support provided by human communities.
For organisations and programme designers, the research suggests that AI should be used to enhance individual-level support rather than as a replacement for community-building or large-scale digital ecosystems.
Although the research was conducted in China, Tucker stated that the findings are likely to be applicable in other settings.
“I think what our research shows is that in the generative AI age, technology can certainly assist with information retrieval, reminders, prompts, all those good things, but we can’t really use it to replace that sense of community,” Tucker said.
Read More