
Google’s enhanced AI chatbot surpasses human doctors in diagnosing skin rashes and analysing medical images
Key Takeaways:
- Google’s experimental AI chatbot, AMIE, has demonstrated superior diagnostic accuracy compared to human physicians in simulated clinical consultations involving images and medical records.
- The updated system incorporates Google’s image-processing large language model Gemini 2.0 Flash, adapted for clinical reasoning without requiring extensive retraining.
- Experts acknowledge the chatbot’s promise but caution that real-world deployment faces significant challenges, and the system’s reproducibility remains limited due to missing technical details.
AI-Powered Diagnosis Using Images
An advanced version of Google’s medical chatbot is showing promising capabilities in interpreting images to diagnose health conditions, including skin rashes from smartphone photographs. The system can now process a broader range of medical imagery—such as electrocardiograms and laboratory reports in PDF format—enhancing its diagnostic abilities across a wider clinical spectrum.
Previously, an earlier version of the artificial intelligence (AI) model had already demonstrated superior diagnostic performance and more empathetic communication than human physicians. The recent upgrade continues this trend, outperforming doctors in interpreting both visual data and patient-reported symptoms.
Introducing AMIE: Articulate Medical Intelligence Explorer
The upgraded system is known as the Articulate Medical Intelligence Explorer (AMIE). It remains experimental and has not yet undergone peer review, with the initial findings published on 6 May on the arXiv preprint server. According to Dr Eleni Linos, Director of the Stanford University Centre for Digital Health, who was not involved in the study, the integration of visual and clinical data “brings us closer to an AI assistant that mirrors how a clinician actually thinks”.
Testing Through Simulated Consultations
To evaluate AMIE’s clinical potential, researchers organised a series of simulated consultations involving 25 trained actors portraying patients. These individuals participated in virtual consultations with both AMIE and human primary-care physicians, covering 105 medical scenarios encompassing diverse symptoms, medical histories, and accompanying images relevant to each case.
After each consultation, both AMIE and the doctors were tasked with generating a diagnosis and treatment plan. Their responses were assessed by a panel of 18 specialists across dermatology, cardiology, and internal medicine, who reviewed transcripts and written summaries.
The outcome was striking: AMIE produced more accurate diagnoses overall, and its performance was notably resilient to challenges such as poor image quality.
Built on Gemini 2.0 Flash: A Shift in Training Methods
The latest iteration of AMIE is based on Google’s Gemini 2.0 Flash, a large language model (LLM) capable of processing both text and images. Rather than retraining the model using a traditional, labour-intensive method with domain-specific datasets, researchers introduced a more efficient adaptation process. They incorporated an algorithm designed to enhance clinical dialogue and reasoning, enabling the model to better replicate the structure of a diagnostic consultation.
To refine the model’s conversational behaviour, developers prompted it to simulate complete diagnostic interactions—alternating between the roles of patient, physician, and independent evaluator. According to Dr Ryutaro Tanno, a researcher at Google DeepMind and co-author of the study, this approach allows the AI to “sort of imbue it with the right, desirable behaviours when conducting a diagnostic conversation”.
“This is much cheaper and potentially more accessible,” Tanno added, highlighting a key advantage of the method compared to previous retraining approaches.
Expert Reactions and Limitations
Simulated patient scenarios are commonly used in medical education to assess human clinical skills. However, Dr Linos cautioned that such simulations cannot fully replicate the nuances of real-world clinical care. “Physicians bring experience, intuition and the ability to physically examine a patient—elements that are hard to replicate in a simulated script,” she said.
Dr Dan Zeltzer, a digital-health researcher at Tel Aviv University, recognised the AI’s potential but expressed concerns about transparency. He noted that the research paper lacks details about the specific code and prompt configurations used, which hinders the ability of others in the field to reproduce the results or build upon them.
Deploying this type of system in everyday clinical practice remains a considerable hurdle, according to Dr Xueyan Mei, an AI scientist at the Icahn School of Medicine at Mount Sinai in New York City. “That being said, we do think large language models for diagnosis would be the way to go in the future,” she remarked.




