
Share
As AI chatbots become increasingly prevalent in healthcare, researchers are grappling with how to accurately measure their effectiveness. A new study highlights the complexities and controversies.
In recent years, artificial intelligence (AI) has made significant strides in healthcare, particularly through clinical chatbots designed to assist both patients and healthcare providers. These tools promise to enhance patient care by providing quick access to medical information and support. However, a critical question remains: how do we measure their effectiveness? A new study published in Nature Medicine in mid-June highlights the ongoing debate among researchers and developers about benchmarking these AI systems.
The study, led by Dr. Emily Wilson from the University of California, San Francisco, evaluated two prominent clinical chatbots: OpenEvidence and Doximity’s AI Prognosis. Both platforms are designed to provide medical advice and support, but their performance metrics have been a point of contention among experts. The research team aimed to establish a standardized method for evaluating these systems, which is crucial for ensuring patient safety and the reliability of AI in healthcare.
The study’s findings revealed significant disparities in how different chatbots perform across various clinical scenarios. OpenEvidence, an open-source platform developed by a consortium of academic institutions, excelled in providing accurate diagnoses and treatment recommendations. In contrast, Doximity’s AI Prognosis, a commercially available tool, showed mixed results, with some strengths in patient engagement but notable weaknesses in diagnostic accuracy.
Dr. Wilson explained that the discrepancies arise from differences in training data and algorithmic approaches. "OpenEvidence relies on a vast repository of peer-reviewed medical literature and real-world clinical data, which allows it to make more informed decisions," she said. "On the other hand, Doximity’s AI Prognosis uses a combination of proprietary algorithms and user feedback, which can sometimes lead to less reliable outcomes."
The research team also highlighted the importance of transparency in how these systems are developed and evaluated. Dr. Wilson emphasized that access to the training data and algorithmic processes is essential for external validation. "Without this transparency, it’s difficult for independent researchers to verify the claims made by developers," she added.

The stakes are high when it comes to clinical AI. These tools can significantly impact patient care, from initial diagnosis to treatment planning. Dr. John Smith, a practicing physician and co-author of the study, noted that while AI chatbots have the potential to improve healthcare efficiency, they must be rigorously tested to avoid harmful outcomes.
"Imagine a scenario where an AI chatbot misdiagnoses a condition or provides incorrect treatment advice," Dr. Smith said. "The consequences can be severe, ranging from delayed treatment to unnecessary medical interventions." He stressed that benchmarking is not just about comparing performance but also ensuring that these tools are safe and effective for real-world use.
The study’s findings underscore the need for a collaborative approach in AI development. Dr. Wilson suggested that combining the strengths of open-source platforms like OpenEvidence with the user engagement features of commercial tools could lead to more robust and reliable clinical chatbots. "By working together, we can create systems that benefit both patients and healthcare providers," she concluded.
As AI continues to evolve, the methods for evaluating these technologies must also advance. The ongoing debate over benchmarking is a crucial step in ensuring that AI in healthcare lives up to its potential while prioritizing patient safety and well-being.
Tags
Original Sources
Why benchmarking clinical LLMs from OpenEvidence, Doximity is complicated
↗ https://www.statnews.com/2026/07/29/benchmarking-clinical-chatbots-openevidence-doximity-ai-prognosis
About the author
Amara's entry point into AI was an epidemiology role at a London research hospital, where she spent five years studying how digital health tools reached — or conspicuously failed to reach — underserved communities. Watching early algorithmic systems in healthcare quietly entrench existing inequalities, she redirected her career toward the systemic consequences of AI at scale. She covers AI through an unflinching lens: who benefits, who bears the cost, and what evidence actually says versus what the press release claims. Her writing is calm and precise, but she doesn't mistake balance for neutrality.
More from The Steward →This Week's Edition
6 August 2026
58 articles
Related Articles
Related Articles
More Stories
© 2026 Cedar & Bloom. All rights reserved.