
Share
As AI ambient scribes gain traction, a new study by Suki highlights the shortcomings of existing evaluation tools, calling for a refresh to better detect LLM-specific errors.
As the adoption of AI ambient scribe technology continues to grow in healthcare, researchers at Suki are raising concerns about the industry's standard method for evaluating clinical AI notes. According to their findings, the Physician Documentation Quality Instrument (PDQI-9), which has been the go-to tool since 2012, may be fundamentally flawed when it comes to assessing modern ambient AI scribes.
Suki, a leading provider of ambient clinical AI solutions, works with over 400 health systems. The company's researchers argue that the PDQI-9 is poorly suited for evaluating AI-generated notes because it focuses on overall note quality-such as organization and conciseness-rather than identifying specific errors unique to large language models (LLMs). These errors can include hallucinations, invented medication doses, and missing diagnoses.
The PDQI-9 evaluation domains are up-to-date, accurate, thorough, useful, organized, concise, synthesized, internally consistent, and cohesive. However, these broad criteria may not be sufficient for detecting the nuanced errors that LLMs can introduce into clinical notes.
Kevin Wang, M.D., Suki's chief medical officer, emphasizes the need for a refresh in evaluation methods: "Base quality rubrics need a refresh in the post-LLM, ambient world." Suki's white paper analyzes the limitations of holistic Likert-based scoring instruments like PDQI-9 and highlights their inability to reliably evaluate AI-generated clinical notes.
In their study, Suki researchers examined 84 paired notes across four specialties. They found significant inconsistencies among reviewers, including:

These issues, according to the researchers, reflect a fundamental mismatch between traditional note-quality rubrics and the types of errors LLMs can make. PDQI-9 scores notes holistically, focusing on factors like organization, conciseness, and cohesiveness-factors that are less relevant when it comes to detecting specific AI-generated errors.
Suki's research underscores the importance of updating evaluation methods to ensure that AI ambient scribes are accurately and reliably assessed. As AI continues to play a larger role in healthcare, it is crucial to have robust tools that can detect and mitigate potential errors, ultimately improving patient care and safety.
Tags
Original Sources
As AI scribe adoption grows, researchers at Suki challenge the industry's quality playbook
↗ https://www.fiercehealthcare.com/ai-and-machine-learning/ai-scribe-adoption-grows-researchers-suki-challenge-industrys-quality
About the author
Kai built ML infrastructure at a Bay Area startup before developing an obsession with transformer architectures and inference optimisation that eventually pulled him out of product work entirely. A stint at a compute research lab sharpened his instinct for what actually matters in a model release versus what is marketing. He writes from the inside — from the perspective of someone who has debugged the systems he is describing at three in the morning. He is allergic to hype and instinctively drawn to the unglamorous plumbing questions that everyone else skips over.
More from The Engineer →This Week's Edition
6 August 2026
58 articles
Related Articles
Related Articles
More Stories
© 2026 Cedar & Bloom. All rights reserved.