
Share
Promising AI demos often run on tidy data that doesn't exist in real hospitals. One data expert explains why fragmented records and imprecise coding can quietly undermine patient care and clinician trust.
Think about the last time you tried to explain a complicated medical history to a new doctor. Details get lost. Dates blur together. A condition your old doctor flagged as "possible" somehow becomes "confirmed" by the time it lands in a new chart. Now imagine an AI system trying to make sense of that same tangled history, at scale, across thousands of patients, with real consequences riding on getting it right.
That's the challenge facing hospitals as they push clinical artificial intelligence beyond flashy demonstrations and into everyday practice. According to Joseph Zabinski, senior vice president of product management at IMO Health, a company that translates physician notes into standardized medical codes, the gap between a polished pilot program and a messy real-world rollout is where many promising AI tools quietly fail.
"Many clinical AI pilots begin with a compelling demonstration and a narrowly defined use case," Zabinski said. The problem is what happens next. Those early demos "often rely on data that are too clean and don't represent the fragmentation, inconsistencies, gaps, drifts and loss of clinical context that actually characterize real-world data in practice."
In plain terms: the AI aced the test because someone handed it a cheat sheet. Real hospital data is nothing like that. It's scattered across systems, recorded in different formats, and shaped by human judgment calls that don't always translate cleanly into computer code.
Even clean data has its own hidden problem. Standard medical coding systems, the ones used across most health systems, often can't capture the fine distinctions that matter most clinically. Severity, acuity, whether a disease is worsening, whether a diagnosis is confirmed or just suspected: these nuances frequently get flattened into broader categories.
That wouldn't be so dangerous if the AI's output looked uncertain when the underlying data was thin. It doesn't. Generative AI tools produce fluent, confident-sounding answers regardless of how shaky the foundation beneath them is. "Generative AI systems can turn these inputs into fluent, plausible outputs, performing well at the surface level while making underlying problems and inconsistency harder to recognize," Zabinski explained.
Picture a student who writes a beautifully argued essay based on a misread assignment. The prose is polished. The conclusion is wrong. That's the risk with clinical AI operating on degraded data: it sounds authoritative even when it shouldn't.
Scaling introduces its own set of headaches, too. A pilot that works in a controlled environment has to eventually connect to EHR histories, lab results, medications, insurance claims, clinical notes and day-to-day workflows. That's when "limited interoperability, weak source attribution and poor workflow integration frequently emerge," Zabinski said. The system that looked flawless in a demo room starts stumbling once it meets the real complexity of a hospital's data ecosystem.

His proposed fix centers on what he calls a clinically governed semantic layer, essentially a translation and oversight system that preserves the actual clinical meaning of information as it moves between different software platforms. Pair that with clear links back to trusted source data and governance teams that include clinicians, technologists, compliance staff and operations leaders, and organizations get a much clearer picture of where a pilot needs more work before it scales.
This matters because the stakes have quietly shifted. Standardizing clinical data used to be mostly an administrative task, supporting billing codes and regulatory reports. Now AI has moved that work to the center of care delivery itself. "AI uses standardized data as a foundation for recommendations, automation, prediction, and other forms of clinical and operational decision support," Zabinski said. Those uses carry real value for patients and providers, but they demand far more precision from the underlying data than billing ever did.
When that precision is missing, the first casualty is patient care. An AI system might confuse a symptom with an actual diagnosis. It might understate how severe a condition is. It might treat a suspected illness as if it were already confirmed. "These errors can influence clinical decision support, follow-up recommendations, test ordering, treatment selection and care coordination," Zabinski warned. Because generative AI rarely signals its own uncertainty, these mistakes can go unnoticed for a long time. And once clinicians start catching the errors, even occasionally, their trust in the tool erodes fast, regardless of how well it performs otherwise.
The damage doesn't stop at the bedside. Distorted clinical meaning ripples into revenue cycle management, quality reporting and population health programs. It can produce claim denials, misplaced patients in quality cohorts and reimbursement that doesn't match the care actually delivered. Staff then spend extra hours reviewing records and fixing errors that never should have happened in the first place. Left unaddressed, those errors can even get baked into dashboards and future AI training sets, meaning today's data problems become tomorrow's inherited mistakes.
So what should hospital leaders look for before scaling a pilot program? Zabinski points to a handful of practical benchmarks. A strong deployment solves a specific, observable problem without adding extra work to already-stretched clinical staff. It draws on clinical knowledge that's been validated and can be traced back to its original sources. It preserves the specificity of the original documentation as that information flows through coding and analytics systems. And crucially, it lets clinicians and auditors actually see which pieces of information shaped a given output, so they can check the AI's reasoning rather than simply trusting it.
Lasting success also depends on who's in the room from day one. The strongest programs bring in clinicians, informaticists, data scientists, compliance officers and operations staff together, rather than treating AI deployment as purely a technical project. And they measure more than whether the output sounds convincing. They track coding accuracy, denial rates, clinical specificity, time saved and whether clinicians are actually adopting the tool in daily practice.
The promise of clinical AI is real: faster insights, better-supported decisions, less administrative burden on overworked clinicians. But that promise only holds if the data feeding these systems retains its clinical meaning from the moment it's captured through every downstream use. Getting that foundation right isn't a technical afterthought. It's the difference between AI that earns clinicians' trust over time and AI that quietly erodes it, one confidently wrong answer at a time.
Tags
Original Sources
Clinical AI pilots falter when data loses meaning
↗ https://www.healthcareitnews.com/news/clinical-ai-pilots-falter-when-data-loses-meaning
About the author
Amara's entry point into AI was an epidemiology role at a London research hospital, where she spent five years studying how digital health tools reached — or conspicuously failed to reach — underserved communities. Watching early algorithmic systems in healthcare quietly entrench existing inequalities, she redirected her career toward the systemic consequences of AI at scale. She covers AI through an unflinching lens: who benefits, who bears the cost, and what evidence actually says versus what the press release claims. Her writing is calm and precise, but she doesn't mistake balance for neutrality.
More from The Steward →This Week's Edition
30 September 2026
28 articles
Related Articles

Blue Shield of California Signals Commercial Expansion of Medicare's ACCESS Chronic Care Model
Health & Science · 5 min

HHS Directs Nearly $250 Million Toward Addiction Treatment and Overdose Prevention
Health & Science · 5 min

UVM Wins $38 Million to Build AI Digital Twins for ICU Patients
Health & Science · 6 min
Related Articles

Blue Shield of California Signals Commercial Expansion of Medicare's ACCESS Chronic Care Model
Health & Science · 5 min

HHS Directs Nearly $250 Million Toward Addiction Treatment and Overdose Prevention
Health & Science · 5 min

UVM Wins $38 Million to Build AI Digital Twins for ICU Patients
Health & Science · 6 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.