
Share
Nearly 750,000 people agreed to share their health records with the NIH's All of Us program. More than 300,000 of those records still can't be used. The reason has nothing to do with consent.
Imagine agreeing to donate your medical history to science, signing every form, trusting the institution with your most personal data, and then having your records simply vanish into a database that can't read them. That's effectively what happened to more than 300,000 participants in the National Institutes of Health's All of Us Research Program.
STAT reported in June that despite persuading 98% of its nearly 750,000 participants to share their electronic health records (EHRs), the program's database has no usable EHR data at all for over 300,000 of them. The program's most recent data release tries to close that gap by tapping into the clinical data-sharing networks hospitals already use to move records between health systems. It's a sensible fix. But the reason it's necessary tells us something important about where American health data infrastructure actually stands in 2026, and it isn't where most people assume.
For the better part of a decade, the dominant theory of health data interoperability was essentially a plumbing problem. Connect the pipes, the thinking went, and the data would flow on its own. That theory wasn't wrong. The 21st Century Cures Act established a patient's legal right to their own data. Certified EHR vendors were required to expose standardized data-sharing interfaces called FHIR APIs. National networks now let hospitals query each other's records almost as easily as placing a phone call. People inside federal health IT offices, hospital systems, and EHR vendors did genuinely hard work to make this happen, often against real institutional resistance. The 98% consent rate in All of Us is proof patients are willing. The pipes, for the most part, are connected.
So why do 300,000 records still not show up?
Retrieving a record and having usable data from that record are two separate problems. The health IT industry solved the first one before it finished the second.
A hospital today can transmit a patient's full chart to a research database in seconds. What actually arrives, though, is rarely a clean, structured dataset ready for analysis. It's a patchwork: some standardized data fields, a scanned referral letter, a faxed prior-authorization form, a PDF export of a visit note with layouts that vary from hospital to hospital and decade to decade. A single patient's chart might span several different systems and document formats, including years of care delivered before anyone thought to make records machine-readable in the first place. The pipe works fine. What flows through it often doesn't arrive in a form any research database can actually use.
This is the part of interoperability that rarely comes up in policy debates, because it isn't really a policy problem. It's a data-engineering problem, and it happens to be a lot closer to being solved than most health system leaders or researchers tend to assume.
Over the past two years, the tools for converting messy, inconsistent, real-world medical documents into structured, validated, research-ready data have matured substantially. The breakthrough isn't one single AI model. It's the pairing of two different approaches working together. The first is probabilistic extraction: AI systems that can read a scanned, handwritten, or inconsistently formatted document roughly the way a trained person would, making educated guesses about what a field means. The second is deterministic validation: rules-based systems that check every extracted field against defined standards before accepting it into a dataset.

Think of it like a junior analyst reading a messy file and flagging their best guesses, paired with a senior reviewer who checks every single guess against a strict rulebook before signing off. Probabilistic AI alone isn't reliable enough for the accuracy thresholds research and clinical use demand. Rules-based systems alone can't handle the sheer variety of formats that real-world medical records take. Together, the two approaches can do what neither can do alone. The gap between them, once measured in years of custom engineering per health system, is now closing in a matter of weeks.
That shift matters for three reasons, and they apply well beyond All of Us. Every research program, health system, and payer facing its own version of the 300,000-record problem has a stake in this.
It's faster than most institutions have priced in. Turning a backlog of unstructured records into usable data no longer requires a multi-year integration project or tearing out legacy systems. It can run alongside what a hospital already has in place, ingesting whatever format a record arrives in, whether that's a fax, a scan, a PDF, or a structured export, without forcing anyone to change how or where their data is stored.
It's also more accurate than the manual alternative, not less. The instinct in research and regulatory circles is to treat automation as a tradeoff against accuracy, with manual chart review as the cautious default. But deterministic validation layers can flag every field that falls outside expected bounds for human review, producing a documented audit trail showing exactly what was extracted, from which source, and at what confidence level. That's often a more rigorous evidentiary record than manual abstraction produces today.
And none of this requires loosening the security standards a program like All of Us has to meet. Records don't need to leave controlled environments or get exposed to consumer AI tools. The compliance frameworks that regulated research already operates under, including SOC 2 and HIPAA, stay intact. If anything, the validation layer is what gives auditors and institutional review boards the traceability they need to trust automated extraction in the first place.
None of this is a knock on how All of Us is handling the problem. Piggybacking on existing data-sharing networks to pull in more records is the right call, and it will close part of the gap. But it will eventually hit the same wall every large-scale EHR retrieval effort runs into. Some meaningful share of what comes back simply won't be clean, structured data. It'll be the messy remainder: the scanned document, the legacy export, the record from a system that never fully adopted modern data standards. Retrieval alone can't fix that.
The real lesson from the 300,000-record gap isn't that interoperability failed. It's that interoperability, as the field has largely defined it, meaning can a record move from one system to another, has substantially succeeded. What's left is a different, more solvable problem than the one the field spent the last decade on. Turning what arrives into what researchers can actually use is no longer the open-ended, multi-year undertaking it was treated as five years ago. For the people whose data is sitting in that gap right now, that difference isn't abstract. It's the difference between their health history contributing to medical research or sitting unread in a database that can't make sense of it. Research programs willing to treat this as the solvable engineering problem it has become will close gaps like this one far faster than the last decade of policy work managed to.
Tags
Original Sources
The Real Gap in All of Us Isn't Consent, It's Conversion - MedCity News
↗ https://medcitynews.com/2026/10/the-real-gap-in-all-of-us-isnt-consent-its-conversion
About the author
Amara's entry point into AI was an epidemiology role at a London research hospital, where she spent five years studying how digital health tools reached — or conspicuously failed to reach — underserved communities. Watching early algorithmic systems in healthcare quietly entrench existing inequalities, she redirected her career toward the systemic consequences of AI at scale. She covers AI through an unflinching lens: who benefits, who bears the cost, and what evidence actually says versus what the press release claims. Her writing is calm and precise, but she doesn't mistake balance for neutrality.
More from The Steward →This Week's Edition
2 October 2026
28 articles
Related Articles

When AI Helps Students Practice But Hurts Them on Test Day
Job Market & Society · 5 min

County Officials Get a Crash Course in AI, But Questions Remain About Who Gets Left Behind
Job Market & Society · 5 min

Stanford HAI Names 15 PhD Researchers to New Data Science Scholars Cohort
Job Market & Society · 5 min
Related Articles

When AI Helps Students Practice But Hurts Them on Test Day
Job Market & Society · 5 min

County Officials Get a Crash Course in AI, But Questions Remain About Who Gets Left Behind
Job Market & Society · 5 min

Stanford HAI Names 15 PhD Researchers to New Data Science Scholars Cohort
Job Market & Society · 5 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.