
Share
A successful AI demo isn't the finish line, it's the first day on the job. Healthcare leaders are learning that piloting an AI agent and trusting it with real patients are two very different tests.
A CEO recently stopped a product lead mid-sentence during a rollout briefing for a new AI Admissions Coordinator. The plan sounded simple enough: pilot, then probation, then full deployment. "What do you mean by probation?" the CEO asked.
That question matters more than it might seem. For anyone who has spent the past year deploying AI systems inside mental health clinics, the difference between a pilot and a probation period feels obvious. It isn't, not yet, and that gap is a real problem. Healthcare organizations know how to test software. They do not yet have a shared playbook for onboarding an AI agent that is expected to do a job from start to finish, the way a new employee would.
This isn't a small operational quirk. It helps explain why so many promising AI projects in healthcare quietly stall. A survey of more than 400 U.S. healthcare leaders found that only 30% of completed AI proofs of concept ever made it into production, according to Bessemer Venture Partners' Healthcare AI Adoption Index. Leaders cited security concerns, data readiness, integration costs, and limited in-house expertise as barriers. A probation period will not solve every one of those problems. But it does address something specific: the moment a team watches a demo succeed and assumes the hard work is finished. In reality, that is often when the real work, monitoring live cases and correcting mistakes as they happen, has only just started.
Think about what integration actually buys you. A recent MedCity article made the case that workflow integration is where many AI deployments fail. That is true as far as it goes. But for an AI agent handling a real job, integration only gets the system to the starting line. Giving it access to the right tools and data does not mean it can use them wisely without someone checking its work.
Consider that AI Admissions Coordinator again. Its job is not just picking up the phone and asking a few questions. It has to know when it has gathered enough information to move forward. It has to route a patient to the right provider, book the correct appointment type, and recognize when a case falls outside its rules. Most calls follow a predictable script, and those are easy. The trouble starts when insurance information is incomplete, or a request does not fit the clinic's scheduling logic, or the next step simply isn't clear. That is exactly where a pilot ends and probation needs to begin.
Probation, in this sense, is the stretch when the clinic actively helps the agent succeed at real, messy work. Smart teams start small: one location, a narrow set of call types, a limited slice of appointment categories. In the early days, someone might review every single call the agent handles. That is not evidence the pilot failed. It's simply how onboarding works, for a person or a piece of software. Whatever the team learns during that review should feed back into the agent's instructions, its integrations, and its escalation rules. As the agent's performance becomes consistent, the clinic can widen the scope of what it handles while pulling back on constant review. Eventually, routine calls should not need a human checking every step. The team's job shifts from watching individual calls to monitoring overall performance metrics, while the agent itself flags anything unusual.

Who signs off on all this matters too, and it is often the wrong person by default. The executive who approves a pilot is not necessarily the person who should decide when probation ends. A CEO or COO is well positioned to weigh the business case, the strategic fit, and how much risk the organization can tolerate. But the workflow owner, someone who actually sees the incomplete cases, the downstream cleanup, and the avoidable escalations, is far better positioned to judge whether the work itself is good. For an Admissions Coordinator, that might be the VP or Head of Admissions. Senior leadership decides whether the organization wants this agent in the first place. The workflow owner decides whether the team can trust it with the job every single day.
Even the scorecard needs rethinking. Traditional software metrics, logins, daily active users, feature adoption, uptime, tell you whether a tool is being used. They tell you almost nothing about whether an AI agent is actually finishing the work it was hired to do. During probation, the numbers that matter are different: the share of eligible cases the agent completes end to end, the accuracy and completeness of that work, how much rework it generates downstream, whether it escalates the right exceptions, and how much human time each completed case still requires.
Those figures reveal how much real, dependable capacity an agent is adding to a team. Sometimes that looks like the output of one full-time employee. Sometimes it looks like five. But that comparison only holds up if the quality of the work matches. Calling an agent "five FTEs of capacity" while staff quietly clean up its mistakes behind the scenes isn't measurement. It's accounting fiction, and it hides real costs.
Graduating from probation should not mean oversight disappears. The Joint Commission's Responsible Use of AI certification requires ongoing monitoring of AI performance and safety across its entire lifecycle, not just at launch. Clinics adopting AI agents should hold themselves to the same standard. When an agent takes on a new responsibility, treat it the way you would treat a human employee's promotion: put that new scope through its own probation period. The same goes for any material change to the underlying model, its instructions, its integrations, or the clinic's own rules. Strong past performance proves the old setup worked. It says nothing about whether the new one will.
None of this means reviewing every case forever. Mature, stable workflows can and should stay exception-based, with humans stepping in only when something unusual happens. But new responsibilities, or any meaningful change to how the agent operates, should bring back closer supervision for a while, with the workflow owner deciding when the agent has earned its autonomy back.
The hiring analogy holds up well here, and it is worth keeping in mind as more clinics adopt these systems. A pilot is the resume, the interview, the work sample. It tells you whether an agent has the necessary skills and can perform under controlled conditions. Probation is its first real stretch on the job, when you find out whether it can handle normal variation, finish work reliably, and know when to ask for help. The pilot offers evidence of capability. Probation is what tells you whether the agent, and the clinic around it, is actually ready for responsibility.
Tags
Original Sources
AI Workers Need a Probation Period, Not Just a Pilot - MedCity News
↗ https://medcitynews.com/2026/09/ai-workers-need-a-probation-period-not-just-a-pilot
About the author
Amara's entry point into AI was an epidemiology role at a London research hospital, where she spent five years studying how digital health tools reached — or conspicuously failed to reach — underserved communities. Watching early algorithmic systems in healthcare quietly entrench existing inequalities, she redirected her career toward the systemic consequences of AI at scale. She covers AI through an unflinching lens: who benefits, who bears the cost, and what evidence actually says versus what the press release claims. Her writing is calm and precise, but she doesn't mistake balance for neutrality.
More from The Steward →This Week's Edition
23 September 2026
29 articles
Related Articles

Gen Z Wants to Stay in Healthcare, But Not Without a Reason to Stick Around
Job Market & Society · 5 min

Healthcare's Leadership Bench Turns Over: Hackensack Meridian, Keck Medicine, ChristianaCare Name New Chiefs
Job Market & Society · 6 min

Engineers Are Earning Less in Real Terms, Even as Demand for Their Skills Grows
Job Market & Society · 5 min
Related Articles

Gen Z Wants to Stay in Healthcare, But Not Without a Reason to Stick Around
Job Market & Society · 5 min

Healthcare's Leadership Bench Turns Over: Hackensack Meridian, Keck Medicine, ChristianaCare Name New Chiefs
Job Market & Society · 6 min

Engineers Are Earning Less in Real Terms, Even as Demand for Their Skills Grows
Job Market & Society · 5 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.