
Share
New survey data shows most enterprises can't trace AI infrastructure failures to their root cause, and that blind spot grows more dangerous as companies hand remediation over to autonomous agents.
Picture a hospital where three different alarms go off at once, but none of the monitors can tell doctors which patient is actually in crisis. That's roughly the situation enterprise IT teams face when an AI training job slows down. The GPU dashboard flags contention. The data pipeline reports a stall. The storage system throws I/O alerts. Each signal is technically true. None of them tells you what actually happened, or where to start fixing it.
That gap between seeing a problem and understanding it is costing companies real money and, increasingly, real safety margin. While infrastructure teams chase symptoms across disconnected dashboards, expensive GPU capacity sits idle. And as enterprises build what the industry calls AI factories, sprawling systems spanning GPUs, storage, networks, data pipelines and hybrid cloud environments, the problem compounds. Most monitoring tools were never built to see across all of it at once.
Virtana's AI Factory Reality Check surveys of U.S. and U.K. enterprise decision makers put a number on that blind spot. Fifty-nine percent of U.S. enterprises and 53% of U.K. enterprises cannot automatically identify the root cause of an AI workload failure across infrastructure domains. That's not a minor inconvenience. It's the difference between fixing a problem and guessing at one.
"Once you get to causality, you can remediate," says Paul Appleby, president and CEO of Virtana. "If you can't get to causality, you're really throwing a dart at a wall to work out what happened and how to fix it."
Part of the problem is cultural. Board pressure pushes companies to show AI progress through deployed infrastructure, so many build first and instrument later, if they instrument at all. The legacy tools doing that instrumentation were designed for stable, predictable systems. AI factories are the opposite. Schedulers shift workloads across hybrid environments constantly, and resource use spikes and drops in bursts nobody planned for.
Threshold-based alerting, the kind that flags a problem when a metric crosses some preset line, only works if you know what "normal" looks like. Sixty-six percent of U.S. enterprises run AI infrastructure without reliable performance baselines. Just 34% of U.S. and 26% of U.K. enterprises describe AI workload performance as highly predictable. At U.S. organizations with more than 50,000 employees, that figure drops to 25%. The companies running the biggest, most consequential AI systems are also the ones with the least ability to anticipate what those systems will do.
Enterprises trying to control premium hardware costs are now rebalancing workloads and consolidating systems while those systems run under live load. Every shift changes how components depend on each other and compete for resources.
"Without system-level observability, organizations can't determine how these changes affect outcomes, cost or reliability," Appleby says. "They're continuously optimizing AI systems they don't fully understand, and introducing risk with every change."
The diagnostic gap shows up differently in each country, but it shows up everywhere. In the U.K., 75% of enterprises rely on automated alerting as their first response to a failure, yet only 47% can automatically trace the root cause across all infrastructure domains. The rest see only one piece of the puzzle, correlate signals by hand, or pull multiple teams together for hours or days to untangle what happened. In the U.S., a quarter of enterprises start incident response with manual investigation across disconnected consoles, essentially detective work performed under time pressure with incomplete evidence.

Appleby describes the cascading failure pattern that makes this so hard to untangle: a storage bottleneck degrades a data pipeline, which stalls a training job, which produces GPU contention. Legacy monitoring registers three separate events in three separate domains. Each is accurate on its own terms. None identifies the actual cause.
Fixing this requires something closer to a living map of the entire system, one that correlates telemetry from every domain in real time and updates continuously as workloads move. Leaders surveyed in both countries ranked unified visibility across AI and infrastructure as their top priority, with automated root cause analysis a close second. That preference held across every role and revenue band surveyed, which suggests this isn't a niche concern. It's a shared recognition that bolting more monitoring tools onto an already fragmented system won't solve a problem rooted in fragmentation itself.
The stakes rise further once autonomous agents enter the picture. Large AI factories will eventually need autonomous remediation, Appleby argues, simply because the complexity involved already exceeds what human teams can coordinate manually. Only 23% of U.K. infrastructure and site reliability engineers who manage AI workloads daily describe performance as highly predictable. Handing control to an agent before fixing the underlying visibility problem doesn't eliminate risk. It accelerates it.
"Without that foundation, AI agents inherit the same blind spots that constrain human operators, and they amplify those failures at machine speed," Appleby says. "An agent that acts on incomplete system context doesn't resolve incidents faster, it creates new ones."
There's also a trust gap inside organizations that deserves attention. In the U.K., 59% of executives believe their organization can automatically identify root cause across all infrastructure domains. Among the infrastructure and SRE engineers who actually field the alerts, only 34% agree. That's a governance problem as much as a technical one: the people approving AI budgets are working from a rosier picture than the people doing the work.
Regulatory exposure raises the cost of getting this wrong. U.K. enterprises operate under GDPR and sector-specific oversight in financial services, healthcare and critical infrastructure, frameworks that demand audit trails and cost attribution. Yet 39% of U.K. enterprises are deprioritizing security and compliance reviews as AI factory demands grow, a trade-off that may look reasonable now and costly later.
"The system that proves an AI factory is performing is the same system that satisfies a regulator, an auditor or a board inquiry," Appleby says. "Sovereignty without observability is a claim that can't be evidenced."
The public narrative suggests Global 2000 companies are deploying AI at massive, confident scale. Appleby pushes back on that. Most of the largest AI infrastructure investments still come from hyperscalers and cloud platforms, not ordinary enterprises. "It's earlier than the level of investment would indicate," he says. "There aren't many examples anywhere in the world of large enterprises that have built, scaled and deployed industrial-scale AI services and are operating them efficiently."
That matters because the instinct to add more compute before fixing visibility and control doesn't solve the underlying problem. It just builds a bigger, costlier version of it. Companies that invest in understanding their systems now, before layering autonomous agents on top, give themselves a real chance to scale safely. Those that don't are gambling with infrastructure, budgets and, increasingly, regulatory standing all at once.
Tags
Original Sources
AI agents can’t fix infrastructure failures they can’t diagnose
↗ https://venturebeat.com/infrastructure/ai-agents-cant-fix-infrastructure-failures-they-cant-diagnose
About the author
Amara's entry point into AI was an epidemiology role at a London research hospital, where she spent five years studying how digital health tools reached — or conspicuously failed to reach — underserved communities. Watching early algorithmic systems in healthcare quietly entrench existing inequalities, she redirected her career toward the systemic consequences of AI at scale. She covers AI through an unflinching lens: who benefits, who bears the cost, and what evidence actually says versus what the press release claims. Her writing is calm and precise, but she doesn't mistake balance for neutrality.
More from The Steward →This Week's Edition
7 October 2026
34 articles
Related Articles

Cisco's Edge Intelligence Tackles the Unsexy Problem of Getting IoT Data Out of the Field
Tools & Engineering · 5 min

The 2026 Nobel Prizes, and the Persistent Gap Science Still Hasn't Closed
Policy & Regulation · 5 min

Trump Declares Anyone Who Says "AI" Instead Of "Super Intelligence" An Enemy
Policy & Regulation · 5 min
Related Articles

Cisco's Edge Intelligence Tackles the Unsexy Problem of Getting IoT Data Out of the Field
Tools & Engineering · 5 min

The 2026 Nobel Prizes, and the Persistent Gap Science Still Hasn't Closed
Policy & Regulation · 5 min

Trump Declares Anyone Who Says "AI" Instead Of "Super Intelligence" An Enemy
Policy & Regulation · 5 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.