
Share
Before AI regulation gets written from scratch, policymakers might look at IEEE 1012, a software safety standard that already sorts risk by consequence and likelihood, offering a tested framework instead of an untested one.
When a piece of software fails, the consequences range from a mildly annoying glitch to a catastrophe that costs lives. That range is not new. Engineers who build medical devices, aircraft systems, and nuclear plant controls have spent decades figuring out how to sort software risks by severity and act accordingly. As lawmakers now scramble to figure out how to regulate artificial intelligence, that existing body of work deserves a much closer look.
Think of it like triage in an emergency room. Not every patient who walks through the door needs the same level of urgent attention. A sprained ankle and a heart attack both require care, but the resources, speed, and scrutiny applied to each look nothing alike. Software safety engineers have long used a similar logic, sorting systems into what they call integrity levels: a shorthand for how much verification, testing, and oversight a piece of software needs based on what could go wrong if it fails.
The IEEE 1012 standard, a long-standing framework for software verification and validation, offers one of the clearest examples of this approach in practice. It maps integrity levels onto a combination of two factors: consequence, meaning how bad the outcome would be if the software failed, and likelihood, meaning how probable that failure actually is. A system where failure could kill people and where that failure is reasonably likely to occur gets the highest integrity level, demanding the most rigorous testing and independent review. A system where failure is both unlikely and low-stakes gets a much lighter touch.
This two-factor approach matters because consequence alone can be misleading. A nuclear reactor control system and a smoke detector could both, in theory, be linked to loss of life. But treating them identically would waste enormous resources on the smoke detector while potentially still under-resourcing the reactor, depending on how each system is actually deployed and how often it might fail in practice.
By pairing consequence with likelihood, the standard avoids that trap. It asks not just "how bad could this be" but "how often might this actually happen." A rare but catastrophic failure mode gets treated differently than a catastrophic failure mode that's baked into normal operating conditions. Similarly, a frequent but low-stakes glitch, like a typo in a weather app, gets treated far more lightly than a rare but frequent-enough-to-worry-about failure in, say, an autonomous vehicle's braking system.
This kind of graduated framework already exists across engineering disciplines that predate the current AI boom by decades. Aerospace engineers use similar tiering for flight control software. Medical device regulators use comparable logic when deciding how much clinical evidence a new diagnostic tool needs before reaching patients. Nuclear plant operators apply analogous thinking to safety-critical control systems. None of these industries invented risk tiering because it sounded good in a policy paper. They invented it because ungraded, one-size-fits-all oversight either strangles innovation with excessive paperwork for low-risk tools or, worse, leaves genuinely dangerous systems under-scrutinized.

AI systems present a similar spread of stakes. A chatbot that occasionally gives an awkward response to a customer service query sits at one end of the spectrum. An AI system used to triage patients in an emergency room, screen loan applications, or make sentencing recommendations in a courtroom sits at the other. Treating both with identical regulatory weight would be both inefficient and dangerous: overregulating the chatbot while potentially underregulating tools that shape people's health, finances, and freedom.
That's the core argument for looking at frameworks like IEEE 1012 as a starting point for AI oversight, rather than building an entirely new risk taxonomy from scratch. The consequence-likelihood matrix is not a perfect fit for every AI application. Machine learning systems, especially those built on large language models, behave differently than the deterministic control software these standards were originally designed for. Predicting the likelihood of failure in a system that generates novel outputs is far messier than predicting failure rates in a braking system with well-defined mechanical tolerances. Still, the underlying logic, scaling scrutiny to match real-world stakes, translates well.
Adapting these standards for AI would require honest engagement with what's genuinely different about machine learning systems. Traditional software failures tend to be reproducible: the same input produces the same broken output every time, which makes testing more tractable. AI systems, particularly generative ones, can behave inconsistently even with identical inputs, and their failure modes are often discovered only after deployment at scale, when millions of users interact with a system in ways developers never anticipated. Any integrity-level framework borrowed from older software standards would need built-in mechanisms for ongoing monitoring, not just pre-deployment testing, to account for that unpredictability.
Regulators drafting AI rules face a genuine choice: build tiered frameworks that already have decades of engineering practice behind them, or start over and risk repeating mistakes that other safety-critical industries solved long ago. Standards bodies like IEEE have already done years of unglamorous work figuring out how to match oversight intensity to actual risk, and there's little reason to discard that groundwork simply because the underlying technology is new.
That doesn't mean simply copying and pasting old standards onto AI systems and calling the job done. It means using proven risk-tiering logic as scaffolding, then building in the flexibility AI-specific unpredictability demands. Getting this balance right matters enormously, not for engineers debating standards documents in conference rooms, but for the people whose health, finances, and safety increasingly depend on systems whose failures are hard to predict and, at the highest integrity levels, impossible to afford.
Tags
Original Sources
Table 1: IEEE 1012 Standard’s Map of Integrity Levels Onto a Combination of Consequence and Likelihood Levels
↗ https://spectrum.ieee.org/regulating-ai-programs-roadmap/table-1-ieee-1012-standards-map-of-integrity-levels-onto-a-combination-of-consequence-and-likelihood-levels
About the author
Amara's entry point into AI was an epidemiology role at a London research hospital, where she spent five years studying how digital health tools reached — or conspicuously failed to reach — underserved communities. Watching early algorithmic systems in healthcare quietly entrench existing inequalities, she redirected her career toward the systemic consequences of AI at scale. She covers AI through an unflinching lens: who benefits, who bears the cost, and what evidence actually says versus what the press release claims. Her writing is calm and precise, but she doesn't mistake balance for neutrality.
More from The Steward →This Week's Edition
6 September 2026
41 articles
Related Articles
Related Articles
More Stories
© 2026 Cedar & Bloom. All rights reserved.