
Share
As language models grow more capable, the benchmarks meant to catch their blind spots are falling behind, leaving safety researchers scrambling to measure risks before they reach the public.
Think about the last time you asked a chatbot a tricky question and got an answer that sounded confident but turned out wrong. Now imagine that same overconfidence baked into a system making decisions about your healthcare, your loan application, or your job screening. That is the quiet worry sitting underneath the AI boom right now: the tools are advancing faster than our ability to verify they are safe.
IEEE Spectrum's reporting on the state of AI makes a point that deserves more attention than it gets. Some of the field's most-used benchmarks, the standardized tests researchers rely on to measure how well an AI system performs, are becoming too easy. When a test stops being challenging, it stops telling us anything useful. That is not a minor technical footnote. It is a warning sign about how we evaluate systems before they get deployed into homes, hospitals, and workplaces.
A commenter on the original piece even joked that language models are getting so good that the real test would be feeding one the Bible, a text famously full of contradictions and puzzles, just to see if the system could make sense of it. The joke lands because it points at something real. If our current yardsticks are maxed out, we need harder, more creative ways to find the cracks in these systems before those cracks show up in the real world.
Here is the plain-language version of the problem. A benchmark is like a driving test. If everyone who takes it passes with a perfect score, you have not proven that all your drivers are excellent. You have proven the test is too easy. The same logic applies to AI. When top models start acing tasks that used to separate strong systems from weak ones, researchers lose the ability to tell which systems are actually reliable and which just look good on paper.
This gap matters most in high-stakes settings. A model that performs beautifully on a benchmark but stumbles on edge cases, unusual inputs, or adversarial prompts crafted to trick it, can still cause real harm once it is out in the wild. Safety testing needs to evolve alongside capability, not lag behind it. Right now, the evidence suggests it is lagging.
Risk management researchers have long argued that you cannot manage what you cannot measure. That principle applies directly here. If the tools we use to measure AI risk are saturated, meaning models are hitting ceiling scores across the board, then organizations deploying these systems are essentially flying without accurate instruments. They may believe a system is safe simply because it passed a test that no longer measures anything meaningful.

This is not an argument against AI progress. It is an argument for keeping pace on the safety side. Every major advance in a powerful technology, from pharmaceuticals to aviation, has required parallel investment in testing infrastructure. AI is no different. The question is whether the field is willing to fund and prioritize that infrastructure with the same enthusiasm it brings to building bigger, flashier models.
There is a governance angle here too. Policymakers around the world are drafting AI regulations right now, and many of those frameworks lean on benchmark performance as a proxy for safety. If the benchmarks themselves are unreliable, then regulations built on top of them inherit that unreliability. A rule that says "systems must score above X on safety benchmark Y" only works if benchmark Y actually distinguishes safe systems from unsafe ones. Saturated tests undermine the whole structure.
None of this is meant to induce panic. It is meant to prompt a fairly practical question: who is responsible for building the next generation of harder, more realistic tests? Right now, much of that work falls to academic labs and a handful of nonprofit research groups, often without the resources that commercial AI labs have. That imbalance, well-funded model development on one side and comparatively underfunded safety evaluation on the other, is itself a risk worth naming.
There is a reader in the original comments who thanked the author simply for "caring about the article information quality." It is a small comment, but it reflects something bigger. People want to trust that the information and tools shaping their lives have been checked carefully. That trust is earned through rigorous testing, not assumed because a company says its product is safe.
The stakes here are not abstract. AI systems already influence decisions about credit, hiring, medical triage, and content moderation. When a company claims its model is safe because it passed a benchmark, the public generally has no independent way to verify that claim. We rely on the benchmark doing its job. If the benchmark is saturated and no longer discriminates between strong and weak performance, that trust is misplaced, even if no one intended to mislead anyone.
Fixing this does not require slowing down AI development, though some caution on rollout speed would not hurt. It requires treating safety evaluation as a first-class engineering problem, deserving the same funding, talent, and urgency as capability research. Harder benchmarks, more adversarial testing, and independent auditing bodies are not obstacles to innovation. They are the guardrails that let innovation proceed without leaving the public to absorb the risk of untested systems. The technology will keep getting more capable. The real test is whether our ability to check it keeps up.
Tags
Original Sources
11. Risks? There are Risks?
↗ https://spectrum.ieee.org/the-state-of-ai-in-15-graphs/ai-risks
About the author
Amara's entry point into AI was an epidemiology role at a London research hospital, where she spent five years studying how digital health tools reached — or conspicuously failed to reach — underserved communities. Watching early algorithmic systems in healthcare quietly entrench existing inequalities, she redirected her career toward the systemic consequences of AI at scale. She covers AI through an unflinching lens: who benefits, who bears the cost, and what evidence actually says versus what the press release claims. Her writing is calm and precise, but she doesn't mistake balance for neutrality.
More from The Steward →This Week's Edition
6 September 2026
41 articles
Related Articles

When the Cloud Fails: Millions of Patient Records Exposed in a String of Healthcare Breaches
Security & Risk · 6 min

A New Startup Sells Access to AI Models With Their Safety Guardrails Stripped Out
Security & Risk · 6 min

A Startup Is Selling Access to AI Models With the Safety Rails Removed
Security & Risk · 6 min
Related Articles

When the Cloud Fails: Millions of Patient Records Exposed in a String of Healthcare Breaches
Security & Risk · 6 min

A New Startup Sells Access to AI Models With Their Safety Guardrails Stripped Out
Security & Risk · 6 min

A Startup Is Selling Access to AI Models With the Safety Rails Removed
Security & Risk · 6 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.