
Share
A16z-backed Vals AI wants to become the trusted referee for AI model performance, with revenue up eightfold and federal contracts in tow. The bigger question is whether verification businesses scale like the models they're grading.
Benchmarking has quietly become one of the AI industry's most valuable forms of marketing. A strong score on the right test can move product decisions, sway enterprise buyers, and generate headlines. The problem, increasingly well documented, is that many of these benchmarks are outdated and gameable. Companies can train against publicly available test sets, effectively memorizing the answers before the exam.
Vals AI, a startup formed in 2024, is positioning itself as the fix. The company has moved quickly from a seed round led by 8VC and Bloomberg Beta last year to a $40 million Series A led by Andreessen Horowitz last month. Revenue is now eight times what it was a year ago, according to the company. Headcount has tripled, from eight employees at the start of the year to 25 today, with plans to add another 10 to 15 hires.
That growth trajectory matters more than the novelty of the product. Independent verification is a business model that AI labs, enterprises, and eventually regulators may need as model deployment scales. Vals is trying to get there first.
Co-founder Rayan Krishnan, 25, built Vals on a straightforward observation: academic benchmarks weren't keeping pace with frontier model advances. His background, a Stanford undergraduate stint, internships at Palantir, and work with Microsoft and Stanford's AI lab, gave him early exposure to the gap between lab-reported capabilities and real-world performance.
Vals differentiates on two fronts. First, it keeps its test materials private, closing the loophole that lets a model train against a public exam and post an inflated score. Second, it grades models on domain-specific task completion rather than abstract knowledge tests. Krishnan draws a contrast with legacy approaches: "Historically, I think evaluation has been done to evaluate intelligence in a very abstract way. Like, do models know enough information to be able to take a bar exam type test?" Instead, Vals asks whether a model "can do work that produces a product of the same quality as a human within every domain."
The company has also built out testing for downside risk, not just capability. Krishnan describes evaluating what happens "if these models ran wild in the world." That includes newer benchmarks covering recursive self-improvement, mental health, cybersecurity, biosecurity, and even a model's ability to apply the Geneva Convention in law-of-armed-conflict scenarios. It's an unusual product roadmap for a 25-person startup, but it signals where Vals thinks demand is heading: safety and compliance verification, not just leaderboard bragging rights.
The revenue model itself is worth pausing on. AI labs pay Vals to test their own models, a dynamic Krishnan compares to a student paying the College Board to sit the SAT. It's a counterintuitive purchase, paying to potentially learn your product underperforms, but it functions as a diagnostic tool for labs trying to improve iteratively, and increasingly as a trust signal for buyers deciding which model to deploy. Vals has also launched a program offering model evaluations to federal agencies, extending its customer base beyond commercial AI labs into government procurement, a market with its own long sales cycles but sticky, high-value contracts once secured.

The timing lines up with a broader shift in how AI companies are funded and valued. SpaceX has gone public. Anthropic is reportedly slated for a public listing later this year, and OpenAI may not be far behind. Krishnan's own framing is direct: as AI becomes "a core part of the economy," he expects benchmarks like his to become "a central part of how these companies submit public filings or talk about the prospective investments they're going to make in AI."
That is a meaningful claim. If independent, private-test benchmarking becomes a reference point for public disclosures or capital allocation decisions, the company that owns that standard captures a structural position in the AI economy, akin to a ratings agency or an accreditation body. Andreessen Horowitz's $40 million bet reads as a wager on exactly that outcome.
The opportunity comes with real risks. Vals is still a 25-person company competing to become the trusted arbiter for an industry with vastly more capital and talent. Model developers themselves, or larger incumbents, could build competing internal or open evaluation frameworks that undercut Vals's independence pitch. Regulatory bodies could also mandate their own testing regimes, sidelining private benchmarking firms altogether.
There's also a conflict-of-interest question baked into the business model. Vals is paid by the same companies whose models it evaluates. That structure isn't unusual, credit rating agencies operate similarly, but it invites scrutiny over incentive alignment as the stakes rise, particularly if benchmark results start feeding into public filings or investment decisions, as Krishnan predicts.
Revenue growth of 8x year-over-year is an impressive headline figure, but it's growth off a small base, and the company hasn't disclosed absolute revenue figures, margins, or customer concentration. A 25-person team scaling into federal contracting and expanded biosecurity and cybersecurity testing domains will face execution risk that a funding round alone doesn't resolve.
Vals AI has identified a real gap: benchmarking hasn't kept up with the pace of frontier model development, and buyers need trustworthy, hard-to-game evaluation. The $40 million Series A, the 8x revenue growth, and the federal agency program are credible early signals. But becoming the "gold standard" for AI benchmarking is an ambitious label to live up to, and the company's long-term value depends on whether it can convert early traction into durable, structural relevance as AI labs go public and capital allocation decisions increasingly hinge on verified performance rather than marketing claims.
Tags
Original Sources
Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking | TechCrunch
↗ https://techcrunch.com/2026/09/19/vals-backed-by-andreessen-horowitz-is-looking-to-become-the-gold-standard-for-ai-benchmarking
About the author
Marcus began tracking AI's market implications in 2016, noticing AI-related patent filings accelerating ahead of earnings upgrades before most of the sell-side had caught on. A former fixed-income quantitative analyst, he spent two decades building models that priced risk across emerging markets before pivoting to cover the economic impact of AI full-time. His writing translates opaque technical developments into clear risk/reward terms — and he's rarely diplomatic about the gap between AI valuations and underlying fundamentals. He believes most market participants still underestimate AI's long-run deflationary effect on knowledge work.
More from The Analyst →This Week's Edition
20 September 2026
20 articles
Related Articles
Related Articles
More Stories
© 2026 Cedar & Bloom. All rights reserved.