
Share
Anastasios Angelopoulos, CEO of Arena and a Berkeley-trained statistician, is bringing his work on rigorous, real-world AI benchmarking to Stanford HAI, raising hard questions about how we actually measure model quality.
Benchmarks are having a credibility crisis, and Stanford HAI is bringing in one of the people best positioned to talk about it. Anastasios Angelopoulos, co-founder and CEO of Arena, the widely used open platform for evaluating AI models through real-world human feedback, is headlining a seminar co-hosted with the AI Measurement Science Center (AIMS). The session is part of a recurring series on academic work in AI evaluation and measurement science, and it's worth paying attention to if you build, ship, or rely on model evaluations for a living.
If you've spent any time comparing LLMs, you've probably used Arena or seen its leaderboards cited in a model release blog post. The platform works by collecting human preference votes on model outputs in the wild, not just on curated academic test sets. That distinction matters more than it sounds. Static benchmarks get memorized, gamed, or simply fail to capture how a model performs on the messy, open-ended prompts real users throw at it. Arena's pitch is that aggregating live human judgment gets you closer to ground truth on actual usefulness.
Angelopoulos isn't a product guy who backed into research. He's the reverse: a theoretical statistician who built a product. He earned his Ph.D. in Computer Science from UC Berkeley, advised by Michael I. Jordan and Jitendra Malik, two names anyone in ML theory or computer vision will recognize instantly. He followed that with a postdoc under Ion Stoica, the Berkeley systems researcher behind Apache Spark and Databricks, before co-founding Arena. That combination, rigorous statistical training plus systems-scale infrastructure experience, shows up directly in how Arena approaches measurement.
Here's the core problem Angelopoulos has spent his career on: how do you get reliable, statistically defensible guarantees out of a black-box model? It's a deceptively hard question. Modern foundation models aren't simple classifiers where you can compute a clean confusion matrix and call it a day. They generate open-ended outputs, they're evaluated by other models or by humans with their own biases, and they're deployed in contexts wildly different from whatever data they were tested on.
A few things make this harder than it looks:

This is where the "measurement science" framing that AIMS uses becomes genuinely useful, rather than just an academic buzzword. It's not enough to report a leaderboard number. You need error bars, you need to understand what population of prompts and users generated that number, and you need methods that hold up under adversarial pressure once people start optimizing for the metric itself. Angelopoulos's research background is explicitly focused on giving black-box AI systems the kind of rigorous reliability guarantees that statisticians would demand from any other measurement instrument, whether that's a clinical trial or a sensor calibration.
That framing also explains why Arena has become influential well beyond academia. Model labs cite it, journalists reference it, and practitioners use it as a sanity check against vendor marketing claims. But influence cuts both ways: the more a benchmark shapes decisions and funding, the more scrutiny its methodology deserves. A platform built on aggregating human votes at scale has to answer real questions about sampling bias, demographic skew in who's voting, and whether popularity among casual users tracks with capability on harder, more specialized tasks. Angelopoulos's statistical background suggests he's thinking about exactly these failure modes, not ignoring them.
If you're an engineer picking between models for a production system, this stuff isn't academic. Leaderboard rank differences that look meaningful can evaporate once you account for statistical uncertainty. A model that "wins" by a point or two on an aggregate human-preference score might be statistically indistinguishable from its competitor once you account for variance in the voting population. Treating a single leaderboard number as gospel is a good way to make a bad engineering decision.
The seminar series itself is structured to dig into exactly this kind of nuance. It's hosted as part of a broader academic program pairing case studies of AI adoption with talks on measurement methodology, and sessions run through the academic year with different speakers each week. The Angelopoulos talk specifically sits at the intersection of two things that don't always get discussed together: the practical, product-facing world of leaderboards that everyone actually uses, and the theoretical statistics that determine whether those leaderboards mean anything.
Arena's human-feedback approach to evaluation has become a de facto standard for comparing frontier models, but its value depends entirely on the statistical rigor behind it. Angelopoulos brings a rare combination: deep theoretical grounding in inference and reliability for black-box systems, paired with hands-on experience running a real-time evaluation platform at scale. For practitioners, the takeaway is straightforward. Don't just read the leaderboard rank, ask about the confidence intervals, the sampling methodology, and whether the benchmark still measures what you think it measures once everyone starts optimizing for it. As AI evaluation matures into its own discipline, talks like this one are a useful reminder that measurement science, not just model architecture, is where a lot of the next hard problems in AI actually live.
Tags
Original Sources
Anastasios Angelopoulos | Measuring AI in the Real World | Stanford HAI
↗ https://hai.stanford.edu/events/anastasios-angelopoulos-measuring-ai-in-the-real-world
About the author
Kai built ML infrastructure at a Bay Area startup before developing an obsession with transformer architectures and inference optimisation that eventually pulled him out of product work entirely. A stint at a compute research lab sharpened his instinct for what actually matters in a model release versus what is marketing. He writes from the inside — from the perspective of someone who has debugged the systems he is describing at three in the morning. He is allergic to hype and instinctively drawn to the unglamorous plumbing questions that everyone else skips over.
More from The Engineer →This Week's Edition
18 September 2026
25 articles
Related Articles

Stanford's CORES Symposium Puts AI's Role in Scientific Reproducibility Under the Microscope
Models & Research · 5 min

Medicare's AI Push Meets a Discovery Problem for Health Chatbots
Products & Applications · 5 min

Microsoft's Internal AI Rollout Offers a Data Point on Enterprise Transformation Economics
Finance & Markets · 5 min
Related Articles

Stanford's CORES Symposium Puts AI's Role in Scientific Reproducibility Under the Microscope
Models & Research · 5 min

Medicare's AI Push Meets a Discovery Problem for Health Chatbots
Products & Applications · 5 min

Microsoft's Internal AI Rollout Offers a Data Point on Enterprise Transformation Economics
Finance & Markets · 5 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.