
Share
On September 24, 2026, Stanford HAI brings together researchers building AI agents that replicate studies, catch errors, and automate research pipelines, alongside a hard question: can we actually trust these systems yet?
Reproducibility has been science's quiet crisis for over a decade. Studies fail to replicate, methods sections leave out crucial details, and researchers burn months trying to reconstruct someone else's pipeline from a sparse GitHub repo. Now AI is being pitched as both part of the problem and part of the fix, and Stanford's Center for Open and Reproducible Science (CORES) is dedicating its 6th annual symposium to sorting out which is which.
The event runs September 24, 2026, as a half-day program at Stanford, structured around keynotes, lightning talks, and a panel discussion. The framing is practical, not philosophical: how do you actually build AI tools that make research more reliable instead of just faster and more prone to silent failure.
That distinction matters. A model that speeds up analysis but hides an error is worse than no automation at all. The symposium's agenda seems built around that tension, pairing optimistic technical demos with a panel specifically asking whether current AI systems clear the trust bar.
The keynote lineup covers three distinct angles on the reproducibility problem, each tackling it from a different layer of the research stack.
Lightning talks fill out the technical program with three shorter presentations. Jiacheng Miao presents Paper2Agent, a project reimagining research papers as interactive, queryable AI agents rather than static PDFs. Zijiao Chen discusses Brain Researcher, an agentic system that builds from knowledge graphs toward reproducible neuroimaging analysis. Ioana Ciuca addresses "cosmic co-discovery," essentially asking where AI-assisted scientific discovery in astronomy should even start.
There's also a shorter slot from Robert Rosenkranz of the Rosenkranz Foundation on AI-enhanced peer review, framed as a quality control mechanism for science broadly, not just a niche tooling improvement.
The panel, titled "Is AI Trustworthy Enough?", is where the symposium's skepticism gets a dedicated stage. Moderated by CORES Associate Director Maya Mathur, it features Sanmi Koyejo, Rob Reich, and Risa Wechsler. Koyejo works on trustworthy ML broadly, Reich brings an ethics and policy lens from Stanford's philosophy department, and Wechsler is a cosmologist, which suggests the panel isn't going to stay narrowly technical. Expect the conversation to range from model reliability metrics to who bears responsibility when an AI-assisted result turns out wrong.

Opening remarks come from James Landay, Stanford HAI's Faculty Director, and Russ Poldrack, CORES Faculty Director, who also presents the closing awards. Those awards cover Open Science, Open Source Software, and Data Sharing, categories that reward the kind of unglamorous infrastructure work that makes reproducibility possible in the first place.
The reproducibility crisis predates AI by a long stretch, going back to well-documented failures to replicate findings across psychology, biomedicine, and other fields in the 2010s. What's changed is the tooling available to address it.
Agentic systems that can read a paper, extract its methodology, and attempt to rerun the analysis represent a genuinely new capability. Paper2Agent and Brain Researcher both point in that direction: turning static research outputs into something interrogable and, ideally, verifiable by machine. If an agent can reconstruct a study's pipeline from its published methods and get the same result, that's a much stronger reproducibility signal than a methods section alone.
But the same agentic capability that makes replication easier also introduces new failure modes. An agent that misreads a methods section, or fills in an ambiguous parameter with a plausible-but-wrong default, can produce a confident, well-formatted result that's simply incorrect. That's the exact tension the "Is AI Trustworthy Enough?" panel is built to interrogate, and it's why the symposium pairs technical demos with a governance-and-ethics conversation rather than treating them as separate tracks.
Schmidt's ImageNetV2 work is a useful reference point here. It showed that benchmark performance can quietly decouple from real generalization when a field optimizes against the same test set for years. Applying that lesson to AI-for-reproducibility tools is an obvious next step: if agentic replication systems become widely used, someone needs to be checking whether they're actually catching errors or just producing confident agreement with whatever the original paper claimed.
The symposium's structure tells you something about where the field currently stands. There's real technical progress: agentic systems for replication, automated neuroimaging pipelines, AI-enhanced peer review proposals. But the organizers clearly felt it necessary to build in a dedicated trust-and-ethics panel rather than let the optimism stand unchallenged. For practitioners building or evaluating these tools, that's the right instinct. The question isn't whether AI can accelerate reproducibility work. It's whether it can do so without introducing new, harder-to-detect failure modes than the ones it's meant to solve.
Tags
Original Sources
CORES Annual Symposium 2026: AI and Scientific Reproducibility | Stanford HAI
↗ https://hai.stanford.edu/events/cores-annual-symposium-2026-ai-and-scientific-reproducibility
About the author
Kai built ML infrastructure at a Bay Area startup before developing an obsession with transformer architectures and inference optimisation that eventually pulled him out of product work entirely. A stint at a compute research lab sharpened his instinct for what actually matters in a model release versus what is marketing. He writes from the inside — from the perspective of someone who has debugged the systems he is describing at three in the morning. He is allergic to hype and instinctively drawn to the unglamorous plumbing questions that everyone else skips over.
More from The Engineer →This Week's Edition
18 September 2026
25 articles
Related Articles

Stanford HAI Talk Puts Arena's Human-Feedback Approach to AI Evaluation Under the Microscope
Models & Research · 5 min

Medicare's AI Push Meets a Discovery Problem for Health Chatbots
Products & Applications · 5 min

Microsoft's Internal AI Rollout Offers a Data Point on Enterprise Transformation Economics
Finance & Markets · 5 min
Related Articles

Stanford HAI Talk Puts Arena's Human-Feedback Approach to AI Evaluation Under the Microscope
Models & Research · 5 min

Medicare's AI Push Meets a Discovery Problem for Health Chatbots
Products & Applications · 5 min

Microsoft's Internal AI Rollout Offers a Data Point on Enterprise Transformation Economics
Finance & Markets · 5 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.