
Share
A physics lab's AI system found a novel trajectory for an 80,000-year interstellar mission, while a fresh batch of puzzles reveals the gap between AI's chess-beating pedigree and its actual reasoning limits.
Puzzles have been a benchmark for machine intelligence since before "machine learning" was even a phrase people used casually. The term itself got popularized in a 1959 paper about an algorithm that learned to play checkers. Chess and Go followed as the canonical test beds, each one eventually falling to a sufficiently trained model.
The pace of progress on puzzles specifically has been startling. In late 2024, even top-tier models could only crack about 18% of New York Times Connections puzzles, the daily word-grouping game that trips up plenty of humans too. By early 2025, some of those same models were solving them almost perfectly, every time.
That kind of jump matters beyond bragging rights. Puzzles are a clean way to isolate where a model's reasoning actually breaks down, separate from the noise of open-ended tasks where "correct" is fuzzy. A new set of seven puzzles designed to stump AI models offers a fresh look at that boundary, letting readers test their own reasoning against the machines' blind spots.
The interesting part isn't that AI got good at Connections. It's why certain puzzle types remain hard even as others fall quickly.
That last point is the one worth sitting with if you build or evaluate models for a living. A benchmark score climbing from 18% to near-perfect in a few months doesn't necessarily mean the underlying reasoning got better across the board. It might just mean the training data caught up to the test.
This is the same dynamic that's playing out in a much higher-stakes arena: interstellar navigation. A nonprofit called the Fermi Explorer Mission announced plans to launch a spacecraft toward Alpha Centauri by the end of 2029. The math involved is almost absurd to contemplate. Alpha Centauri sits 4.4 light-years away, and even under the mission's best-case scenario, the spacecraft could take up to 80,000 years to arrive.
What makes the mission notable from an engineering standpoint isn't the destination. It's the route. The spacecraft will follow a novel trajectory discovered by an AI system built by PSI, a physics research lab. Finding efficient paths through gravitational fields across interstellar distances is a search problem with an enormous solution space, precisely the kind of thing where AI-assisted optimization can outperform traditional analytical methods by exploring possibilities a human physicist wouldn't think to try.

That's the throughline connecting a word puzzle to a decades-spanning space mission: both are about search spaces, and about whether a model's apparent competence generalizes or is narrowly scoped to the exact shape of the problem it was trained or tuned on. A model that nails Connections after fine-tuning on similar puzzles hasn't necessarily gotten smarter at ambiguity in general. An AI that finds a viable trajectory to another star system has solved a well-defined physics optimization problem, which is a different kind of achievement than general reasoning, but a genuinely impressive one on its own terms.
There's a broader context here too, and it's not just an academic curiosity. As missions like Fermi Explorer push the reasons-to-go-to-space conversation forward, a wave of new books is questioning why humans bother with any of this in the first place. Space anthropologist Deana L. Weibel frames exploration as a form of pilgrimage in her book The Ultraview Effect. Civilian astronaut Eiman Jahangir recounts a Blue Origin flight in A Heart for Space. Former NASA astronaut Leroy Chiao puts it more bluntly in Dinner with an Astronaut: people simply "need to know what's on the other side."
Whatever the motivation, the tools for getting there are increasingly AI-assisted. That's true whether the "there" is another star system or just a harder class of puzzle nobody's solved reliably yet.
The rapid climb in puzzle-solving benchmarks, from 18% to near-perfect on Connections within months, says more about targeted fine-tuning than about a generalized reasoning breakthrough. Practitioners evaluating models should treat sharp benchmark jumps with some skepticism until they understand what specifically changed in training.
Puzzles that require lateral reframing or ambiguous, culturally loaded clues remain a better proxy for reasoning limits than pattern-heavy formats, which are more susceptible to targeted fine-tuning.
The Fermi Explorer Mission's use of an AI-discovered trajectory for its Alpha Centauri run shows a more mature and narrower application of the same underlying capability: search over enormous solution spaces. It's a reminder that "AI found something a human couldn't" is often less about general intelligence and more about brute-force exploration done well.
Both threads point to the same practical lesson. Test where a model actually struggles, not just where it dazzles, and be precise about what kind of problem you're actually asking it to solve.
Tags
Original Sources
The Download: AI puzzles and a path to our nearest star system
↗ https://www.technologyreview.com/2026/09/02/1143283/the-download-ai-puzzles-alpha-centauri-mission
About the author
Kai built ML infrastructure at a Bay Area startup before developing an obsession with transformer architectures and inference optimisation that eventually pulled him out of product work entirely. A stint at a compute research lab sharpened his instinct for what actually matters in a model release versus what is marketing. He writes from the inside — from the perspective of someone who has debugged the systems he is describing at three in the morning. He is allergic to hype and instinctively drawn to the unglamorous plumbing questions that everyone else skips over.
More from The Engineer →This Week's Edition
3 September 2026
45 articles
Related Articles

Predicting Antenna Interference on Complex Platforms Before Anyone Builds Anything
Models & Research · 5 min

HIMSS Overhauls Digital Health Scoring System to Reckon With AI's Reach
Policy & Regulation · 6 min

What an Airport's AI Failures Can Teach Hospitals About Playing It Safe
Products & Applications · 5 min
Related Articles

Predicting Antenna Interference on Complex Platforms Before Anyone Builds Anything
Models & Research · 5 min

HIMSS Overhauls Digital Health Scoring System to Reckon With AI's Reach
Policy & Regulation · 6 min

What an Airport's AI Failures Can Teach Hospitals About Playing It Safe
Products & Applications · 5 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.