
Share
A physics lab's AI system charted a novel path to our nearest star system, underpinning a nonprofit's 2029 launch plan, while a fresh batch of AI puzzle benchmarks shows just how fast (and unevenly) models are improving.
Puzzles have been baked into AI research since the field's earliest days. The term "machine learning" itself got popularized back in 1959, in an article about an algorithm learning to play checkers. Chess and Go followed as the classic test beds. Now there's a new front: the humble word puzzle, and it's revealing something interesting about where large language models actually stand.
Take the New York Times' Connections puzzle, a grid-based word association game. In late 2024, even the best models solved only 18% of them correctly. By early 2025, some models were solving them nearly perfectly, every single time. That's not incremental progress, that's a step function. And it raises a question practitioners should care about: what changed under the hood to produce that jump, and does it generalize?
Researchers have now assembled seven puzzles that have tripped up models at various points, specifically to probe where AI still falls short compared to humans. The value here isn't just novelty. Watching where a model succeeds or fails on a constrained, verifiable task gives you a much cleaner signal about reasoning capability than open-ended benchmarks do. Puzzles have fixed answers. There's no ambiguity in grading, no room for a model to talk its way into a passing score. That makes them a useful diagnostic even as they double as entertainment. If you want to see how you stack up against the machines, MIT Technology Review has published a test built from these seven puzzles.
Meanwhile, in a much bigger application of algorithmic problem-solving, a nonprofit called the Fermi Explorer Mission announced it plans to launch a spacecraft toward the Alpha Centauri system by the end of 2029.
The scale here is genuinely hard to grasp. Alpha Centauri sits 4.4 light-years from Earth. Even with an optimistic mission profile, the spacecraft could take up to 80,000 years to arrive. That's longer than recorded human history, several times over.
What makes the mission notable from a technical standpoint isn't the launch date. It's the trajectory. The path the spacecraft will follow was discovered by an AI system built by PSI, a physics research lab, and it's described as a genuinely novel route, not a refinement of existing orbital mechanics playbooks. Trajectory design for interstellar or even interplanetary missions is a brutal optimization problem: you're balancing gravitational assists, fuel constraints, launch windows, and travel time across a search space that's too large for humans to fully explore by hand. That's exactly the kind of problem where a well-tuned optimization or search-based AI system can outperform traditional methods, by finding a solution humans wouldn't have thought to try.
A few things worth flagging for anyone following this:

It's a strange but fitting pairing of stories: AI getting startlingly good at solving contained, human-designed puzzles on one hand, and AI charting a path across the literal galaxy on the other. Same underlying capability, wildly different stakes.
A few other items from the AI world this week are worth a practitioner's attention, even in brief.
OpenAI is reportedly restricting access to its next model, internally called Astra, after rating it a "critical" cyber risk. Testing apparently showed the model could automate cyberattacks, and OpenAI says it's the first model to cross that internal threshold. The company plans additional security measures before wider release. This follows other recent safety stumbles at the company, including a Hugging Face-related hack that raised questions about internal culture.
On the coding front, Google is reportedly close to shipping a new model, possibly named Gemini 3.8 Flash, that narrows the coding capability gap with OpenAI and Anthropic. If the reports hold up, this matters for anyone choosing between model providers for code generation workloads, though it's worth noting that not everyone agrees AI coding tools are living up to the hype in production settings.
And in the safety research world, Marius Hobbhahn, CEO of Apollo Research, made a pointed comment to the Guardian about the urgency of alignment work: if you build something vastly smarter than you, it better be on your side. That's not a new sentiment in AI safety circles, but it lands differently against the backdrop of models crossing "critical" risk thresholds in live testing.
The puzzle benchmarks tell a clear story: models went from solving fewer than one in five Connections puzzles to nearly acing them within a matter of months, which says more about how quickly targeted capabilities can improve than it does about general intelligence. Puzzles remain a useful, low-noise way to spot exactly where models diverge from human reasoning.
The Alpha Centauri mission is a reminder that AI-driven optimization isn't just winning at games or writing code. It's finding genuinely novel solutions to physics problems that have stumped human trajectory designers, even if the payoff (a spacecraft arriving in 80,000 years) is about as far from immediate as it gets. Combined with the same week's news on cyber-risk model restrictions, it's a snapshot of an industry pushing capability limits on every axis at once: language, physics, and security, sometimes all in the same news cycle.
Tags
Original Sources
The Download: AI puzzles and a path to our nearest star system
↗ https://www.technologyreview.com/2026/09/02/1143283/the-download-ai-puzzles-alpha-centauri-mission
About the author
Kai built ML infrastructure at a Bay Area startup before developing an obsession with transformer architectures and inference optimisation that eventually pulled him out of product work entirely. A stint at a compute research lab sharpened his instinct for what actually matters in a model release versus what is marketing. He writes from the inside — from the perspective of someone who has debugged the systems he is describing at three in the morning. He is allergic to hype and instinctively drawn to the unglamorous plumbing questions that everyone else skips over.
More from The Engineer →This Week's Edition
6 September 2026
41 articles
Related Articles
Related Articles
More Stories
© 2026 Cedar & Bloom. All rights reserved.