
Share
As transformers show their age, a wave of startups is pushing the boundaries to develop the next generation of large language models (LLMs+).
Back in 2017, Google researchers introduced transformers with "Attention Is All You Need," revolutionizing how we process text data. Fast forward nine years, and transformers are the backbone of every major large language model (LLM) on the market. But as Justin Dangel, cofounder and CEO of Subquadratic, points out, "Transformers are one of the most important innovations in computer science, but they're starting to show their age."
Recent advancements in LLMs, such as reasoning models and handling large inputs, often feel like workarounds rather than natural extensions of transformer technology. This has led a growing number of scientists and engineers to ask what's next for these powerful models.
The key strength of transformers lies in dense attention, which encodes the meaning of text blocks into numerical representations. However, this mechanism also introduces significant limitations. For instance, dense attention requires quadratic computational complexity with respect to input length, making it inefficient for processing very long sequences. This inefficiency is a major bottleneck as LLMs are increasingly expected to handle complex tasks and large datasets.
These limitations have prompted researchers and startups to explore alternative architectures and techniques that can overcome these challenges. Some of the most promising approaches include sparse attention, linear transformers, and hybrid models that combine elements from different paradigms.

Several startups are leading the charge in developing the next generation of LLMs, collectively known as LLMs+. These companies are not just tweaking existing transformer architectures but are fundamentally rethinking how language models can be built and optimized.
While these innovations are exciting, they also come with challenges. For instance, sparse attention and linear transformers require careful tuning and optimization to ensure they perform well on a wide range of tasks. Hybrid models can be more complex to implement and train, requiring significant expertise and resources.
Despite these hurdles, the potential rewards are substantial. As Dangel notes, "The companies that crack this problem will have a significant advantage in the AI landscape." The race is on, and the future of language models could look very different from what we know today.
In the coming years, we can expect to see more breakthroughs in LLM technology, driven by both established players and innovative startups. The next generation of LLMs+ will likely be characterized by greater efficiency, scalability, and versatility, paving the way for new applications and use cases in natural language processing and beyond.
Tags
Original Sources
These startups are chasing the next big thing in LLMs
↗ https://www.technologyreview.com/2026/08/10/1141511/these-startups-are-chasing-the-next-big-thing-in-llms
About the author
Kai built ML infrastructure at a Bay Area startup before developing an obsession with transformer architectures and inference optimisation that eventually pulled him out of product work entirely. A stint at a compute research lab sharpened his instinct for what actually matters in a model release versus what is marketing. He writes from the inside — from the perspective of someone who has debugged the systems he is describing at three in the morning. He is allergic to hype and instinctively drawn to the unglamorous plumbing questions that everyone else skips over.
More from The Engineer →This Week's Edition
17 August 2026
113 articles
Related Articles

Suki Researchers Challenge Traditional AI Note Evaluation Methods in Healthcare
Models & Research · 3 min

The Path to Distributed Artificial Superintelligence: Connecting AI Agents for Better Coordination
Models & Research · 4 min

LLM Security Flaw Exposed and Geothermal Power Revived
Models & Research · 4 min
Related Articles

Suki Researchers Challenge Traditional AI Note Evaluation Methods in Healthcare
Models & Research · 3 min

The Path to Distributed Artificial Superintelligence: Connecting AI Agents for Better Coordination
Models & Research · 4 min

LLM Security Flaw Exposed and Geothermal Power Revived
Models & Research · 4 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.