
Share
CoreWeave's new Forge platform merges experiment tracking, agent evaluation, and post-training into one infrastructure stack, betting that unifying the ML lifecycle beats stitching together point solutions for teams shipping real models.
CoreWeave just rolled out Forge, a platform that stitches together Weights & Biases experiment tracking, agent evaluation tooling, and serverless infrastructure for inference, post-training, and sandboxed compute, all running on CoreWeave's own GPU backend. If you've used Weights & Biases on their Multi-tenant Cloud before, the migration is painless: your projects, runs, artifacts, and account carry over automatically.
The pitch here is consolidation. Instead of bolting together separate tools for tracking experiments, evaluating agents, fine-tuning models, and serving inference, Forge tries to put all of it under one roof with shared infrastructure underneath. For teams tired of context-switching between five different dashboards to answer "why did this run regress," that's a meaningful quality-of-life improvement, assuming the integration actually holds up in practice.
Forge splits into two broad categories: tools for tracking and improving your work, and infrastructure for running it.
On the tracking and evaluation side:
On the infrastructure side, Forge offers three serverless products:

CoreWeave is also pointing people toward a Builder Resource Center with demos, guides, code samples, and events for working across these products, which suggests they're expecting a ramp-up period as teams figure out how the pieces fit together.
The interesting architectural bet in Forge is treating post-training and inference as serverless from the jump. That's not a new idea in the broader industry, but tying it directly to an experiment-tracking and evaluation layer that was previously a standalone product (W&B) is a notable move. It suggests CoreWeave wants to own more of the ML lifecycle on its infrastructure rather than just being the GPU layer underneath someone else's stack.
Serverless RL specifically is worth watching. RL-based post-training for LLMs has historically required a decent amount of custom infrastructure, reward modeling pipelines, rollout generation, and the compute orchestration to make them efficient. Wrapping that into a serverless offering alongside more standard SFT is a sign this workflow is maturing from research-lab bespoke setups into something closer to a managed product.
The agent evaluation angle, split across Weave and Agent Lens, also reflects where a lot of practitioner pain currently sits. Evaluating a single model's output against a benchmark is relatively well understood at this point. Evaluating a multi-step agent that calls tools, makes decisions, and chains actions together is a much messier problem, and tooling for it is still catching up to the pace at which people are actually deploying agents in production.
Whether Forge delivers on the "no infrastructure management" promise will come down to execution details that aren't fully spelled out yet: latency on the serverless inference endpoints, pricing for sustained post-training workloads, and how well ARIA's automated analysis holds up against hand-rolled debugging when a run goes sideways. For teams already in the Weights & Biases ecosystem, though, the migration path looks low-friction, and that alone lowers the bar for trying it.
Tags
Original Sources
About the author
Kai built ML infrastructure at a Bay Area startup before developing an obsession with transformer architectures and inference optimisation that eventually pulled him out of product work entirely. A stint at a compute research lab sharpened his instinct for what actually matters in a model release versus what is marketing. He writes from the inside — from the perspective of someone who has debugged the systems he is describing at three in the morning. He is allergic to hype and instinctively drawn to the unglamorous plumbing questions that everyone else skips over.
More from The Engineer →This Week's Edition
9 October 2026
34 articles
Related Articles

DOE Expands AI Assistant That Runs Particle Accelerators to 16 Institutions
Models & Research · 6 min

OpenAI's Latest Math Proofs Skip Key Transparency Steps Its Own Advisers Recommended
Models & Research · 5 min

OpenAI's Mass Math Release Raises a Basic Question: Who Do You Actually Talk To?
Models & Research · 5 min
Related Articles

DOE Expands AI Assistant That Runs Particle Accelerators to 16 Institutions
Models & Research · 6 min

OpenAI's Latest Math Proofs Skip Key Transparency Steps Its Own Advisers Recommended
Models & Research · 5 min

OpenAI's Mass Math Release Raises a Basic Question: Who Do You Actually Talk To?
Models & Research · 5 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.