
Share
Koa, built on Nvidia's open-weight Nemotron and trained on synthetic sales and support data, signals enterprises don't need Claude or GPT-class reasoning for every agent task, just the right task-specific one.
Salesforce dropped its biggest AI announcement of Dreamforce this week, and it's not another agent framework or dashboard. It's a reasoning model called Koa, and the interesting part isn't that it exists, it's what it's built on and why.
Koa is post-trained on top of Nvidia's open-weight Nemotron model, and it's purpose-built for sales, marketing, and customer-support work. That's a meaningfully different bet than what the frontier labs are pushing. OpenAI and Anthropic want enterprises pumping proprietary data, prompts, and feedback into general-purpose models, often at significant cost. Salesforce is betting its customers would rather have a smaller, cheaper, task-specific model that never touches their raw customer data in the first place.
Here's the pitch, broken down:
That last point is the quiet one that matters most to enterprise buyers. Security and compliance teams don't love routing sensitive workflows through third-party APIs they don't fully control. A model that lives inside Salesforce's existing trust boundary sidesteps that entire conversation.
"We've built many small task-specific language models, which are part of Agentforce's portfolio," Jayesh Govindarajan, EVP of Salesforce AI, told TechCrunch. "But reasoning has always been something that we've relied on the frontier model providers for. Until now."
Before Koa, any Agentforce task requiring multi-step reasoning, the kind of long-running chain-of-thought work that simple task models can't handle, got routed through the gateway to Claude or ChatGPT. That worked, but it meant every complex query left Salesforce's infrastructure and hit the token-metered pricing of a frontier API. Koa is meant to intercept a chunk of that traffic and handle it in-house, more cheaply.
The obvious question is why Salesforce didn't just train its own base model from scratch, or lean on one of the strong open-weight options already out there, like Alibaba's Qwen. Govindarajan's answer is pretty blunt: provenance.
"One of the reasons we hadn't done this before, train our own enterprise-grade frontier model, we always wanted to, but the challenge has always been the lack of a pre-trained base model to start with," he said. "Until Nemotron came along, there was no sovereign American pre-trained model that was available, one, and two, that was state of the art, and, three, that had clear data provenance. We have no idea what Qwen trains on."

That's a real constraint for enterprise buyers, not just a talking point. If you're selling into regulated industries, "we don't know what's in the training data" is a dealbreaker, regardless of how good the benchmarks look. Nemotron gives Salesforce a base model with a known, auditable training pipeline, built by a company that already sells them the chips.
Post-training Koa didn't involve touching any real Salesforce customer data either. Instead, the two companies built synthetic training scenarios designed to mimic what actual customer interactions look like.
"We actually simulated a customer service environment with a persona customer service professional, including irate customers that call into the customer service center, all the way to a sales professional who's trying to close a deal," Govindarajan said.
That's a sensible approach for a company sitting on enormous amounts of customer data it legally and contractually can't use for training. Synthetic persona generation lets you capture the shape of real interactions, angry callers, negotiation patterns, escalation paths, without ever exposing an actual customer record to a training run. It also sidesteps the murkier legal territory other AI companies have wandered into around data provenance and consent.
On the inference side, Nvidia's pitch is about efficiency rather than raw capability. Kari Ann Briski, Nvidia's VP of Generative AI Software for Enterprise, described Nemotron's architecture as optimized for token efficiency specifically.
"It's kind of the trifecta of things that you need to have: sovereign AI, time to first token, efficient reasoning, for the tokenomics of it all," Briski told TechCrunch.
Time-to-first-token and token-per-task efficiency are the metrics that actually determine your AI bill at scale, more than raw benchmark scores do. If Koa genuinely needs fewer tokens to complete the same customer-support ticket or sales workflow that a frontier model would handle, that compounds fast across millions of agent calls a month.
Worth noting: Salesforce isn't torching its frontier-model relationships here. It also just announced Claudeforce, a partnership letting Anthropic's Claude serve as an AI interface while customer data stays inside Salesforce's own systems of record. So the strategy isn't "replace frontier models," it's "route intelligently." Simple, high-volume tasks go to Koa. Complex or open-ended reasoning can still go to Claude, but under Salesforce's data controls rather than a raw API call.
Koa is less a shot at Anthropic or OpenAI and more a bet that enterprise AI is bifurcating into two markets: general-purpose frontier reasoning for open-ended problems, and cheap, governed, task-specific models for the repetitive high-volume work that actually drives AI spend. Nemotron's open weights and clean data provenance gave Salesforce a credible path to build the latter without waiting on, or paying, the frontier labs. Whether other enterprise software vendors follow the same playbook, building on open-weight bases with synthetic post-training rather than licensing frontier APIs outright, is the thing worth watching over the next year.
Tags
Original Sources
Salesforce and Nvidia's new reasoning model is everything the AI labs should fear | TechCrunch
↗ https://techcrunch.com/2026/09/15/salesforce-and-nvidias-new-reasoning-model-is-everything-the-ai-labs-should-fear
About the author
Kai built ML infrastructure at a Bay Area startup before developing an obsession with transformer architectures and inference optimisation that eventually pulled him out of product work entirely. A stint at a compute research lab sharpened his instinct for what actually matters in a model release versus what is marketing. He writes from the inside — from the perspective of someone who has debugged the systems he is describing at three in the morning. He is allergic to hype and instinctively drawn to the unglamorous plumbing questions that everyone else skips over.
More from The Engineer →This Week's Edition
16 September 2026
31 articles
Related Articles
Related Articles
More Stories
© 2026 Cedar & Bloom. All rights reserved.