
Share
Announced at VMware Explore 2026, the new platform bundles inference, agentic AI, and traditional workloads onto one private cloud stack, with 150+ supported models, tokenomics tooling, and Zero Trust security baked in from day one.
Broadcom used VMware Explore 2026 in Las Vegas to launch VMware Private AI Cloud, and the pitch is simple even if the engineering underneath isn't: stop shipping enterprise data out to wherever the model lives, and instead bring the model to the data. For any organization that's been wrestling with data residency requirements, GPU cost sprawl, or the general chaos of bolting AI onto existing infrastructure, that's a meaningful reframe.
The platform is built on VMware Cloud Foundation (VCF) 9 and is designed to run inference workloads, agentic applications, and standard enterprise apps side by side on a single private cloud. Ram Velaga, president of Broadcom's Infrastructure Software Group, framed it as a convergence moment: "VMware Private AI Cloud is the inflection point where enterprise private cloud and private AI infrastructure stop operating as separate disciplines and become one, enabling production inference workloads and agentic AI with the data sovereignty, compliance posture, and cost predictability their business demands."
That convergence claim is worth unpacking, because it's really three separate problems Broadcom is trying to solve at once: hardware costs, operational overhead, and token economics.
On the hardware side, VCF 9 leans on NVMe memory tiering and cluster-wide storage deduplication to cut costs, and it supports heterogeneous clusters spanning GPUs, CPUs, and accelerators from multiple vendors plus OEM/ODM server hardware. Translation for practitioners: you're not locked into a single silicon vendor's roadmap just to run production AI. On the tokenomics front, the platform adds token monitoring, multi-tenant model sharing, GPU/vGPU tracking, and an AI metrics observability dashboard, giving ops teams visibility into where compute and tokens are actually going instead of guessing from a monthly bill.
The centerpiece of the infrastructure story is VMware AI Factory, described as the software-defined foundation of Private AI Cloud. It automates the path from bare metal to a deployed model, plus Day 2 operations, the ongoing patching, scaling, and lifecycle work that tends to eat engineering time after initial rollout. Broadcom says the goal is faster time to first model deployment and tighter control over tokenomics, which matters because a lot of enterprise AI projects stall not at the proof-of-concept stage but during the slow grind of productionizing.
Paired with AI Factory is what Broadcom calls Model as a Service for VCF: customers can now run more than 150 open source and commercial models on-premises, including Nemotron 3, Gemma 4, cotomi, Qwen, and GLM 5.2. That's a notably broad menu, and it signals Broadcom is trying to position VCF as model-agnostic infrastructure rather than a platform that locks customers into one provider's stack. Models get delivered as a service to internal user communities through VCF's built-in services, which is the kind of detail that matters if you're trying to avoid standing up separate serving infrastructure for every model your data science team wants to try.
Security gets equal billing here, and it's aligned to NIST CSF 2.0, the cybersecurity framework many enterprises already use for compliance mapping. The approach is defense-in-depth: automated non-disruptive updates, plus VMware vDefend handling virtual patching and hypervisor-level lateral security through microsegmentation to enforce Zero Trust. Avi Load Balancer adds web application firewall and API protection on top.

A few specific security pieces stand out:
That last category, agentic-specific defenses, points to where Broadcom clearly sees the real risk emerging. Autonomous agents don't behave like traditional applications. They can act unchecked, exceed their intended scope, or just misread instructions, and none of the usual perimeter security assumptions hold up cleanly.
To address that, Broadcom is introducing AgentMinder, described as a central control plane for governing autonomous agents at scale. It treats each agent as an enterprise-grade identity, binding its authority to a specific mission, an approved toolset, and authorized resources. It enforces runtime policy, so every tool invocation gets checked against least-privilege access, and it logs everything for compliance-grade auditability. If you've been worried about agents quietly accumulating permissions or wandering outside their intended lane, this is Broadcom's answer.
Tanzu Platform rounds out the agentic story with two pieces: AI-ready data foundations that let data owners build pipelines across structured and unstructured data, publishing governed data products to a marketplace that agents and developers can discover without data leaving the enterprise; and agent foundations that enforce a deny-by-default runtime, meaning agents get zero access to APIs, networks, MCP servers, or the internet unless explicitly granted. A new isolated credential store shields credentials from agents entirely, which is a direct, sensible countermeasure against credential theft and prompt injection attacks. You can't leak what you literally cannot see.
VMware Private AI Cloud is Broadcom's attempt to collapse three previously separate concerns, private cloud infrastructure, AI model serving, and agentic AI governance, into one operational surface. The hardware flexibility and 150-plus model catalog make a real case for cost-effective on-premises AI at scale, while the security architecture, particularly AgentMinder and the deny-by-default agent runtime, addresses a genuine and growing gap in how enterprises control autonomous systems. Whether the "one platform" pitch holds up in practice will depend on how cleanly these pieces integrate once IT teams start deploying them for real, but the direction, keeping models close to data instead of the reverse, tracks with where enterprise AI concerns are actually headed.
Tags
Original Sources
Broadcom Announces VMware AI Factory, Enabling Faster Time to Production AI and Greater Control Over AI Tokenomics
↗ https://www.broadcom.com/company/news/releases/broadcom-introduces-vmware-private-ai-cloud
About the author
Kai built ML infrastructure at a Bay Area startup before developing an obsession with transformer architectures and inference optimisation that eventually pulled him out of product work entirely. A stint at a compute research lab sharpened his instinct for what actually matters in a model release versus what is marketing. He writes from the inside — from the perspective of someone who has debugged the systems he is describing at three in the morning. He is allergic to hype and instinctively drawn to the unglamorous plumbing questions that everyone else skips over.
More from The Engineer →This Week's Edition
1 September 2026
22 articles
Related Articles

Agentic AI Is Reshaping the Analytics Stack, But Judgment Remains a Human Asset
Products & Applications · 5 min

Stryker Gets FDA Clearance for Surgical Navigation System Built on Apple Vision Pro
Products & Applications · 5 min

Solera Health Adds Noom to Its Curated Weight Management Network
Products & Applications · 5 min
Related Articles

Agentic AI Is Reshaping the Analytics Stack, But Judgment Remains a Human Asset
Products & Applications · 5 min

Stryker Gets FDA Clearance for Surgical Navigation System Built on Apple Vision Pro
Products & Applications · 5 min

Solera Health Adds Noom to Its Curated Weight Management Network
Products & Applications · 5 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.