
Share
At its first big Surface event since Copilot+ PCs launched, Microsoft bet on "hybrid intelligence": local models, a new containment system for agents, and a laptop with RTX GPU power built in.
Six months ago I built a small app that ran a local model on my Surface Laptop 7 to make sense of troubleshooting signals without sending data to the cloud. I was proud of it. Then Microsoft ran a 137-billion-parameter coding model locally with Airplane mode on, live on stage, and my little app suddenly felt very small.
That happened yesterday at 19:00 Swedish time, when Microsoft went live from San Francisco with its Windows and Surface event, the first major one since the Copilot+ PC launch in May 2024. Satya Nadella, Windows chief Pavan Davuluri and NVIDIA's Jensen Huang shared the stage. Their message: the next chapter of the PC is about agents, and a lot of the intelligence is moving back onto the device in front of you.
Microsoft calls this hybrid intelligence. Windows becomes the place where agents run locally when that makes sense and reach for the cloud when they need to, with the security and manageability organizations already expect. The reasoning is practical, not philosophical. Models keep growing. Our needs are growing faster than our cloud budgets. Every token sent to the cloud costs money and time, and some data simply shouldn't leave the device. Mix local and cloud intelligently, and you spend cloud capacity only where it's actually needed.
Four pieces have to work together to make this real, according to Microsoft: capable local models, smart routing between local and cloud, a fast runtime, and silicon that can carry the load. Yesterday they showed all four.
Surface Laptop Ultra is built around NVIDIA's RTX Spark superchip: a Blackwell RTX GPU with up to 6,144 cores paired with a Grace CPU with up to 20 cores. Microsoft optimized Windows for the chip first, then designed the entire device around it, with mechanical, thermal, acoustic and material engineers working alongside industrial designers from day one.
The specs back up the ambition. Up to 128 GB of unified memory shared between CPU and GPU. Up to 1 petaflop of AI performance. Models over 120 billion parameters running locally. A 15-inch mini-LED PixelSense Ultra touchscreen at up to 2,000 nits, 120 Hz, 262 PPI, though touch only, no pen support. Cooling got 2.5 times the thermal capacity of the previous Surface Laptop. The SSD is user-removable. Battery life hits 15 hours of local video playback. It's under 18mm thin and under 4.5 lbs, which is less ultrabook and more portable workstation.
Pricing starts at $2,599 in the US, with pre-orders open now and shipping from October 16. The 128 GB configuration I actually want lists at $5,899.99. In Sweden, Surface Laptop Ultra for Business starts at 42,309 SEK, ordered by phone through the Microsoft Store.
One detail that mattered more to me than it should: Magnetic Connect. Back in May 2025, Microsoft dropped the magnetic Surface Connect port from the 12-inch Surface Pro and 13-inch Surface Laptop in favor of USB-C everywhere. I argued it was worth the trade. Turns out we didn't have to choose. Magnetic Connect is a standard USB-C port with a magnet built in, still carrying data and video. My charger drawer, and my clumsier colleagues, are equally relieved.

For desks instead of backpacks, there's the Surface RTX Spark Dev Box: same RTX Spark platform, same 1 petaflop and 128 GB unified memory ceiling, part of the Project Zenith family. It ships preloaded with VS Code, Git, GitHub CLI, GitHub Copilot, WSL, Python, Node and PowerShell 7, costs $5,999, sells only through Microsoft.com in the US, and ships in November. The chassis has 1,000 air vents, a nod to its 1,000 teraflops. Someone on that team has a sense of humor.
Agents don't behave like regular apps. They run continuously, call tools, write code, touch files, often with nobody watching every step. Microsoft put it plainly: "An agent cannot be its own security authority." That's the thinking behind Microsoft Execution Containers (MXC), now generally available on Windows 11. You declare which files and network destinations an agent may use, and Windows enforces that boundary externally, so neither the agent nor any code it generates can grant itself more access.
There are four isolation levels: process containers for fast, lightweight containment of generated code (ideal for coding agents, cross-platform on Windows, macOS and Linux); session containers for long-running agents needing their own desktop and account (Windows only); WSL containers for agents living in the Linux toolchain (Windows only); and MicroVMs for higher-risk workloads needing a hardware-backed boundary (Windows and Linux, experimental).
The standout feature is Learning mode: process containers can generate a report of everything an agent tried to access, letting you build least-privilege policy from evidence instead of guesswork. Anyone who's rolled out ASR rules or App Control will recognize exactly why that matters.
Adoption is already underway. GitHub Copilot, OpenAI Codex, OpenClaw, Replit, LM Studio, NVIDIA OpenShell and Unsloth AI support MXC today, with Anthropic Claude Code, Perplexity, Manus and Box on the way. Meta's Muse is coming to Windows as a native app with MXC built in. Jensen Huang said on stage that MXC "is going to revolutionize how agents are built and deployed." Still missing, but coming: Intune policy for MXC process containers, agent identity through Microsoft Entra to distinguish agent activity from user activity, and Agent 365 controls for local agents.
On the model side, MAI Code 1.1 Flash is Microsoft's coding model: 137 billion total parameters, 6.8 billion active at any time, now squeezed to 3-bit precision for an 80% size reduction while keeping coding quality and the full 256K context window. On Surface Laptop Ultra it uses at most about 75.5 GB of memory at full context, and reads prompts at over 900 tokens per second at 64K context. GitHub's HydraFusion will handle routing between local and cloud models in GitHub Copilot app, CLI and VS Code, arriving in experimental preview later in October. More local models are coming too, including an NVIDIA Nemotron model under 20 GB and DeepSeek V4 Flash, alongside llama.cpp support in Windows ML.
Copilot+ PCs already run over 2 trillion inferences locally every month, and more than 40% of business laptops being built today qualify as Copilot+ PCs. Over the coming months, Copilot gains local context (understanding your files and recent activity with permission), local actions (organizing files, running diagnostics, troubleshooting, coding) and local models that defer to the cloud only when needed.
For anyone managing Windows fleets, the priorities are concrete. Start testing MXC now using the SDK, configuration schema and samples on GitHub, and run at least one agent in Learning mode to see what it actually tries to access. Prepare for agent identity in Entra, which will let you restrict an agent's access without blocking the human using the device. Treat RTX Spark machines like Windows on Arm devices and test your VPN clients, security agents and line-of-business apps before ordering them at scale. Buy for memory, not clock speed: a 24 GB machine and a 128 GB machine are fundamentally different tools for local AI. And check DFCI support before rolling out any new Surface model, because management basics still come first. Microsoft says more detail on hybrid intelligence for organizations is coming at Ignite in November, which is worth a spot on your calendar regardless of where you land on the hardware debate.
Tags
Original Sources
Windows Got Hybrid Intelligence and I neeeed a Surface Ultra - Mr T-Bone´s Blog
↗ https://www.tbone.se/2026/10/08/windows-got-hybrid-intelligence-and-i-neeeed-a-surface-ultra
About the author
Kai built ML infrastructure at a Bay Area startup before developing an obsession with transformer architectures and inference optimisation that eventually pulled him out of product work entirely. A stint at a compute research lab sharpened his instinct for what actually matters in a model release versus what is marketing. He writes from the inside — from the perspective of someone who has debugged the systems he is describing at three in the morning. He is allergic to hype and instinctively drawn to the unglamorous plumbing questions that everyone else skips over.
More from The Engineer →This Week's Edition
9 October 2026
34 articles
Related Articles

QuantumScape Brings Solid-State Batteries Into the Data Center Rack
Products & Applications · 5 min

Chinese Developer Shuts Down ARTEX AI Agent After Link to South Korean Bank Hacks
Products & Applications · 5 min

UNC Chapel Hill Opens Funding Round for Faculty and Staff AI Projects
Products & Applications · 6 min
Related Articles

QuantumScape Brings Solid-State Batteries Into the Data Center Rack
Products & Applications · 5 min

Chinese Developer Shuts Down ARTEX AI Agent After Link to South Korean Bank Hacks
Products & Applications · 5 min

UNC Chapel Hill Opens Funding Round for Faculty and Staff AI Projects
Products & Applications · 6 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.