Microsoft's Hybrid Intelligence: How Windows Routes AI Between Local and Cloud
Microsoft's new Windows routing system decides whether coding agent tasks run locally or in the cloud, using on-device models.

What is Microsoft’s hybrid intelligence routing system?
Hybrid intelligence is Microsoft’s term for a system built into Windows and GitHub Copilot that automatically decides whether an AI coding task runs on your local machine or gets sent to a cloud model. Microsoft showed it off at its Windows and Surface event in San Francisco, framing it as three pieces working together: a router that picks the right model for the job, local models small enough to run on a laptop, and new hardware (built around Nvidia’s RTX Spark chip) powerful enough to run them. The headline demo involved a coding agent that processed 1.6 million tokens during a single session at no cost, because nearly all of that traffic stayed on the device.
TL;DR
- Microsoft demoed a GitHub Copilot session where a coding agent processed roughly 1.6 million tokens locally and only about 10,000 tokens through the cloud, a ratio of around 150 to 1.
- The “auto” model picker in Copilot now chooses between cloud and local models on its own. The developer never explicitly chose where the task would run.
- A common pattern emerged in the demo: a cloud model handled planning (one of OpenAI’s GPT models), while local sub-agents running Microsoft’s own on-device model did the repetitive, token-heavy work.
- Microsoft is shrinking models aggressively through quantization, including a coding model cut to 3-bit precision and a version of DeepSeek reportedly run at an average of 1.6 bits, raising real questions about whether quality holds up on hard tasks.
- Windows ML, Microsoft’s runtime for running models across CPUs, GPUs, and NPUs, is adding support for Llama.cpp, which opens the door to running open and fine-tuned models from the broader community rather than waiting on Microsoft-packaged versions.
- A new sandboxing layer called Microsoft execution containers (MXC) lets organizations set policies on what an agent can access, with every action attributed to the agent’s own identity rather than the logged-in user.
- The feature starts rolling out in the GitHub Copilot app on October 15, with local context, local actions, and local models expanding over the following months.
Other agents start typing. Remy starts asking.
Scoping, trade-offs, edge cases — the real work. Before a line of code.
How does the router decide between local and cloud models?
In the demo, the presenter left Copilot’s model picker set to “auto.” For a routine cleanup task on the Windows Terminal codebase, auto routed the job to Microsoft’s own local coding model, visibly spinning up the laptop’s GPU. To prove the point, the presenter put the machine into airplane mode and the task kept running, since nothing needed to leave the device.
A harder task, sorting through a batch of GitHub issues tagged “needs repo,” triggered a different pattern. A cloud model (reportedly one of OpenAI’s GPT-series models) handled the planning step, then spun up multiple sub-agents that ran locally on Microsoft’s on-device model to do the bulk of the reading and investigation. This is the same division of labor showing up across the agentic AI space generally: an expensive, capable model does the thinking, and cheaper local models chew through large volumes of context.
That division is also where the 1.6 million token figure comes from. Long-running coding agents tend to pull in huge amounts of context (file contents, issue threads, logs) relative to what they actually output. Routing that bulk reading to a free, local model instead of a metered cloud API is where the cost savings show up.
What happens when the local model gets it wrong?
This is the open question Microsoft didn’t answer on stage. Every routing system is only as good as its fallback behavior. If a local model can’t actually handle the task it’s been handed, does the system detect that and escalate to the cloud, or does the user just get a worse answer with no indication anything went wrong?
Microsoft didn’t demonstrate that failure path, which means it’s still unclear whether hybrid intelligence will degrade gracefully or just degrade. This is a familiar problem for anyone who has worked with model routing systems generally, where a router’s real test isn’t the easy cases, it’s the moment it misjudges a task.
Which models actually run locally, and what are the tradeoffs?
Microsoft named three models designed to run on-device:
- Microsoft AI code 1.1 flash, a coding model first introduced at Microsoft’s Build conference, now quantized down to 3 bits. That’s roughly an 80% size reduction, and it’s reportedly running with a 256k context window fully on-device.
- A coming Neotron model quantized to 2 bits, running in about 20 gigabytes of memory, which Microsoft positioned as a good match for the RTX Spark hardware.
- A version of DeepSeek run at an average of 1.6 bits in roughly 60 gigabytes of memory, with Microsoft claiming on stage that it outperforms on coding and reasoning benchmarks, without specifying which benchmarks or by how much.
Remy is new. The platform isn't.
Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.
The quantization numbers are aggressive. Pushing a model down to 3 bits or lower cuts memory and compute requirements dramatically, but it also raises real questions about quality, especially for coding tasks where precision matters. Microsoft’s claim that the DeepSeek figure is an “average” suggests a mixed approach: keeping the layers that matter most at higher precision while pushing less-active layers much lower. Whether that preserves real-world performance is the thing to watch for. Three areas worth scrutinizing as these models reach independent testing: whether coding pass rates hold up against higher-precision versions of the same model, whether quality falls apart at long context lengths, and whether tool calls stay clean under agentic workloads. Broken structured output is often the first casualty of heavy quantization, and a broken tool call generally means a failed task for an agent.
There’s also a cost to long context windows that’s easy to overlook: at 256k tokens, the KV cache itself consumes significant memory on top of the model’s weights, which likely means the larger-memory configurations of this hardware are necessary to actually use that context length in practice.
Why does adding Llama.cpp to Windows ML matter?
Windows ML is Microsoft’s runtime for deploying AI models across different hardware (CPUs, GPUs, NPUs) without developers needing to write separate code paths for each. Bringing Llama.cpp support into it is a meaningful shift because it opens Windows up to the broader open-model ecosystem rather than locking users into Microsoft-packaged models.
Practically, that means when a new open model ships with a GGUF quantized version, often within hours of release, Windows users don’t have to wait for Microsoft to officially support it. That includes fine-tuned or uncensored community variants that Microsoft would be unlikely to package itself. For developers building Windows apps that ship with a local model, it also means one runtime that handles whatever hardware a given user has, rather than needing to target specific chips.
What’s still unclear is whether third-party or custom models can be plugged into the Copilot router itself, or whether routing stays limited to Microsoft’s own models. If users can eventually swap in any compatible open model (a mixture-of-experts model, for instance) and have the OS route to it automatically, that would meaningfully expand what “hybrid intelligence” means in practice.
What is Microsoft doing about agent security and sandboxing?
Running capable models locally raises an obvious question: how do you stop an autonomous agent from doing something you didn’t intend? Microsoft’s answer is Microsoft execution containers (MXC), a runtime sandbox comparable in concept to Nvidia’s sandboxing approach or Docker-style containers.
MXC lets organizations define policies for what an agent can access, with Windows enforcing those policies while the agent runs. Notably, every action an agent takes is attributed to the agent’s own identity, not the logged-in user’s. That separation matters for auditing and tracing what happened, and it also means a user could revoke an agent’s access to a specific action without losing their own access to the system.
Is the RTX Spark hardware worth the price?
The hardware side of this announcement centers on the RTX Spark, built with Nvidia and available in laptop and desktop form factors, with a Windows version of Nvidia’s DGX station teased for later this year. Configurations go up to 128 gigabytes of unified memory, with full CUDA support running natively on Windows, no Linux dual-boot required. Independent coverage has placed the GPU performance in the range of an RTX 5070, with around 20 CPU cores and a power draw under 80 watts.
Pricing reportedly starts around $2,600, climbing with memory configuration. Given current pressure on memory and RAM pricing industry-wide, that’s a real cost barrier, and one that’s likely to apply across local AI hardware generally rather than being specific to Microsoft or Nvidia’s choices.
Remy doesn't write the code. It manages the agents who do.
Remy runs the project. The specialists do the work. You work with the PM, not the implementers.
Whether this is a good value depends entirely on the use case. For developers who already burn through cloud API costs running long-context agentic workflows, offloading the bulk of token volume to a $2,600+ local machine could pay for itself. For casual users, it’s a steep price for a capability still getting worked out.
Frequently Asked Questions
What is Microsoft’s hybrid intelligence?
It’s Microsoft’s system, shown at its Windows and Surface event, for automatically routing AI tasks between local, on-device models and cloud models, depending on task difficulty and available hardware.
When does hybrid intelligence launch?
Microsoft said the feature begins rolling out in the GitHub Copilot app on October 15, with local context, actions, and models expanding over the following months.
Does this replace cloud AI models entirely?
No. The demoed workflow used a cloud model for planning complex tasks and local models for the bulk of repetitive, token-heavy work. It’s a division of labor, not a full replacement.
Can I run my own choice of model instead of Microsoft’s?
That’s unclear. Windows ML now supports Llama.cpp, which should make running open and fine-tuned models on Windows easier, but it isn’t confirmed whether the Copilot router itself can be pointed at non-Microsoft models.
How much does the RTX Spark hardware cost?
Pricing reportedly starts around $2,600 and increases with memory configuration, with current memory market pressures likely pushing higher-memory versions well above that starting price.
