Microsoft Brings Llama.cpp to Windows: Run DeepSeek Locally on RTX
Microsoft is adding llama.cpp to Windows ML, letting quantized DeepSeek and Microsoft AI Code models run locally on RTX hardware.

Microsoft is building local AI routing directly into Windows, pairing llama.cpp support with heavily quantized versions of DeepSeek and its own coding models so they run on your laptop’s GPU instead of a data center. At its recent Windows and Surface event, Microsoft showed a coding agent sending 1.6 million tokens into a local model during a single session, at zero API cost, because the work happened on-device rather than in the cloud.
TL;DR
- Microsoft demoed a GitHub Copilot session where a coding agent processed 1.6 million tokens locally on a laptop, with only about 10,000 tokens round-tripping to a cloud model for planning.
- The new “auto” model picker in Copilot now routes tasks between local and cloud models automatically, based on task difficulty, rather than developers manually choosing where the work runs.
- Microsoft’s own AI Code 1.1 Flash model runs locally at 3-bit quantization with a 256k context window, cutting its footprint by roughly 80% compared to full precision.
- A DeepSeek V3.1 variant at 1.6-bit average quantization was shown running in about 60GB of memory, a quantization level Microsoft claims still holds up on coding and reasoning benchmarks.
- Windows ML is gaining native llama.cpp support, meaning GGUF-format open models become usable on Windows almost as soon as they’re released, without waiting for Microsoft to repackage them.
- A new sandboxing layer called Microsoft Execution Containers (MXC) lets organizations set policies on what local agents can access, with every action attributed to the agent rather than the logged-in user.
- The hardware behind this, RTX Spark laptops and desktops, starts around $2,600 and scales up to 128GB of unified memory with full CUDA support on Windows, no Linux dual-boot required.
Everyone else built a construction worker.
We built the contractor.
One file at a time.
UI, API, database, deploy.
What did Microsoft actually announce?
At its San Francisco event, Microsoft tied together three things: new RTX-based hardware, a runtime capable of running open models locally, and a routing layer in GitHub Copilot that decides where a given task should execute. The headline demo had a Copilot agent investigating GitHub issues on the Windows Terminal codebase. A simple cleanup task got routed to Microsoft’s own AI Code Flash model running locally, visible as a GPU spike on screen, and the presenter even put the machine into airplane mode to prove it kept working without network access.
A harder task, going through a batch of issues tagged “needs repo,” triggered a different pattern. A cloud model (described as an OpenAI model in the GPT family) handled the planning, then dispatched three sub-agents that ran locally on the Microsoft AI Code model to do the bulk of the reading and investigation. That split is where the 1.6 million token figure comes from: the local sub-agents consumed the overwhelming majority of tokens, while only roughly 10,000 tokens worth of output went back to the cloud model, a ratio around 150 to 1.
This mirrors a pattern already common in agentic workflows: an expensive model does the thinking, cheaper models do the high-volume reading and grunt work. Microsoft is now building that pattern into the operating system itself.
How does the local/cloud routing actually work?
Microsoft calls this “hybrid intelligence,” with the Windows team describing it as intelligent routing between models, a runtime, and the chipset. In practice, the GitHub Copilot model picker has an “auto” setting that now chooses between cloud models and local models depending on the task, not just between different cloud providers.
The system isn’t fully transparent about failure handling. If a local model can’t handle a task, it’s unclear whether the router detects that and escalates to the cloud, or whether the user just gets a worse answer with no visibility into what went wrong. That’s a meaningful gap, since every routing system lives or dies on how it behaves when it guesses wrong. Microsoft didn’t demonstrate that failure path on stage.
The routing feature ships inside the GitHub Copilot app starting October 15, with local context, local actions, and additional local models expected to roll out over the following months.
Which models can you run locally, and how small did they get?
Three models were shown running locally during the event, each quantized far below typical production precision:
- Microsoft AI Code 1.1 Flash, the coding model introduced at Microsoft’s Build conference, running at 3-bit quantization with a 256k token context window, a roughly 80% size reduction versus the original.
- An upcoming Nemotron-based model running at 2-bit quantization in about 20GB of memory, positioned as a fit for lower-memory RTX Spark configurations.
- A DeepSeek V3.1 variant running at an average of 1.6 bits in roughly 60GB of memory, with Microsoft claiming performance on coding and reasoning benchmarks that would have been cloud-only a year ago.
Plans first. Then code.
Remy writes the spec, manages the build, and ships the app.
The 1.6-bit figure is an average rather than a flat rate: important layers get kept at higher precision while less-activated layers get pushed much lower. That’s a standard mixed-precision quantization strategy, but it raises real questions. Does the coding pass rate hold up against a higher-precision version of the same model? Does quality degrade at long context lengths? And do tool calls stay well-formed, since structured output is often the first thing to break under aggressive quantization, and a broken tool call generally means a failed agentic task.
Long context adds its own cost beyond the model weights themselves: the KV cache at 256k tokens takes up substantial memory on top of the quantized weights, which likely pushes practical use toward the higher-memory RTX Spark configurations rather than entry-level ones.
Why does llama.cpp support in Windows ML matter?
Windows ML is Microsoft’s runtime for deploying models across GPUs, NPUs, and CPUs. Bringing llama.cpp into that runtime is arguably the most consequential technical change in the announcement, because it opens up compatibility with the broader open-model ecosystem rather than locking users into Microsoft-curated models.
Concretely, it means that when a new open model is released and a GGUF quantized version shows up (often within hours), Windows users can run it without waiting for Microsoft to package or approve it. That includes fine-tuned, specialized, or uncensored community variants that Microsoft would never officially support. For developers building Windows apps that ship with a local model, it also means one runtime that handles whatever hardware a given user has, rather than writing separate code paths for different GPUs or NPUs.
One open question is whether third-party models can plug into Copilot’s own routing system, or whether that router is restricted to Microsoft’s models. If any local model could be wired into the OS-level router, swapping in something like a mixture-of-experts model as a system-wide default, that would be a significant capability for developers rather than just end users.
Is local execution actually safe to let loose on your machine?
Running agentic coding tools locally raises an obvious question: what stops an agent from doing something it shouldn’t? Microsoft’s answer is Microsoft Execution Containers (MXC), a runtime sandbox comparable in spirit to Nvidia’s OpenShell or Docker-style sandboxing. Organizations can define policies for what an agent is allowed to access, and Windows enforces those policies while the agent runs.
The more interesting detail is agent identity: every action an agent takes is attributed to the agent itself, not to the logged-in user. That makes auditing and tracing agent behavior possible, and it also means a user can revoke an agent’s access to a specific action mid-task while retaining their own access, rather than losing control of the whole session.
What hardware does this run on, and what does it cost?
The local models were demonstrated on RTX Spark hardware, Nvidia GPUs integrated into new Windows laptops and desktops with up to 128GB of unified memory and full CUDA support, no Linux required. Independent coverage has put the GPU performance roughly in line with an RTX 5070, paired with 20 CPU cores and an 80-watt power draw. Pricing starts around $2,600, climbing with added memory. Lower-memory configurations (down to 24GB) and a Windows version of Nvidia’s DGX Station were also teased for later release. Rising memory and component costs mean these prices are likely to drift upward regardless of what Microsoft or Nvidia intend.
Frequently Asked Questions
Can I run DeepSeek locally on a Windows laptop today?
Remy is new. The platform isn't.
Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.
Microsoft demonstrated a 1.6-bit quantized DeepSeek V3.1 variant running locally in about 60GB of memory on RTX Spark hardware. Broader availability depends on the Windows ML and llama.cpp rollout timeline, with GitHub Copilot’s local routing features arriving starting October 15.
What is the difference between 3-bit and 1.6-bit quantization?
Both are compression techniques that reduce model size by storing weights with fewer bits of precision than the original (typically 16-bit) training format. Lower bit counts shrink memory requirements further but risk more quality loss; 1.6-bit setups described here use mixed precision, keeping important layers higher and less-used layers lower, averaging to 1.6 bits overall.
Why does llama.cpp support matter for Windows users?
It lets Windows run GGUF-format open models, the most common format for community-quantized models, without waiting for Microsoft to officially package them. New model releases become usable on Windows almost immediately after a GGUF quant appears.
Does local AI replace the cloud entirely?
No. Microsoft’s demo used a hybrid pattern where a cloud model handled planning and complex reasoning while local models handled high-volume, lower-complexity work like reading through repository issues. The router decides which tasks go where.
What is Microsoft Execution Containers (MXC)?
It’s a runtime sandbox for AI agents running on Windows, letting organizations set access policies that Windows enforces while an agent operates. Actions taken by an agent are attributed to the agent’s own identity rather than the logged-in user’s.

