Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
RTX Spark priceSurface Laptop Ultra specsNvidia RTX Spark

RTX Spark and Surface Laptop Ultra: Pricing and Specs Explained

Microsoft's RTX Spark-powered Surface laptops start around $2,600, with up to 128GB unified memory for running local AI models.

Edited by Luis Chavez-Mattos, Director of Product RSS
RTX Spark and Surface Laptop Ultra: Pricing and Specs Explained

What is the RTX Spark and why does it matter for local AI?

The RTX Spark is Nvidia’s small-form-factor AI computer built around a GPU roughly equivalent to an RTX 5070, paired with a CPU and unified memory architecture designed to run large language models directly on the device. Microsoft showed it off as the engine behind new Surface hardware, including the Surface Laptop Ultra, at a Windows and Surface event in San Francisco. The pitch is simple: instead of sending every AI request to the cloud, your laptop can run serious models itself, for free, with no API bill and no internet connection required.

That’s not a hypothetical. In one demo Microsoft showed a coding agent process 1.6 million tokens during a single session investigating GitHub issues on the Windows Terminal codebase. The cost was zero, because the bulk of that work ran locally on the machine rather than through a cloud API.

TL;DR

  • The RTX Spark pairs an Nvidia GPU (roughly RTX 5070 class, around 20 CPU cores, under 80 watts) with unified memory configurations that reportedly range from 24GB up to 128GB.
  • Surface devices built on RTX Spark start at around $2,600, with pricing climbing as memory capacity increases, partly due to ongoing RAM and memory cost pressures across the industry.
  • Microsoft’s “hybrid intelligence” router in GitHub Copilot can now send simple tasks to a local model and harder tasks to a cloud model like GPT-5, splitting the work automatically rather than leaving the choice to the developer.
  • In a demo task, 1.6 million tokens went in and roughly 10,000 came out, a ratio near 150 to 1, almost all of it processed locally by a small on-device model while a cloud model handled the planning.
  • Microsoft is shipping three local models: a coding model called MAI Code 1.1 Flash at 3-bit quantization, an upcoming Nitron-style model at 2-bit for leaner hardware, and a heavily compressed version of DeepSeek at around 1.6-bit average precision.
  • Llama.cpp is being built into Windows ML, Microsoft’s cross-hardware runtime, which means open GGUF model files should work on day one without waiting for official Microsoft packaging.
  • A new sandboxing layer called Microsoft Execution Containers (MXC) lets organizations set policies on what local agents can access, with every action attributed to the agent rather than the logged-in user.

Remy doesn't build the plumbing. It inherits it.

Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something.

200+
AI MODELS
GPT · Claude · Gemini · Llama
✓
1,000+
INTEGRATIONS
Slack · Stripe · Notion · HubSpot
✓
MANAGED DB
AUTH
PAYMENTS
CRONS

Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want.

How much does the RTX Spark hardware actually cost?

Pricing for the RTX Spark powered Surface laptops starts around $2,600. That’s the entry point for the base configuration, and the price rises from there as you add memory. Microsoft and Nvidia are offering multiple memory tiers, with the high end reaching 128GB of unified memory, and a smaller 64GB tier has already appeared in Nvidia’s own DGX Spark lineup. Reports also point to configurations as low as 24GB aimed at lighter workloads.

The climbing price with more memory isn’t really a Microsoft or Nvidia choice. Memory and RAM costs have been rising industry-wide, and unified memory architectures like this one are especially sensitive to that, since memory capacity is the main lever for how large a model you can actually load and run. If you want to run bigger, higher-quality models locally, you pay for the RAM to hold them, and right now that RAM is expensive.

Beyond the laptop, Microsoft also announced a desktop tier built on the same chip, and teased a Windows version of Nvidia’s DGX Station arriving later this year, aimed at giving more people access to this class of local compute.

How does Microsoft’s local-cloud routing actually work?

The feature doing the real work here is what Microsoft calls “intelligent routing,” built into GitHub Copilot’s model picker. When set to “auto,” the picker no longer chooses only between cloud models. It can now also route a task to a model running on your own device.

In a live demo, a simple cleanup task on the Windows Terminal codebase got sent to Microsoft’s local coding model, MAI Code 1.1 Flash, and the presenter showed the GPU usage spike on screen. She then put the entire machine into airplane mode, and the task kept running, proof that it wasn’t touching the network at all.

A harder task told a different story. When asked to work through a batch of GitHub issues tagged “needs repro,” the system used a cloud model (reportedly one of OpenAI’s GPT-5 family) to do the planning, then dispatched three sub-agents to execute that plan locally using the on-device coding model. That’s where the 1.6 million token count came from: the local sub-agents did almost all of the reading and grinding, while only about 10,000 tokens round-tripped to the cloud. That’s a roughly 150-to-1 split favoring local compute, and it’s a pattern likely to show up more broadly in agentic workflows, where context windows balloon with reading and tool output but the actual generated response stays small.

What’s unclear, and wasn’t demonstrated, is what happens when the router guesses wrong. If a task is too hard for the local model, does the system detect that and escalate to the cloud, or does it just return a worse answer with no signal that anything went wrong? That failure-handling question matters a lot for anyone planning to rely on this kind of auto-routing for real work.

Which models actually run locally, and are they any good?

REMY IS NOT
  • ✕a coding agent
  • ✕no-code
  • ✕vibe coding
  • ✕a faster Cursor
IT IS
✓a general contractor for software

The one that tells the coding agents what to build.

Microsoft named three local models at the event, each pushed to fairly aggressive quantization levels to fit on-device:

MAI Code 1.1 Flash, Microsoft’s coding model first introduced at Build, now compressed to 3-bit, cutting its size by roughly 80%, and running with a 256k context window fully on-device. Whether a 3-bit quantized model holds up for genuinely difficult coding tasks is an open question. Aggressive quantization tends to degrade exactly the kind of precision coding work demands.

A new Nitron-style model running at 2-bit precision in about 20GB of memory, explicitly positioned as a good fit for lower-memory RTX Spark configurations.

A compressed version of DeepSeek running at an average of 1.6-bit precision in around 60GB of memory. Microsoft’s on-stage claim was that a year ago, this level of capability only existed in frontier cloud models. The “1.6-bit average” detail matters: it implies mixed precision, where important layers stay higher-precision and less-active layers get pushed lower, rather than a flat 1.6-bit cut across the whole model.

The real test for any of these is threefold: does coding accuracy hold up against the full-precision version, does performance degrade at long context lengths, and do tool calls stay clean under agentic use. Structured output is often the first casualty of heavy quantization, and a broken tool call usually means a failed task in any agent pipeline. There’s also a memory cost to long context itself: a 256k context window requires a sizable KV cache on top of the model weights, which likely pushes serious users toward the higher-memory configurations of this hardware rather than the entry-level tier.

Why does Llama.cpp support inside Windows ML matter?

Windows ML is Microsoft’s runtime for deploying models across GPUs, NPUs, and CPUs. Folding Llama.cpp into it is a practical move: it means open models in GGUF format should work the moment they’re released, without waiting for Microsoft to officially package them. That includes fine-tuned, uncensored, or niche community variants that Microsoft would otherwise never convert or support itself.

It also simplifies life for developers building Windows apps with embedded local models, since they can target one runtime that adapts to whatever hardware the end user has, rather than writing separate code paths for different GPU and NPU combinations. One open question is whether third-party models can be plugged directly into Copilot’s router, or whether routing stays limited to Microsoft’s own picks. If any local model can be wired into that router, it would meaningfully expand what “local AI on Windows” means in practice.

What is Microsoft Execution Containers (MXC)?

As local models gain more autonomy to take real actions, Microsoft introduced MXC, a runtime sandbox comparable in spirit to Nvidia’s and Docker’s container-based isolation approaches. Organizations can define policies for what an agent is allowed to access, and Windows enforces those policies while the agent runs.

Cursor
ChatGPT
Figma
Linear
GitHub
Vercel
Supabase
goremy.ai

Seven tools to build an app. Or just Remy.

Editor, preview, AI agents, deploy — all in one tab. Nothing to install.

The notable detail is agent identity: every action an agent takes gets attributed to that agent specifically, not to the logged-in user. That separation matters for auditing and for control. If an agent starts doing something it shouldn’t, it can be blocked and lose access to that action, while the human user’s own permissions stay untouched.

Frequently Asked Questions

How much does the RTX Spark cost?

Surface laptops built on the RTX Spark chip start at around $2,600 for the base configuration. Price increases with memory capacity, with configurations going up to 128GB of unified memory.

What GPU is inside the RTX Spark?

Reported specs put the GPU at roughly RTX 5070 performance level, paired with around 20 CPU cores, and a power draw under 80 watts.

Can the RTX Spark run large models like DeepSeek locally?

Microsoft demonstrated a heavily quantized version of DeepSeek running at roughly 1.6-bit average precision in about 60GB of memory. How well that compressed version performs against the full-precision model on real coding and reasoning tasks hasn’t been independently verified yet.

Does Windows now support open-source AI models natively?

Yes, through Llama.cpp being integrated into Windows ML, Microsoft’s runtime for running models across different hardware. This should let GGUF-format open models run shortly after release without needing official Microsoft packaging.

What happens when a local AI model can’t handle a task?

Microsoft didn’t demonstrate this failure case. It’s unclear whether the system automatically detects when a local model is struggling and escalates to a cloud model, or whether it simply returns a weaker answer without flagging the issue.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.