Pi 1.0's Code Mode: How Native MCP and Model Routing Actually Work
Pi 1.0 adds native MCP via code mode, custom model routing, deferred tool loading, and Pi Durable for long-running agent sessions.

What is Pi’s code mode and why does it matter for MCP?
Code mode is Pi’s answer to a problem that has annoyed anyone using the Model Context Protocol: tool descriptions and raw results flood the model’s context before it has done any real work. Instead of dumping every MCP tool definition and every intermediate response in front of the model, Pi 1.0 lets the agent write a small JavaScript program that calls tools, combines results, and returns only the useful output. The model never has to read every field of every API response just to filter a list down to five relevant items. That filtering happens inside the script, and only the compact result comes back into the conversation.
TL;DR
- Pi 1.0 adds native MCP support through code mode, letting agents write JavaScript that orchestrates tool calls instead of flooding context with raw tool output.
- A custom model router can be registered as a selectable model, choosing a different physical model (and thinking level) per request based on your own policy logic.
- Deferred tool loading lets extension authors mark tools as discoverable-on-demand rather than loaded upfront, keeping context lean as you add more integrations.
- Cache warming can refresh an eligible Anthropic prompt cache during long tool runs, but only when Pi estimates the avoided cache-miss cost clears a 5-cent threshold.
- Code mode also opens up specialist models like a Jev classifier and image generators, callable from scripts without appearing in the normal chat model picker.
- Pi Durable is a separate, experimental package for long-running, resumable multi-conversation agent applications with forkable state and typed JSON documents.
- The final 1.0 release reports a roughly 40% prompt token reduction in one specific GPT-5.6 test case (about 5,300 tokens down to 3,300), not a blanket cost cut across all tasks.
How does code mode change tool orchestration?
Pi previously avoided native MCP support on principle. The team’s criticism was that typical MCP integrations shove tool descriptions and large text results straight into the model’s context window, which is an expensive way to get a model from “I need some data” to “here’s the answer.” Code mode fixes this by wrapping MCP calls in a JavaScript sandbox. The agent writes a script that calls one or more tools, often in parallel, groups or filters the results, and returns a condensed report. If you wanted a summary of recent activity on a project, a script could pull issues, pull commits, and pull pull requests concurrently, then combine them into one short answer instead of pasting three raw API responses into the chat.
The sandbox has real boundaries worth understanding. The JavaScript itself runs isolated, but the tools it calls keep their own permissions: a shell tool can still edit files, an external integration can still take an action against a live system. If a script fails partway through, any tool calls that already completed are not rolled back automatically. Code mode also includes explicit storage helpers for JSON data that needs to persist across calls, since ordinary JavaScript variables don’t survive between separate executions. Error handling is still the developer’s job, not something code mode solves for free.
For setup, Pi supports MCP servers running locally over standard input/output and remote servers over HTTP, configurable at the user or project level. Servers using code mode exposure (the default) have their tools discovered on demand rather than placed directly in front of the model. You can opt individual tools into direct exposure if you want the model to see them immediately, and OAuth is supported for servers that require account sign-in.
What can you do with Pi’s custom model routing?
One of the more practical additions in Pi 1.0 is the ability to register a virtual model that acts as a router. Instead of manually switching models mid-task, an extension can register one selectable “model” that, before each request, picks the actual physical model that should handle it. A simple policy might route difficult planning steps to a stronger, more expensive model and hand routine implementation work to something cheaper. The routing logic can also inspect conversation state to choose an appropriate reasoning or thinking level.
Pi keeps the real underlying model and its usage visible in the session, so you can see exactly what the router chose and how much it cost, even though the decision logic lives entirely in your extension. The launch demo showed an extension that planned with Claude Opus and implemented with GPT, using the Jev classifier to decide when to switch, then showed the handoff happening from Opus to a GPT variant mid-session along with the resulting cost and cache numbers.
- ✕a coding agent
- ✕no-code
- ✕vibe coding
- ✕a faster Cursor
The one that tells the coding agents what to build.
Routing isn’t free, though. Switching physical models mid-session can break prompt cache reuse, and running a classifier to make the routing decision adds latency to every request. A router only pays off if the end-to-end task still finishes correctly and the total cost (including the routing overhead) actually comes in lower than just using one solid model the whole way through. Testing a router against a single-model baseline on real tasks, rather than assuming the cheaper path wins by default, is the sensible way to validate it.
What are deferred tools and why do they help?
As you connect more MCP servers and extensions, you accumulate more tool descriptions and argument schemas, and loading all of them upfront burns context before the agent has even started on your actual request. Pi 1.0 introduces a three-way distinction for tools: those declared directly to the model, those reached only through code mode, and deferred tools that get discovered and activated through search when needed.
Some tools can also be restricted to direct model calls only, which makes sense for operations that coordinate other tools or need user interaction in the loop. In practice, this means a small local code edit doesn’t need to carry detailed instructions for every unrelated integration you happen to have installed. The tradeoff is that discovery quality depends on how well tools are described, since the agent has to find the right one through search rather than seeing it listed explicitly. Testing discovery with ordinary, naturally-phrased requests (not just prompts that name the tool outright) is a reasonable way to check this works before relying on it.
Is Pi’s new caching and theming worth paying attention to?
Cache warming addresses a narrow but real pain point for Anthropic-based workflows with heavy context: a long tool run can outlast a prompt cache’s lifetime, forcing an expensive cache miss on the next turn. Pi can issue a refresh to keep an eligible cache warm, but only when it estimates the refresh will save at least 5 cents in avoided cache-miss cost. That estimate requires knowing the cache lifetime for the model in question. Warming can run during active sessions by default, with an option to also warm during idle periods between runs, or be turned off entirely. The refresh itself uses resources and shows up in session accounting, even though it stays out of the model’s actual conversation context. Checking session details (through the /session command) before deciding whether idle warming fits a given workload is worth doing, since someone checking in every few minutes has very different needs from someone who leaves a session open all weekend.
On the interface side, Pi 1.0 ships a new default terminal theme called “system” that reads the terminal’s own foreground, background, and ANSI colors and builds its interface around that palette, adjusting contrast for readability and re-querying if the terminal switches between light and dark. Full-screen mode is also now the default, keeping the prompt editor and status area fixed while conversation scrolls within the window; the older scroll-back behavior is still available as “regular” mode through settings or a startup flag.
What is Pi Durable and who is it for?
Seven tools to build an app. Or just Remy.
Editor, preview, AI agents, deploy — all in one tab. Nothing to install.
Pi Durable is a separate, experimental package built for applications that need agent conversations to survive restarts, not just single coding sessions. A single harness can run multiple conversations at once, each with its own model, tools, instructions, and working directory. Conversations can be forked at a specific point and continued independently, and multiple clients can connect to the same harness to follow progress or add input. It also supports typed JSON application data that commits alongside transcript updates, which is useful for building something like a shared investigation workspace where several people track the same underlying task through separate threads.
Durability comes from checkpointing: a process can stop and a new process can reopen the same storage and resume unfinished work. But replay is conservative by design. A tool interrupted mid-execution is only automatically replayed if it explicitly declares that replay is safe; otherwise the agent just gets told what was interrupted and what output exists so far. That distinction matters a lot in practice, since retrying a read-only search and retrying an action that already changed an external system are very different risks. Request IDs prevent a retried submission from being double-counted, but that doesn’t guarantee every external side effect happens exactly once.
Storage options include memory, SQLite, and JSONL, with memory storage not surviving a restart at all. The default SQLite setup doesn’t guarantee the newest commits survive a host or power failure, and only one process can own a storage backend at a time. The APIs are explicitly experimental, and model context limits still apply: Durable’s compaction example demonstrates summarization and recovery from context overflow, but it doesn’t grant a model unlimited memory.
Frequently Asked Questions
What is Pi’s code mode in simple terms?
It’s a JavaScript sandbox where the agent writes small programs to call MCP tools, combine their results, and return a compact answer, instead of pasting raw tool output directly into the model’s context.
Does Pi’s custom model routing save money automatically?
Not automatically. A router can send cheaper requests to cheaper models, but switching models can break prompt caching and adds routing overhead, so it needs to be tested against a single-model baseline on real tasks before assuming it’s cheaper overall.
Is Pi Durable ready for production use?
No. It’s described as an experimental foundation, the APIs are expected to change, default SQLite storage doesn’t guarantee durability through a power failure, and only one process can own a storage backend at a time.
Does deferred tool loading reduce token usage for every task?
It reduces upfront context usage in setups with many installed tools, but the actual savings depend on the task and setup. The one reported figure (about 40% fewer prompt tokens in a specific GPT-5.6 test) came from a particular configuration, not a general guarantee.
Can code mode call models other than the main chat model?
Yes. Code mode can call specialist models like the Jev classifier and image generation models using the session’s existing credentials. These don’t appear in the regular chat model picker and are reached only through code mode or extensions.

