Microsoft FrogNano 4B: A Tiny Coding Model You Can Run Locally
Microsoft's FrogNano 4B is a 4B coding model built on Qwen 3.5, trained with RL. Here's how it runs locally and handles real bugs.

What is Microsoft FrogNano 4B?
FrogNano 4B is a small open-weight coding model from Microsoft, released under the model ID microsoft/FrogNano-4B-2609 on Hugging Face. It is built on top of Qwen 3.5 4B and trained with reinforcement learning rather than distillation from a larger teacher model. Microsoft’s own description puts the training data at around 1,500 synthetic software engineering tasks, paired with a lightweight agent harness called Leaf that restricts the model to just five tools: read, write, edit, glob, and bash. The pitch is simple: a coding agent small enough to fit on one consumer or prosumer GPU, instead of requiring a cluster or an API key for a frontier model.
TL;DR
- FrogNano 4B is a 4 billion parameter coding model from Microsoft, built on Qwen 3.5 4B and released as
microsoft/FrogNano-4B-2609with MIT licensing. - It was trained with reinforcement learning on roughly 1,500 synthetic software engineering tasks rather than copying outputs from a bigger model.
- The model pairs with Leaf, a minimal agent harness that exposes only five tools (read, write, edit, glob, bash), which shapes how it behaves when run in other harnesses.
- In a hands-on test, it correctly diagnosed and fixed a real SQL aggregation bug in a multi-service Docker app (Postgres, FastAPI, Nginx, Redis), though its reasoning path was messy and inefficient.
- It struggled with a single-file HTML/JavaScript app (an animated kebab-cooking demo), suggesting its strength is unevenly distributed across languages.
- It performed well on a Python task, building a full expense tracker with a SQLite backend, a working CLI, and passing unit tests it wrote itself.
- On benchmark charts, FrogNano sits near coding-specialist models roughly seven times its size, though larger specialist models (~30B) and frontier models still outperform it.
Other agents start typing. Remy starts asking.
Scoping, trade-offs, edge cases — the real work. Before a line of code.
How was FrogNano 4B trained?
Microsoft trained FrogNano using reinforcement learning instead of supervised fine-tuning on a bigger model’s outputs. The process, as described by Microsoft and demonstrated in testing, works like a self-generated curriculum: the model creates its own coding tasks, discards the ones it always solves (too easy) and the ones it never solves (too hard), and keeps the tasks where it wins only some of the time. It then trains on that narrow band of tasks with RL, improves, and writes a harder batch of tasks for its new skill level. This loop reportedly repeats five times, so the model is always training at the edge of its own ability rather than on a fixed, pre-built dataset.
This self-play style curriculum is paired with the Leaf harness during training, which only gives the model five tools: read, write, edit, glob, and bash. That constraint matters. A model trained to solve problems with a narrow, predictable toolset will reason differently than one trained with open-ended tool access, which shows up later when you run it in a different agent framework.
How do you run FrogNano 4B locally?
FrogNano 4B is small enough to serve with standard local inference stacks. In testing, it was run using SGLang, with VRAM usage driven largely by the KV cache rather than the model weights themselves. Reducing the context length brings memory requirements down to a level that fits on a single commodity GPU, the kind of card an individual developer or small team might already own rather than a multi-GPU server.
Because Leaf is built with Kubernetes-style deployment in mind, running FrogNano outside that harness means substituting a more accessible agent framework. In testing, the Hermes agent framework was used instead, with tool calling enabled so the model could still read, write, and edit files and run shell commands, mirroring the five-tool setup it was trained on. Microsoft also published a quantized GGUF version of the model, which reduces storage and memory needs further for anyone running it on more constrained hardware.
Does FrogNano 4B actually fix real bugs?
In one test, FrogNano was pointed at a live multi-service web app, a ferry occupancy dashboard built with a Postgres 16 database, a FastAPI backend, an Nginx frontend, and Redis caching, all running through Docker Compose. The bug: the backend counted booking rows instead of summing passenger counts per booking, so a booking covering a family or tour group of six people was counted as a single passenger. The practical effect was that ferries showing as nearly empty were actually near capacity, a real safety and overbooking risk in a live system.
Given a goal to find and fix the bug, FrogNano identified the problem early using a tool-call chain, correctly spotting the row-count-versus-passenger-sum issue. From there its process got messy: it tore down and rebuilt the database multiple times, misattributed the cause at points, and ran a few broken scripts before correcting course. Eventually it fixed the underlying SQL query, and a hard refresh of the dashboard confirmed the fix: occupancy figures that had read as 5% full with 114 seats available corrected to roughly 76.7% full with 28 seats remaining, matching the real passenger counts.
The inefficiency during the process is likely tied to the harness mismatch. FrogNano was trained inside Leaf’s constrained five-tool environment, and running it in a different agent framework changes how cleanly it executes multi-step debugging, even when it eventually reaches the right answer.
Where does FrogNano 4B fall short?
Testing surfaced a clear language skew. Asked to build a single self-contained HTML file simulating a kebab rotating on a gas-powered skewer, the model produced a file with errors that didn’t render anything in the browser. Even after being told the file wasn’t working and attempting a fix, the output still failed to load. This points to FrogNano being considerably stronger in Python than in HTML/JavaScript, a reasonable outcome for a model trained primarily on backend-style software engineering tasks rather than frontend or browser-based rendering problems.
By contrast, a Python task went smoothly: building an expense tracker application with a SQLite backend. The model built a working command-line tool, wrote five unit tests that passed, seeded demo data, and provided the commands needed to use it. In practice, the CLI worked as advertised: initializing the database, adding food and transport expenses, listing all entries, and correctly totaling expenses by category, all without manual debugging.
Is FrogNano 4B worth running over a bigger model?
On Microsoft’s benchmark comparisons, FrogNano sits far to the small end of the parameter-count axis while still outperforming its own Qwen 3.5 4B base model and a larger 9 billion parameter sibling model. On the main coding benchmark referenced, it lands near coding-specialist models around seven times its size. It doesn’t beat dedicated ~30 billion parameter coding specialists or frontier models, which remain well ahead on raw benchmark scores.
The practical case for FrogNano isn’t that it replaces larger models. It’s that a 4B model running on a single commodity GPU can handle real debugging tasks in a production-shaped app (multi-service, containerized, SQL-backed) and build working Python tools with tests, without API costs or cloud dependency. For Python-centric backend work on constrained hardware, that’s a meaningful capability gap closed. For frontend or multi-language work, it’s clearly not there yet.
Frequently Asked Questions
What base model is FrogNano 4B built on?
FrogNano 4B is built on Qwen 3.5 4B, fine-tuned by Microsoft using reinforcement learning rather than supervised distillation.
What is the Leaf harness?
Leaf is Microsoft’s lightweight agent harness for FrogNano, giving the model exactly five tools: read, write, edit, glob, and bash. It’s designed with Kubernetes-based deployment in mind, which is why local testing often substitutes a different agent framework like Hermes.
How much VRAM does FrogNano 4B need?
VRAM usage is driven mainly by the KV cache size, not the 4B parameter weights alone. Reducing the context length brings the model down to a footprint that fits on a single commodity GPU, and a quantized GGUF version is available for even lower memory use.
Is FrogNano 4B good at frontend code?
Remy doesn't write the code. It manages the agents who do.
Remy runs the project. The specialists do the work. You work with the PM, not the implementers.
Testing found it significantly weaker at HTML/JavaScript tasks, failing to produce a working single-file browser demo even after attempting a fix. Its strength appears concentrated in Python and backend-style software engineering.
Can FrogNano 4B fix bugs in real applications?
Yes, in testing it correctly diagnosed and fixed a passenger-count aggregation bug in a multi-service Dockerized app (Postgres, FastAPI, Nginx, Redis), though it took an inefficient, roundabout path before arriving at the correct fix.


