Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Limit 1B ViolettoParadigma AI modelsmall math model

Limit 1B Violetto: Testing Paradigma's Tiny Math Model Locally

A local test of Limit 1B Violetto, Paradigma's 1B math specialist, covering install steps, VRAM use, and its tendency to drift off-topic.

Edited by Luis Chavez-Mattos, Director of Product RSS
Limit 1B Violetto: Testing Paradigma's Tiny Math Model Locally

What is Limit 1B Violetto?

Limit 1B Violetto is a 1 billion parameter dense transformer built by Paradigma, trained from scratch on under 300 billion curated tokens with a 131k context window. It is the company’s first released model, and it is designed to do one thing: solve hard math problems. It is only lightly instruction tuned, which means it doesn’t behave like a general chat assistant. It works best when given a single difficult math problem per turn, and according to Paradigma’s own documentation, it can take a vague or unrelated prompt and reframe it as a math problem instead of answering what was actually asked.

TL;DR

  • Limit 1B Violetto is a 1 billion parameter math-specialist model from Paradigma, trained on under 300 billion tokens with a 131k context window.
  • The model is lightly instruction tuned rather than chat tuned, so it performs best with one hard math prompt per turn instead of open-ended conversation.
  • In local testing on an RTX A6000, the model loaded using just under 2GB of VRAM, making it runnable on far less hardware than most reasoning models.
  • On competition math benchmarks like AIME 2026 and HMMT 2026, Paradigma’s own charts show it matching models 25 to 30 times larger, though it falls behind on older 2025 contests and clearly behind on research-style benchmarks like ArchiveMath.
  • Hands-on testing confirmed a strong scope drift problem: asked about photosynthesis or to write a haiku, the model answered a math-flavored version of the question instead, in one case inventing its own definition of what a haiku is.
  • The weights and the VLLM plugin used to serve the model are released under Apache 2.0, making the whole stack open to inspect and modify.
  • Paradigma’s benchmark comparisons are self-reported, with the company noting that training compute is estimated, RL is excluded, and some competitor scores come from model cards or MathArena rather than Paradigma’s own reruns.

Remy doesn't build the plumbing. It inherits it.

Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something.

200+
AI MODELS
GPT · Claude · Gemini · Llama
✓
1,000+
INTEGRATIONS
Slack · Stripe · Notion · HubSpot
✓
MANAGED DB
AUTH
PAYMENTS
CRONS

Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want.

How do you install and run Limit 1B Violetto locally?

The setup differs from the usual “pull a GGUF and run it” workflow because Paradigma ships a custom VLLM plugin alongside the model. The general process:

  1. Clone Paradigma’s GitHub repo and open it locally.
  2. Create a new virtual environment.
  3. Install the Limit plugin, VLLM, PyTorch, and the required transformers version from the project’s lock file.
  4. Start the VLLM server, which downloads the model weights on first run.

The model files themselves are lightweight. On Hugging Face, the repository (paradigma-inc/limite-1b-violetto) contains a single safetensors checkpoint along with a custom modeling file, a chat template, and configuration files, consistent with a compact, purpose-built model rather than a general foundation model with multiple checkpoint shards.

Testing was done on Ubuntu with a single Nvidia RTX A6000 (48GB VRAM), though the hardware demands turned out to be minimal. Once served, the model consumed just under 2GB of VRAM, a fraction of what the A6000 offers and well within reach of consumer GPUs.

How well does it actually perform on math problems?

In direct testing, the model handled a simple algebra problem correctly and quickly. For a harder, competition-style problem (pulled from Paradigma’s own blog, with a known answer of 1080), the model spent a long time in its visible chain-of-thought before arriving at the correct answer. The reasoning trace showed it identifying that only pairs where a equals b, with a coprime to 2025, satisfied the problem, then using Euler’s totient function to count them, a legitimate and fairly sophisticated approach for a model this small.

That result lines up with Paradigma’s published benchmark table, which compares Limit 1B against models ranging from 1 billion to 753 billion parameters across seven math benchmarks. On the newest competition tests, specifically AIME 2026, HMMT 2026, and a benchmark referred to as “beyond AIME,” the 1 billion parameter model reportedly sits at or near the top of the pack, ahead of every other small model on the list and competitive with models 25 to 30 times its size. It doesn’t catch the very largest models, like GLM 5.2, on the hardest problems, which is expected given the parameter gap.

The picture is less flattering elsewhere. On older 2025 contest problems, the model slips behind a handful of smaller rivals. On ArchiveMath, a benchmark built from research-style mathematics rather than competition puzzles, it falls clearly behind. The pattern suggests Paradigma optimized specifically for competition-style problem solving rather than general mathematical reasoning, and that specialization shows up in the gaps.

Plans first. Then code.

PROJECTYOUR APP
SCREENS12
DB TABLES6
BUILT BYREMY
1280 px · TYP.
yourapp.msagent.ai
A · UI · FRONT END

Remy writes the spec, manages the build, and ships the app.

It’s worth repeating that these are Paradigma’s own numbers. The company’s footnotes acknowledge that training compute figures are estimated, reinforcement learning stages are excluded from some comparisons, and the pipeline stages counted differ across models. Several competitor scores in the table are marked as pulled from model cards or MathArena rather than rerun independently by Paradigma. None of this invalidates the results, but it means the compute-efficiency claims (a chart plotting MMM 2026 score against estimated training compute showing Limit 1B reaching similar accuracy to models trained with one to three orders of magnitude more compute) should be read as the vendor’s framing rather than a verified third-party benchmark.

Why does the model struggle outside of math?

This is where Limit 1B Violetto’s specialist design becomes obvious. Paradigma’s own documentation warns that the model may not perform general assistant tasks, and testing confirmed exactly that.

Asked a non-math question about photosynthesis, the model didn’t answer it. Instead, it described cell division and energy transfer, drifted into unrelated territory, and even invented a fabricated term at the end of its answer. Asked to write a haiku about the sea, the model never produced an actual haiku. Instead, its reasoning trace showed it inventing its own definition of a haiku as a “star-shaped writing” structure, then delivering a boxed description based on that invented definition rather than a poem.

This isn’t random failure. It’s a direct consequence of how the model was built. Because Limit 1B is only lightly instruction tuned and trained overwhelmingly on math-focused data, it appears to interpret most inputs through a mathematical lens regardless of what’s actually being asked. The model’s internal reasoning in the haiku test even showed signs of its bundled chat template nudging it toward step-by-step reasoning, a structure suited to math problems but poorly matched to creative or factual requests.

Is Limit 1B Violetto worth using?

For anyone specifically looking for a lightweight, locally hostable model to throw hard math and competition-style problems at, it’s a genuinely interesting option. A 1 billion parameter model that lands near the top of AIME 2026 and HMMT 2026 comparisons, while running in under 2GB of VRAM, is a notable efficiency result, even accounting for the vendor’s own caveats about how those compute comparisons were calculated.

It is not a general-purpose assistant, and treating it like one will produce odd results. It won’t reliably answer science questions, write creative content, or hold a normal conversation. Paradigma seems to have made that tradeoff deliberately, and the Apache 2.0 license on both the weights and the VLLM plugin means developers can inspect, modify, or fine-tune it for narrower use cases if the out-of-the-box scope limitations are a dealbreaker.

Frequently Asked Questions

What is Limit 1B Violetto used for?

It’s designed specifically for solving advanced and competition-level math problems, not for general chat, writing, or factual Q&A.

How much VRAM does Limit 1B Violetto need?

In local testing on an RTX A6000, the loaded model used just under 2GB of VRAM, making it runnable on modest consumer hardware.

Is Limit 1B Violetto open source?

Yes. Both the model weights and the VLLM plugin used to serve it are released under the Apache 2.0 license.

Why does the model answer non-math questions incorrectly?

Because it’s only lightly instruction tuned and trained mostly on math data, it tends to reinterpret non-math prompts (like questions about photosynthesis or requests for a poem) through a mathematical or step-by-step reasoning lens, often producing irrelevant or invented answers.

How does it compare to larger math models?

REMY IS NOT
  • ✕a coding agent
  • ✕no-code
  • ✕vibe coding
  • ✕a faster Cursor
IT IS
✓a general contractor for software

The one that tells the coding agents what to build.

On newer competition benchmarks like AIME 2026 and HMMT 2026, Paradigma’s own data shows it performing close to models 25 to 30 times larger, though it falls behind on older contests and research-style math benchmarks like ArchiveMath.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.