Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Kolibri-1 benchmarksAleph Alpha KolibriMoE model comparison

Kolibri-1 Benchmarks: How Aleph Alpha's Model Stacks Up

Kolibri-1's model card compares it to Qwen3.5, GLM-4.7 Flash and other MoE models. Here's what the specs and evaluation setup actually show.

Edited by Luis Chavez-Mattos, Director of Product RSS
Kolibri-1 Benchmarks: How Aleph Alpha's Model Stacks Up

What is Kolibri-1 and how does it compare to other MoE models?

Kolibri-1 is Aleph Alpha’s open-weight mixture-of-experts (MoE) reasoning model, built for German and English and released under Apache 2.0. It has 78 billion total parameters but activates only 3.46 billion per token, putting it in the same active-parameter class as other small-MoE models like Qwen3.5 35B-A3B, GLM-4.7 Flash 30B-A3B, and Nemotron 3 Nano 30B-A3B. Aleph Alpha’s own model card benchmarks Kolibri-1 against these peers, marking the best score among MoE models at each parameter tier while greying out dense models that activate several times more compute per token.

TL;DR

  • Kolibri-1 activates 3.46B parameters out of 78B total, which places it in direct competition with other sub-4B-active MoE models rather than larger dense systems.
  • The model card explicitly separates MoE models from dense models in its evaluation tables, since dense architectures use several times more active compute per token and aren’t a fair apples-to-apples comparison.
  • Named competitors in the 3B-active tier include Qwen3.5 35B-A3B, GLM-4.7 Flash 30B-A3B, and Nemotron 3 Nano 30B-A3B, with additional comparisons against larger MoE and dense models at 4-6B, 12B, 27B, and 70B active-parameter tiers.
  • Kolibri-1 is built specifically for German and English, using a custom tokenizer tuned for German morphology, a deliberate narrowing of scope compared to multilingual models like Qwen or GLM.
  • The model supports explicit reasoning effort levels (low, medium, high, or disabled) and Hermes-style tool calling, both of which factor into its post-training evaluation suite.
  • Training used 20 trillion tokens of bilingual pretraining data plus mid-training and long-context extension stages, run on 768 NVIDIA B200 GPUs over roughly three weeks for the main pretraining phase.
  • The published model card data cuts off mid-table, so some exact benchmark scores (accuracy percentages per eval) aren’t fully available in the source document, even though the comparison structure and competitor list are clear.

Plans first. Then code.

PROJECTYOUR APP
SCREENS12
DB TABLES6
BUILT BYREMY
1280 px · TYP.
yourapp.msagent.ai
A · UI · FRONT END

Remy writes the spec, manages the build, and ships the app.

What models does Kolibri-1 get compared against?

Aleph Alpha organizes its evaluation tables by active-parameter tier, not total parameter count, which is the more meaningful comparison for MoE architectures since only a fraction of total weights fire per token. At the 3B-active tier, the direct competitors are:

  • Qwen3.5 35B-A3B
  • GLM-4.7 Flash 30B-A3B
  • Nemotron 3 Nano 30B-A3B

The table also includes “Kolibri Origin,” which appears to be an internal reference or earlier checkpoint used as a baseline alongside the released model. Beyond the 3B tier, the card lists comparison groups at 4-6B active parameters, 12B active parameters, and separate columns for 27B and 70B dense models (Gemma-class and Llama-class systems, based on the parameter sizes shown). Those dense models are visually greyed out and excluded from the “best in group” marking, since a 70B dense model spending far more compute per token isn’t competing on equal footing with a 3B-active MoE model.

This structure matters for anyone evaluating open models for production use: raw benchmark leaderboards that mix dense and MoE scores without accounting for active parameters can make small MoE models look weaker than they are relative to their actual inference cost.

Why does Aleph Alpha separate MoE and dense models in its benchmarks?

Mixture-of-experts models trade memory for compute. Kolibri-1 needs roughly 78 GB of memory to hold its full weight set (in FP8), but only activates 3.46 billion parameters for any given token, which keeps inference compute and latency low. A dense model with similar benchmark scores might activate 10, 20, or more times the parameters per token, meaning it costs substantially more to serve at the same throughput.

Aleph Alpha’s model card makes this explicit by bolding the best score in each row overall but underlining the best score specifically among MoE models in each parameter group, while greying out dense models entirely from that secondary comparison. This is a reasonable way to benchmark efficiency-focused models: it tells you how Kolibri-1 performs against models you’d actually choose between if active-parameter cost and serving hardware are constraints, rather than against every model that happens to score well regardless of inference cost.

What architecture choices does Kolibri-1 make?

Kolibri-1 is a 50-layer transformer MoE model using a 4:1 ratio of sliding-window attention (SWA) to grouped-query attention (GQA), trained with the Muon optimizer and a technique called Exact Quantile Balancing across 384 experts per layer (1 shared, 6 routed per token). The SWA-heavy design is the mechanism behind its long-context efficiency: most attention layers only look at nearby tokens, while a smaller number of global-attention layers handle long-range context, which keeps the memory and compute cost of long sequences from exploding.

The model was pretrained on sequences of 16,384 tokens, mid-trained up to 65,536, and extended to a native 262,144-token context in a final long-context training phase. Because positional encoding only applies to the sliding-window layers, Aleph Alpha says context can in principle extend further without position scaling tricks, and they validated quality up to 1,048,576 tokens, though they recommend staying at or under 262,144 tokens for latency-sensitive or complex tasks.

Precision-wise, the model ships in FP8 (float8_e4m3fn) weights with dynamically quantized activations and an FP8 KV cache, while embeddings, the LM head, normalization layers, and the MoE router stay in bfloat16. That’s a deliberate choice to shrink memory footprint and speed up serving without degrading the components most sensitive to quantization error.

Is Kolibri-1 worth considering for German-language use cases?

For teams specifically working in German or bilingual German/English contexts, Kolibri-1’s focus is the main differentiator. Aleph Alpha built a tokenizer tailored to German word structure rather than relying on a general multilingual tokenizer, and the pretraining mix itself is heavily bilingual: about 62.5% English, 23.9% German, and 13.6% code across 20 trillion training tokens. That’s a narrower language scope than Qwen3.5 or GLM-4.7 Flash, which are built for broader multilingual coverage, but Aleph Alpha frames this narrowing as intentional: depth in two languages over breadth across many.

Whether that tradeoff is “worth it” depends on the use case. If a deployment is English-only or spans many languages beyond German and English, the narrower focus doesn’t offer the same advantage. If German-language quality and sovereignty (Aleph Alpha is a German AI company and EU GPAI Code of Practice signatory) matter to a project, Kolibri-1 is one of relatively few open-weight MoE models purpose-built around that need.

How was Kolibri-1 trained and what did it cost?

Pretraining ran on 768 NVIDIA B200 GPUs (96 nodes of 8 B200s each) for 21 days, totaling about 392,000 GPU-hours, followed by 5 days of mid-training (90,000 GPU-hours) and about 13 hours of long-context extension (10,000 GPU-hours). Total estimated compute for pretraining alone comes to 6.4×10²³ FLOPs. Aleph Alpha also reports an estimated energy consumption of about 950 megawatt-hours across pretraining, mid-training, and long-context stages, excluding supervised fine-tuning, reinforcement learning, and idle/peak states.

Post-training combined supervised fine-tuning on a bilingual mix of open-source and synthetic data with reinforcement learning across reasoning, agentic, and instruction-following environments, which is where the tool-calling and explicit reasoning-effort capabilities come from.

Frequently Asked Questions

What is Kolibri-1’s active parameter count?

Kolibri-1 has 78 billion total parameters but activates only 3.46 billion per token, placing it in the same efficiency class as Qwen3.5 35B-A3B and GLM-4.7 Flash 30B-A3B.

Which models does Kolibri-1 benchmark against?

Aleph Alpha’s model card compares Kolibri-1 directly against Qwen3.5 35B-A3B, GLM-4.7 Flash 30B-A3B, and Nemotron 3 Nano 30B-A3B in the 3B-active-parameter tier, with additional comparisons against larger MoE models and dense models like Gemma and Llama-class systems at higher parameter tiers.

Why does the benchmark table separate MoE and dense models?

Because dense models activate several times more parameters per token than MoE models of similar total size, comparing raw benchmark scores without accounting for active compute would overstate how expensive it is to match a dense model’s performance with an MoE model, or understate the efficiency of a strong MoE score.

What languages does Kolibri-1 support?

Everyone else built a construction worker.
We built the contractor.

🦺
CODING AGENT
Types the code you tell it to.
One file at a time.
🧠
CONTRACTOR · REMY
Runs the entire build.
UI, API, database, deploy.

Kolibri-1 is built specifically for German and English, using a custom tokenizer optimized for German word structure, rather than targeting broad multilingual coverage.

What context length does Kolibri-1 support?

Its native trained context is 262,144 tokens after a long-context training phase, but Aleph Alpha validated quality and serving efficiency up to 1,048,576 tokens, while recommending 262,144 or fewer for latency-sensitive or complex tasks.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.