Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Kolibri 1 benchmarkGPT-OSS comparisonGLM-4.5 Air

Kolibri 1 Benchmarks: How It Stacks Up Against GPT-OSS and GLM

Kolibri 1's MoE design and bilingual training compared against GPT-OSS, GLM-4.7 Flash, Nemotron, and Qwen on size and architecture.

Edited by Luis Chavez-Mattos, Director of Product RSS
Kolibri 1 Benchmarks: How It Stacks Up Against GPT-OSS and GLM

What is Kolibri 1 and why does it matter for benchmarking?

Kolibri 1 is a mixture-of-experts (MoE) reasoning model from Aleph Alpha, built specifically for German and English. It has 78 billion total parameters but activates only 3.46 billion per token, putting it in direct competition with other small-active-parameter MoE models like GLM-4.7 Flash (30B-A3B), Nemotron 3 Nano (30B-A3B), and Qwen3.5 (35B-A3B). The comparison matters because these models all chase the same goal: near-large-model capability at a fraction of the inference cost, with Kolibri adding a specific bet on bilingual depth over broad multilingual coverage.

TL;DR

  • Kolibri 1 activates 3.46B of its 78B total parameters per token, placing it in the same active-parameter class as GLM-4.7 Flash, Nemotron 3 Nano, and Qwen3.5, all of which also use roughly 3B active parameters despite having larger total parameter counts.
  • The model’s evaluation tables separate MoE models from dense models (like 27B and 70B dense checkpoints), because dense models activate several times more compute per token and aren’t a fair apples-to-apples comparison on efficiency.
  • Aleph Alpha built a custom tokenizer tailored to German word structure, a design choice aimed at making German text processing efficient without hurting English performance, which matters for evals run across both languages.
  • Kolibri supports a native 262,144 token context after a dedicated long-context training phase, with validated quality up to just over 1 million tokens, which is relevant to any benchmark involving long-document or retrieval tasks.
  • The architecture uses 4:1 sliding-window to full attention (SWA:GQA), a hybrid pattern that keeps long-context inference affordable while only a minority of layers attend across the entire sequence.
  • Training used 384 experts per layer with 1 shared and 6 routed experts, a comparatively fine-grained MoE setup that differs from how some competing models split expert counts.
  • Kolibri ships with explicit reasoning effort levels (low, medium, high, or none) and Hermes-style tool calling, both configurable through the chat template, which affects how it performs on reasoning-heavy benchmark categories.

Everyone else built a construction worker.
We built the contractor.

🦺
CODING AGENT
Types the code you tell it to.
One file at a time.
🧠
CONTRACTOR · REMY
Runs the entire build.
UI, API, database, deploy.

How does Kolibri 1’s architecture compare to GPT-OSS, GLM, and Qwen?

Kolibri 1 uses a 50-layer transformer MoE design with 384 experts per layer, of which 1 is shared and 6 are routed per token. Its active parameter count, 3.46B, lands it squarely among the current wave of efficient MoE models rather than against heavyweight dense models. GLM-4.7 Flash runs at 30B total parameters with roughly 3B active (30B-A3B), Nemotron 3 Nano follows the same 30B-A3B pattern, and Qwen3.5 scales up to 35B total with roughly 3B active (35B-A3B). Kolibri’s 78B total parameter count is larger than all three, meaning it holds more capacity in memory while activating a similar compute budget per token.

This is the core trade-off MoE architectures make: more total parameters means a larger memory footprint (Kolibri’s FP8 weights alone need about 78 GB), but the per-token compute cost stays close to a much smaller dense model. Aleph Alpha’s own evaluation tables explicitly group MoE models together and mark dense models (27B and 70B checkpoints) separately, noting that dense models activate several times as many parameters per token and therefore aren’t a fair comparison point on efficiency, even if their raw task scores are competitive.

Why does Kolibri 1 separate English and German evals?

Kolibri was trained on a bilingual corpus: roughly 62.5% English, 23.9% German, and 13.6% code, out of a 20 trillion token pre-training set. That’s a deliberate narrowing of scope. Most competing open models (GPT-OSS, GLM, Qwen) are trained as broadly multilingual systems, covering dozens of languages with English dominant and other languages, including German, as secondary coverage.

Aleph Alpha’s bet is that depth in two languages beats breadth across many. Part of that bet is a custom tokenizer designed around German word structure (German’s long compound words and inflections tokenize inefficiently under tokenizers built primarily for English or Chinese). A tokenizer mismatch shows up directly in benchmark efficiency: the same German sentence can cost meaningfully more tokens on a generic tokenizer, which affects both inference cost and, in some cases, effective context usage. Evaluating English and German separately lets the benchmark tables show whether this specialization actually produces better scores in German, rather than just better token economics.

How does the mixture-of-experts design affect benchmark results?

The number and structure of experts shapes how a model handles different task types. Kolibri’s 384-expert, 6-routed-plus-1-shared-expert setup is more fine-grained than some competitors’ expert counts, which in theory lets the router specialize more narrowly for a given input, at the cost of more complex routing behavior during training and inference.

Kolibri also uses a 4:1 ratio of sliding-window attention to full (grouped-query) attention layers. Most layers only look at nearby tokens, while a smaller fraction attend across the entire context. This keeps inference cheap even at long context lengths, but it also means benchmark tasks that require connecting distant pieces of information rely on a minority of the model’s layers to do that work. For benchmarks that stress long-context retrieval or multi-document reasoning, this architectural choice is as relevant as raw parameter counts.

Does context length give Kolibri 1 an edge on long-document benchmarks?

Kolibri was pre-trained on 16,384-token sequences, mid-trained on 65,536 tokens, and then passed through a dedicated long-context training phase using 262,144-token sequences, which Aleph Alpha calls its native context length. Because positional encoding is only applied in the sliding-window layers, the model can in principle be extended to arbitrary context lengths without position scaling tricks. Aleph Alpha says it has validated quality and serving efficiency up to 1,048,576 tokens (just over a million), though it recommends staying at or below 262,144 tokens for latency-sensitive or complex-task deployments.

This matters for any benchmark involving retrieval-augmented generation or long-document processing, categories Kolibri is explicitly positioned for. A model’s benchmark score on a long-context needle-in-haystack test or multi-document QA task depends heavily on whether the underlying attention mechanism was actually trained and tuned at that length, not just technically capable of accepting that many tokens.

Is Kolibri 1 worth considering over GPT-OSS or GLM for bilingual use cases?

For teams working primarily in German and English, Kolibri 1’s narrower language focus and dedicated tokenizer are a real differentiator, something broadly multilingual models like GPT-OSS, GLM-4.7 Flash, or Qwen3.5 don’t optimize for in the same way. Its MoE efficiency profile (3.46B active out of 78B total) puts it in a reasonable cost bracket for self-hosting, provided you can meet its hardware floor: Aleph Alpha lists a minimum of two A100 80GB or two H100 SXM5 GPUs, with one H200, B200, or B300 as single-card alternatives.

The trade-off is breadth. If a workload needs strong performance across many languages beyond German and English, a broadly multilingual model is still the safer default. Kolibri’s Apache 2.0 license and open weights make it straightforward to test directly against GPT-OSS, GLM, Nemotron, or Qwen on your own bilingual workload rather than relying solely on published benchmark tables.

Frequently Asked Questions

What does “active parameters” mean and why does it matter for comparing these models?

Active parameters refers to how many of a model’s total parameters are actually used to process each token, which is what drives inference cost and speed. Kolibri 1 has 78B total parameters but only activates 3.46B per token, similar to GLM-4.7 Flash and Nemotron 3 Nano’s roughly 3B active parameter designs, so comparing active parameters rather than total parameters gives a more accurate picture of real-world serving cost.

Is Kolibri 1 a dense model or a mixture-of-experts model?

Kolibri 1 is a mixture-of-experts model, using 384 experts per layer across 50 layers, with 1 shared and 6 routed experts activated per token. This is distinct from dense models like 27B or 70B checkpoints, which activate all or most of their parameters for every token and therefore cost more compute per token despite sometimes having fewer total parameters.

What languages does Kolibri 1 support?

One coffee. One working app.

You bring the idea. Remy manages the project.

WHILE YOU WERE AWAY
✓Designed the data model
✓Picked an auth scheme — sessions + RBAC
✓Wired up Stripe checkout
✓Deployed to production
Live at yourapp.msagent.ai

Kolibri 1 is built specifically for German and English, trained on a corpus that’s roughly 62.5% English, 23.9% German, and 13.6% code. This is narrower than fully multilingual models like Qwen3.5 or GLM, which is a deliberate design choice Aleph Alpha describes as prioritizing depth over broad language coverage.

How long of a context can Kolibri 1 handle?

Kolibri’s native context length after long-context training is 262,144 tokens, and Aleph Alpha recommends staying within that range for latency-sensitive or complex tasks. The model has been validated for quality and serving efficiency up to 1,048,576 tokens, and its architecture allows extension beyond that in principle since positional encoding is confined to its sliding-window attention layers.

What hardware is needed to run Kolibri 1?

Kolibri 1’s FP8 weights take up roughly 78 GB of memory. Aleph Alpha lists a minimum requirement of two A100 80GB or two H100 SXM5 GPUs, or a single H200, B200, or B300 card, with two H100 SXM5 or two H200 GPUs recommended for better performance.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.