Kolibri-1: Aleph Alpha's Open-Weight German-English MoE Model
Kolibri-1 is Aleph Alpha's 78B MoE model with 3.5B active params, 1M context, and Apache 2.0 license, built for German and English.

What is Kolibri-1?
Kolibri-1 is an open-weight mixture-of-experts (MoE) language model released by the German company Aleph Alpha on October 3, 2026. It has 78 billion total parameters but activates only about 3.46 billion per token, a design that keeps inference cheap while retaining the capacity of a much larger network. The model focuses specifically on German and English, supports context windows up to roughly 1 million tokens, and ships under the Apache 2.0 license.
TL;DR
- Kolibri-1 is a 78B-parameter MoE model from Aleph Alpha with only 3.46B active parameters per token, making it cheap to run relative to its total size.
- The model supports a context window up to 1,048,576 tokens, though Aleph Alpha recommends staying at or below 262,144 tokens for serving efficiency on complex tasks.
- It was trained on 20 trillion tokens of bilingual German-English-code data, plus 3.44T tokens of mid-training and 201B tokens dedicated to long-context extension.
- Kolibri-1 grew out of an earlier, never-released model called Kolibri Origin, which Aleph Alpha built first to validate its training pipeline before scaling up.
- In about three months, the architecture scaled from 30B to 78B total parameters while active parameters per token barely moved, from 3.27B to 3.46B.
- The model runs on FP8 weights and needs at least two A100 80GB or H100 GPUs, with four-figure VRAM usage (around 140GB) reported in real-world testing.
- It is released under Apache 2.0, supports explicit reasoning modes and tool calling, and Aleph Alpha is a signatory of the EU’s GPAI Code of Practice.
Remy doesn't build the plumbing. It inherits it.
Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something.
Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want.
How does Kolibri-1’s architecture work?
Kolibri-1 is a 50-layer transformer built on a mixture-of-experts design with 384 experts per layer, of which one is shared and six are routed for any given token. Only a small slice of the full parameter set fires for each token, which is the core trick behind MoE models: you get the knowledge capacity of a dense 78B model without paying the compute cost of running all 78B parameters every time.
Attention uses a 4:1 ratio of sliding-window attention to grouped-query attention. Most layers only look at a local window of nearby text, while a minority of layers attend across the entire context. This split is what makes long-context inference affordable. Because positional encoding is applied only in the sliding-window layers, Aleph Alpha says the context can in principle extend indefinitely without needing position-scaling tricks. The company validated quality and serving efficiency up to 1,048,576 tokens, though it recommends capping at 262,144 tokens for latency-sensitive or complex workloads.
Weights are stored in FP8 (float8_e4m3fn) with 128x128 blocks and dynamically quantized activations, while embeddings, the LM head, normalization layers, and the MoE router stay in bfloat16. The model also includes an FP8 KV cache. Training used Muon as the optimizer and a technique called Exact Quantile Balancing to keep the mixture of experts stable.
Why does the Kolibri Origin backstory matter?
Before Kolibri-1, Aleph Alpha built an internal-only model called Kolibri Origin, used purely to shake out bugs in the training pipeline. The company never released it publicly, but the jump from Origin to Kolibri-1, which happened over roughly three months, is a useful data point on how fast a lab can scale a validated recipe.
Total parameters nearly tripled, from about 30 billion to 78 billion. Active parameters per token, by contrast, moved only slightly, from 3.27 billion to 3.46 billion. That’s the entire point of the MoE architecture: Aleph Alpha could make the model dramatically more capable without making it dramatically more expensive to serve. Alongside that, training data scaled from 7.5 trillion tokens to 20 trillion, trained context length jumped from 64,000 tokens to 262,144, and the number of experts per layer tripled from 128 to 384. The knowledge cutoff also moved forward, with Aleph Alpha listing June 18, 2026 for both English and German implicit knowledge.
Aleph Alpha has also published a diagram describing its training methodology as iterative rather than one-shot. Both pre-training and post-training follow the same loop: pick a data recipe and architecture (or reward scheme, for post-training), train a small development model, evaluate it, and repeat that cycle many times before committing to the expensive full-scale run. The final large training run, in this framing, is the last step in a long chain of cheap experiments rather than a single bet.
What are the hardware and deployment requirements?
Remy is new. The platform isn't.
Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.
Kolibri-1’s FP8 weights have a memory footprint of about 78GB, but running the model in practice requires more headroom than that for context, batching, and the KV cache. Aleph Alpha’s own minimum spec is two A100 80GB GPUs, two H100 SXM5 GPUs, one H200, one B200, or one B300. The recommended setup bumps to two H100 SXM5s, two H200s, one B200, or one B300. Independent testing on a dual-GPU local rig reported around 140GB of VRAM in use.
Kolibri-1 requires the aleph-alpha-inference package, which provides a vLLM plugin, and Aleph Alpha distributes a container image as well. Serving the model with reasoning and tool-calling enabled looks like a standard vLLM launch command with added flags for an FP8 KV cache, a kolibri1 reasoning parser, and automatic tool-call parsing. Pushing past the default 262,144-token context to the full 1,048,576-token range requires explicit overrides for max model length and position embeddings. Aleph Alpha recommends sampling with temperature=1.0, top_p=0.97, and top_k=128.
How does Kolibri-1 perform on real tasks?
Aleph Alpha’s own benchmark tables group Kolibri-1 against other MoE models with similar active-parameter counts, roughly in the 3B active-parameter class, including GLM-4.7 Flash 30B-A3B, Nemotron 3 Nano 30B-A3B, and Qwen3.5 35B-A3B. Dense models with several times more active parameters per token are shown for reference but not treated as direct competitors, since they operate on fundamentally different compute economics.
Independent hands-on testing (conducted by a YouTube creator covering the release) used coding and instruction-following stress tests rather than standard benchmark suites. In one test, the model was asked to write a Blender Python script to construct a complete indoor trampoline park scene, including a trampoline grid, foam pit, dodgeball court, and obstacle course, and to export it, a task that reportedly took 40 to 60 minutes. In another, it built a playable third-person Godot game in a single pass with no external assets. A third test asked for a single self-contained HTML file using plain canvas and JavaScript to render a cinematic animated kebab scene with flames, smoke, embers, and reflections, a prompt designed to check whether the model could track dozens of simultaneous instructions without dropping any. The tester reported the model delivered nearly all the specified elements correctly.
A separate German-language test built around a five-part philosophical prompt (translation, explanation, constrained rewriting, and a comparison of what gets lost between versions) found the model’s German output notably stronger and more idiomatic than its constrained English output, with the final “what was lost in translation” step identified as the weakest link.
Is Kolibri-1 worth using?
For teams that need strong German-language performance with Apache 2.0 licensing and no API dependency, Kolibri-1 is a reasonable candidate, provided you have the GPU budget. The MoE design means inference cost stays close to a 3-4B active-parameter model even though the full checkpoint needs about 78GB of memory just to load. That tradeoff favors teams with existing multi-GPU infrastructure over anyone looking to run a sovereign model on consumer hardware.
The long-context support (up to 1 million tokens, with a practical recommendation of 262,144) and built-in tool-calling and reasoning-effort controls make it suitable for agentic workflows and retrieval-augmented generation, particularly in bilingual German-English settings. Whether it’s “worth it” compared to similarly sized MoE models like GLM-4.7 Flash or Qwen3.5 depends heavily on whether German-language quality is a priority, since that’s the explicit design focus Aleph Alpha built the model around rather than broad multilingual coverage.
Frequently Asked Questions
Seven tools to build an app. Or just Remy.
Editor, preview, AI agents, deploy — all in one tab. Nothing to install.
What does Kolibri-1 stand for and who makes it?
Kolibri is the German word for hummingbird. It’s built by Aleph Alpha Research GmbH and released by Aleph Alpha GmbH, a German AI company that has been developing language models for several years.
How many parameters does Kolibri-1 actually use per token?
Kolibri-1 has 78 billion total parameters but activates only about 3.46 billion per token through its mixture-of-experts routing, which uses one shared expert and six routed experts out of 384 per layer.
What hardware do I need to run Kolibri-1?
Aleph Alpha lists a minimum of two A100 80GB or two H100 SXM5 GPUs (or a single H200, B200, or B300), with two H100s, H200s, or a single B200/B300 recommended. Real-world testing reported around 140GB of VRAM usage on a dual-GPU setup.
Is Kolibri-1 free to use commercially?
Yes. The model is released under the Apache 2.0 license, which permits commercial use, modification, and redistribution.
How does Kolibri-1 compare to Kolibri Origin?
Kolibri Origin was an internal, unreleased model Aleph Alpha built to test its training pipeline. Over about three months, the team scaled the architecture from roughly 30 billion to 78 billion total parameters, trained on nearly three times more data (7.5 trillion to 20 trillion tokens), and tripled the number of experts, all while keeping active parameters per token almost unchanged.





