Aleph Alpha released Kolibri 1 (Kolibri) — a bilingual English–German mixture-of-experts Transformer with roughly 78.1B total parameters and about 3.46B active per token — under Apache 2.0 on Hugging Face as Aleph-Alpha/Kolibri-1. The company frames it as a sovereign open-weight model for regulated work (public administration, industrials, aerospace): trained on infrastructure in Germany and Finland, with native long context, controllable reasoning effort, tool calling, and an explicit grounding/abstention story. Same-day Hacker News reaction was loud — roughly 532 points and 303 comments on the announcement thread.
Vendor-reported benches look strong on math and code. 9to6AI has not hands-on tested Kolibri 1; treat those scores as Aleph Alpha’s own harness results, not independent verification.
Image credit: Aleph Alpha
Key points
- What shipped: Kolibri 1 — open-weight English–German MoE reasoning model; Apache 2.0 weights on Hugging Face (
Aleph-Alpha/Kolibri-1). - Size: ~78.1B total params; ~3.46B active per token; 384 experts, 6 routed + 1 shared; 50 MoE layers.
- Context: Native up to 262,144 tokens; quality/serving validated to 1,048,576; Aleph Alpha recommends ≤262k for efficiency and complex tasks.
- Languages: Bilingual by design. Blog: ~21.3% of pre-training tokens are German. HF card notes ~23.9% German in the filtered corpus mix — different measurement, same bilingual intent.
- Knowledge cutoff: English and German — June 2026 (June 2026 per model card).
- Controls: Reasoning effort
none/low/medium/high; Hermes-style tool calling via thekolibri1parser. - Serve path:
aleph-alpha-inference(vLLM plugin) or containerghcr.io/aleph-alpha/aleph-alpha-inference. - Hardware (vendor): ~78 GB FP8 footprint; minimum examples include 2×A100 80GB, 2×H100, or 1×H200 (and similar).
- Benches (vendor-reported, high effort): e.g. AIME 2025 EN 96.9, GPQA Diamond EN 84.3, LiveCodeBench v6 85.9, SWE-Bench Verified 66.4.
- Grounding: Abstention training + Merlin-Arthur protocol — model taught to say it doesn’t know when context doesn’t support an answer.
- Caveat: 9to6AI has not hands-on tested; compare against your own German RAG, agent, and long-doc workloads.
What shipped
Kolibri 1 is the public release of Aleph Alpha’s latest MoE after an internal predecessor (Kolibri Origin: ~30.6B total / ~3.27B active, shorter context, not publicly released). The public card and blog put the following on the table for builders:
- Product: Kolibri 1 (Kolibri)
- Org: Aleph Alpha
- Architecture: MoE Transformer, 50 layers
- Total / active params: ~78.1B / ~3.46B per token
- Experts: 384 total; 6 routed + 1 shared
- Languages: English, German
- License: Apache 2.0
- Weights: Hugging Face
Aleph-Alpha/Kolibri-1 - Native context: 262,144 tokens
- Extended context: Validated to 1,048,576
- Reasoning effort: none, low, medium, high
- Tool calling: Yes (Hermes-style,
kolibri1parser) - Inference:
aleph-alpha-inference/ vLLM plugin;ghcr.io/aleph-alpha/aleph-alpha-inference - Approx. VRAM (FP8): ~78 GB
- Knowledge cutoff: EN/DE June 2026
Official announcement: aleph-alpha.com — Kolibri has landed. Model card: huggingface.co/Aleph-Alpha/Kolibri-1. Full detail: tech report PDF.
What changed
Open weights, permissive license. Apache 2.0 plus a full HF upload is the practical ship for teams that need to inspect, fine-tune, or air-gap — not another gated API-only drop.
Sparse MoE aimed at serving cost. Activating ~3.5B of ~78B per token is Aleph Alpha’s bet on quality vs decode cost. Memory still holds the full expert set (~78 GB FP8), so this is “cheap per token,” not “fits on a laptop.”
Long context without pretending every length is free. Native training to 256k, extrapolation validated to 1M, with an explicit recommendation to stay at ≤262k for efficiency and hard tasks. That honesty matters more than a headline “1M context” alone.
Bilingual by design, not English-plus-a-bit-of-German. The blog puts German at ~21.3% of pre-training tokens (organic web, rephrasing, limited translation). The HF card’s ~23.9% German share describes the filtered corpus mix — cite both carefully if you compare data cards. Tokenizer work (UniBPE) targets German morphology/compounds while keeping English efficient.
Controllable reasoning + tools. Four effort levels let you trade latency/cost for depth. Tool calling ships with a dedicated parser path in their vLLM plugin — agentic RAG and function-calling stacks are in scope, not afterthoughts.
Grounding as a product feature. Merlin-Arthur (Merlin vs Morgana contexts) trains abstention when evidence is missing. For regulated RAG, “I don’t know” is often more valuable than a fluent guess. Again: vendor evaluation — verify on your documents.
Sovereignty framing. Built and trained under European/German law on Germany/Finland infrastructure; positioned for public sector and industrials that care about supply-chain story, on-prem freedom, and EU AI Act / GDPR alignment. Marketing language — useful if that is your buyer; irrelevant if you only care about open MoE quality.
What it means for builders
If you need strong German + English on-prem. Kolibri is aimed at bilingual orgs that refuse to send internal docs to a US-hosted API. Start with your own German public-sector, legal, or industrial RAG suite — not only EN leaderboards.
If you are MoE-shopping on ~3B active. Compare against other sparse open models at similar active size on your mix of math, code, tool use, and German. Vendor tables put Kolibri near or above several larger-active peers on selected math/code rows — treat that as a hypothesis to falsify.
If long documents are the job. Prototype at 64k–256k first. Use the 1M path only when you have measured latency, KV cost, and quality; Aleph Alpha itself recommends ≤262k for complex work.
If agents and tools matter. Wire Hermes-style tools through aleph-alpha-inference / vLLM with the kolibri1 reasoning and tool parsers. Check multi-turn tool recovery on your harness — BFCL and τ-bench style numbers in the blog are still vendor-run.
If hardware budget is real. Plan for ~78 GB FP8 weights plus KV. The published minimums (2×A100 80GB, 2×H100, 1×H200, etc.) are a floor for getting it running, not a throughput guarantee at 256k.
If compliance is the buyer, not the engineer. Document training geography, license, and abstention behavior for procurement — then still run red-team and grounding evals. Sovereignty claims do not replace application-layer filters.
What to watch
- Independent (non-vendor) benches on German and bilingual agentic RAG.
- Real on-prem throughput at 128k–256k with the official inference stack.
- Whether Merlin-Arthur abstention holds up on messy enterprise corpora, not curated suites.
- Ecosystem packaging beyond Aleph Alpha’s vLLM plugin (llama.cpp, other runtimes, managed hosts).
- Follow-on sizes or vertical post-trains from the same Model Factory pipeline.



