Kolibri 1: Aleph Alpha’s Open-Weight English-German MoE

Aleph Alpha released Kolibri 1 (Kolibri) — a bilingual English–German mixture-of-experts Transformer with roughly 78.1B total parameters and about 3.46B active per token — under Apache 2.0 on Hugging Face as Aleph-Alpha/Kolibri-1. The company frames it as a sovereign open-weight model for regulated work (public administration, industrials, aerospace): trained on infrastructure in Germany and Finland, with native long context, controllable reasoning effort, tool calling, and an explicit grounding/abstention story. Same-day Hacker News reaction was loud — roughly 532 points and 303 comments on the announcement thread.

Vendor-reported benches look strong on math and code. 9to6AI has not hands-on tested Kolibri 1; treat those scores as Aleph Alpha’s own harness results, not independent verification.

Image credit: Aleph Alpha

Key points

  • What shipped: Kolibri 1 — open-weight English–German MoE reasoning model; Apache 2.0 weights on Hugging Face (Aleph-Alpha/Kolibri-1).
  • Size: ~78.1B total params; ~3.46B active per token; 384 experts, 6 routed + 1 shared; 50 MoE layers.
  • Context: Native up to 262,144 tokens; quality/serving validated to 1,048,576; Aleph Alpha recommends ≤262k for efficiency and complex tasks.
  • Languages: Bilingual by design. Blog: ~21.3% of pre-training tokens are German. HF card notes ~23.9% German in the filtered corpus mix — different measurement, same bilingual intent.
  • Knowledge cutoff: English and German — June 2026 (June 2026 per model card).
  • Controls: Reasoning effort none / low / medium / high; Hermes-style tool calling via the kolibri1 parser.
  • Serve path: aleph-alpha-inference (vLLM plugin) or container ghcr.io/aleph-alpha/aleph-alpha-inference.
  • Hardware (vendor): ~78 GB FP8 footprint; minimum examples include 2×A100 80GB, 2×H100, or 1×H200 (and similar).
  • Benches (vendor-reported, high effort): e.g. AIME 2025 EN 96.9, GPQA Diamond EN 84.3, LiveCodeBench v6 85.9, SWE-Bench Verified 66.4.
  • Grounding: Abstention training + Merlin-Arthur protocol — model taught to say it doesn’t know when context doesn’t support an answer.
  • Caveat: 9to6AI has not hands-on tested; compare against your own German RAG, agent, and long-doc workloads.

What shipped

Kolibri 1 is the public release of Aleph Alpha’s latest MoE after an internal predecessor (Kolibri Origin: ~30.6B total / ~3.27B active, shorter context, not publicly released). The public card and blog put the following on the table for builders:

  • Product: Kolibri 1 (Kolibri)
  • Org: Aleph Alpha
  • Architecture: MoE Transformer, 50 layers
  • Total / active params: ~78.1B / ~3.46B per token
  • Experts: 384 total; 6 routed + 1 shared
  • Languages: English, German
  • License: Apache 2.0
  • Weights: Hugging Face Aleph-Alpha/Kolibri-1
  • Native context: 262,144 tokens
  • Extended context: Validated to 1,048,576
  • Reasoning effort: none, low, medium, high
  • Tool calling: Yes (Hermes-style, kolibri1 parser)
  • Inference: aleph-alpha-inference / vLLM plugin; ghcr.io/aleph-alpha/aleph-alpha-inference
  • Approx. VRAM (FP8): ~78 GB
  • Knowledge cutoff: EN/DE June 2026

Official announcement: aleph-alpha.com — Kolibri has landed. Model card: huggingface.co/Aleph-Alpha/Kolibri-1. Full detail: tech report PDF.

What changed

Open weights, permissive license. Apache 2.0 plus a full HF upload is the practical ship for teams that need to inspect, fine-tune, or air-gap — not another gated API-only drop.

Sparse MoE aimed at serving cost. Activating ~3.5B of ~78B per token is Aleph Alpha’s bet on quality vs decode cost. Memory still holds the full expert set (~78 GB FP8), so this is “cheap per token,” not “fits on a laptop.”

Long context without pretending every length is free. Native training to 256k, extrapolation validated to 1M, with an explicit recommendation to stay at ≤262k for efficiency and hard tasks. That honesty matters more than a headline “1M context” alone.

Bilingual by design, not English-plus-a-bit-of-German. The blog puts German at ~21.3% of pre-training tokens (organic web, rephrasing, limited translation). The HF card’s ~23.9% German share describes the filtered corpus mix — cite both carefully if you compare data cards. Tokenizer work (UniBPE) targets German morphology/compounds while keeping English efficient.

Controllable reasoning + tools. Four effort levels let you trade latency/cost for depth. Tool calling ships with a dedicated parser path in their vLLM plugin — agentic RAG and function-calling stacks are in scope, not afterthoughts.

Grounding as a product feature. Merlin-Arthur (Merlin vs Morgana contexts) trains abstention when evidence is missing. For regulated RAG, “I don’t know” is often more valuable than a fluent guess. Again: vendor evaluation — verify on your documents.

Sovereignty framing. Built and trained under European/German law on Germany/Finland infrastructure; positioned for public sector and industrials that care about supply-chain story, on-prem freedom, and EU AI Act / GDPR alignment. Marketing language — useful if that is your buyer; irrelevant if you only care about open MoE quality.

What it means for builders

If you need strong German + English on-prem. Kolibri is aimed at bilingual orgs that refuse to send internal docs to a US-hosted API. Start with your own German public-sector, legal, or industrial RAG suite — not only EN leaderboards.

If you are MoE-shopping on ~3B active. Compare against other sparse open models at similar active size on your mix of math, code, tool use, and German. Vendor tables put Kolibri near or above several larger-active peers on selected math/code rows — treat that as a hypothesis to falsify.

If long documents are the job. Prototype at 64k–256k first. Use the 1M path only when you have measured latency, KV cost, and quality; Aleph Alpha itself recommends ≤262k for complex work.

If agents and tools matter. Wire Hermes-style tools through aleph-alpha-inference / vLLM with the kolibri1 reasoning and tool parsers. Check multi-turn tool recovery on your harness — BFCL and τ-bench style numbers in the blog are still vendor-run.

If hardware budget is real. Plan for ~78 GB FP8 weights plus KV. The published minimums (2×A100 80GB, 2×H100, 1×H200, etc.) are a floor for getting it running, not a throughput guarantee at 256k.

If compliance is the buyer, not the engineer. Document training geography, license, and abstention behavior for procurement — then still run red-team and grounding evals. Sovereignty claims do not replace application-layer filters.

What to watch

  • Independent (non-vendor) benches on German and bilingual agentic RAG.
  • Real on-prem throughput at 128k–256k with the official inference stack.
  • Whether Merlin-Arthur abstention holds up on messy enterprise corpora, not curated suites.
  • Ecosystem packaging beyond Aleph Alpha’s vLLM plugin (llama.cpp, other runtimes, managed hosts).
  • Follow-on sizes or vertical post-trains from the same Model Factory pipeline.

Sources

Leave a Comment

Your email address will not be published. Required fields are marked *