4 October 2026

Aleph Alpha Kolibri-1 is worth testing for self-hosted AI, not migrating to

Aleph Alpha released open weights for a bilingual mixture-of-experts model with context up to one million tokens. Operators should focus instead on its 78-billion-parameter memory footprint and dedicated serving stack.

Virtual Arc · Editorial image

What shipped

Aleph Alpha released Kolibri-1 on October 3 with its weights and configuration under Apache 2.0. The bilingual mixture-of-experts model has 78.1 billion total parameters, activates 3.46 billion per token, and supports controlled reasoning and tool calls.

Its context has been validated up to 1,048,576 tokens, but the vendor recommends no more than 262,144 for efficient serving and complex tasks. Capacity plans should use that recommendation, not the million-token headline.

The operator math

Sparse activation reduces computation per token, but it does not remove the need to keep the full model in memory. The BF16 weights occupy roughly 156 GB, while the official runtime requires Aleph Alpha's vLLM plugin, currently aligned with vLLM 0.29.

This is not a simple endpoint swap. A team takes responsibility for GPU capacity, runtime upgrades, service availability and the cost of idle infrastructure.

What we would do

Virtual Arc would begin with a limited shadow evaluation on real German-English documents, retrieval workloads and tool calls. The existing hosted model would remain the control, with no production traffic moved initially.

Migration is justified only if Kolibri-1 wins on cost per completed task, P95 latency and reliability, or if data control repays the infrastructure premium. Otherwise, we would wait for a more mature runtime and broader hosting support.

Our take

Virtual Arc's view is that Kolibri-1 is important enough to enter a controlled production trial, but not strong enough to trigger an automatic migration from hosted models. The figure of 3.46 billion active parameters suggests inexpensive inference, yet all 78.1 billion weights must remain available in memory; the BF16 release occupies about 156 GB, and the official serving path depends on a dedicated vLLM plugin currently tied to one minor version. The cost has not disappeared. It has moved from a vendor invoice into GPU capacity, on-call ownership, upgrades and idle-hardware risk. We would run a two-week shadow trial only for document-heavy German-English RAG and tool-calling workflows where data control has measurable value. We would cap context at the recommended 262k tokens and track cost per accepted result, P95 latency, valid tool calls, grounded abstention and GPU utilisation. Without a clear win on those measures, we would not migrate.

Sources
  1. Kolibri Has Landed: A Sovereign Open-Weight Model
  2. Aleph Alpha Kolibri-1-BF16 Model Card
  3. Aleph Alpha Inference vLLM Plugin
  4. Heidelberg AI company launches Kolibri language model

← All posts