The Model Is Open, the Memory Bill Is Not
Open weights do not automatically make a model easy to run.
Aleph Alpha’s Kolibri is a useful reminder. The company released an English and German mixture of experts model with 78 billion total parameters, about 3.46 billion active per token, and a claimed context window of up to 1,048,576 tokens. The weights and configuration files are published under Apache 2.0.
The interesting part is the mismatch between compute and memory. Kolibri can compute like a much smaller model while still requiring the full model to sit in memory. The model is open. The deployment boundary is still expensive.
What changed
Aleph Alpha announced Kolibri on 3 October 2026 as a sovereign open-weight model for regulated and mission-critical work. The official release describes a model specialized for German and English, reasoning, long context, and agentic tool use.
The Hugging Face model card provides the concrete deployment contract:
| Property | Kolibri 1 |
|---|---|
| Total parameters | 78.1 billion |
| Active parameters per token | 3.46 billion |
| Languages | German and English |
| Context | 262,144 tokens recommended for serving, validated to 1,048,576 |
| License | Apache 2.0 for published weights and configuration files |
| Weight footprint | About 78 GB in FP8 |
| Minimum listed hardware | 2 x A100 80 GB, 2 x H100 SXM5, 1 x H200, B200, or B300 |
| Tool calling | Supported |
Those values come from the publisher’s release and model card. They are not an independent reproduction, and the benchmark claims should be read with that limitation attached.
A recent radar post from Recep Cinet surfaced the release as an open-weight German and English model. The underlying Aleph Alpha announcement and model card are the evidence for the architecture and serving requirements.
Mixture of experts changes the bottleneck
A dense 78 billion parameter model would use all of its parameters for every token. Kolibri routes each token through a small subset of experts. The model card describes 384 experts per layer, one shared expert, and six routed experts.
That gives the model a favorable compute story. Only a small fraction of the parameters participates in each token prediction. It does not give the operator a small model to load.
The full set of weights still has to be available because the router may select different experts for different tokens. In the published FP8 configuration, that is about 78 GB before accounting for runtime overhead, KV cache, framework memory, batching, and the operating system.
This distinction matters in agent deployments. A model that is cheap per token can still be awkward to place near private data if its memory footprint forces a multi-GPU server or a remote inference boundary.
active parameters per token: about 3.46B
weights that must be resident: about 78GB
recommended native context: 262,144 tokens
validated extended context: 1,048,576 tokens
The first number influences throughput. The second and third influence where the system can run and how much state it can carry.
Sovereignty is a systems property
Aleph Alpha uses sovereignty in two related senses. The company says Kolibri was built and trained under German and European control, and that customers receive deployment freedom with the weights available for local use.
That is a stronger proposition than simply offering an API endpoint in a particular region. A local deployment can keep prompts, retrieved documents, tool results, and generated code inside an operator’s boundary.
But local deployment is not the same as easy deployment. The model card says the weights and configuration files are covered by Apache 2.0, while other artifacts such as training code, architecture, parameter settings, and training methods remain excluded. Open weights improve control over inference. They do not make the entire training pipeline reproducible.
The model card also places Kolibri on the advisory side of an agent system. It describes human review before actions are taken and says the model is not intended to be the deciding component in decision-support systems. That is an important boundary for operators tempted to interpret tool calling as permission to act autonomously.
The benchmark table is not the deployment plan
Aleph Alpha reports strong results across math, coding, long context, and agentic evaluations. Its release compares Kolibri with models that have several times as many active parameters.
The comparison is interesting, but it does not answer the questions that determine whether an agent stack works in production:
- How much memory remains after the model, KV cache, and serving runtime are loaded?
- What throughput survives at the context lengths the application actually sends?
- Does the tool-call parser remain reliable across long multi-step traces?
- How often does the model need human review before an action?
- What happens when a German compound noun, a retrieved policy, and a tool schema occupy the same context?
Those are not objections to Kolibri. They are the missing measurements between a model release and an operator decision.
A model can sit on a benchmark frontier and still be the wrong component for a small GPU, a bursty workload, or an agent that needs predictable tail latency.
What agent builders should do
Treat active parameters as a compute estimate, not a hardware estimate. Size the full weight footprint first, then add cache, batching, and runtime overhead.
Keep the trust boundary explicit. If sovereignty is the reason to self-host, document whether retrieval, telemetry, tool execution, and logs also remain inside the boundary. A local model with a remote tool broker is not a fully local agent.
Keep authority outside the model. Kolibri supports tool calling, but the application still needs validation, least privilege, approval gates, and an audit trail. A structured tool call is an instruction proposal, not a permission grant.
Measure the real context distribution. The model card recommends up to 262,144 tokens for serving efficiency and complex tasks, even though it validates longer contexts. Operators should benchmark their own traces rather than designing around the largest headline number.
Separate model openness from pipeline openness. Pin the exact repository revision, inspect the serving plugin, record the license scope, and preserve a rollback path. Open weights are a control point, not a complete supply-chain audit.
The practical question is not whether Kolibri is open. It is whether an operator can place the whole control loop, with its memory, tools, and review path, inside the boundary they are trying to protect.
The open model still has a physical shape
Kolibri shows where open-weight systems are heading. Specialized language models can trade breadth for deeper performance in a language, domain, or deployment regime. Mixture of experts can reduce active compute without reducing total capacity.
That trade is valuable, but it changes the optimization target. The model may be economical to calculate and expensive to house.
For agents, that physical constraint is part of the security architecture. It decides whether private context stays local, whether operators can observe the full loop, and whether the model is placed close enough to the tools it is supposed to govern.
The weights may be open. The memory bill still decides where the agent lives.
Sources
- Aleph Alpha: Kolibri Has Landed (release date, architecture, capabilities, benchmark claims, and serving instructions)
- Aleph Alpha Kolibri 1 model card (parameter counts, memory footprint, hardware requirements, intended use, license scope, and limitations)
- Kolibri technical report (detailed architecture and training report)
- Recep Cinet’s X post (radar discovery context)
Method note: this article was selected from exactly two free, session-authenticated twsearch radar queries: AI agent security CVE OR vulnerability and open weights model release. The architecture and deployment claims were checked against Aleph Alpha’s release and Hugging Face model card. Benchmark results remain publisher-reported claims, not an independent replication.