Kolibri AI is Kolibri-1, a free open-weights language model from Aleph Alpha in Heidelberg, Germany, released on 3 October 2026 under the Apache 2.0 licence. It's a mixture-of-experts model with 78.1B total parameters but only 3.46B active per token, it speaks German and English, and it was built with European compliance in mind. You'll also see people spell it "Colibri". Same model.
📺 Watch: Germany's NEW AI Model is INSANE (FREE!)
🔥 Get the Agent OS as a free bonus: AI Profit Boardroom members get the full Agent OS zip, prompt libraries, daily tutorials and weekly live coaching calls. → Get inside
"Kolibri" is German for hummingbird. Small active footprint, fast wings. That's the pitch.
I covered it in my video "Germany's NEW AI Model is INSANE (FREE!)" on 5 October 2026. I read through the benchmarks and the tech report on camera. I did not run it locally myself, and I'll be straight with you about why further down. This guide covers what it is, the specs, why it matters, the benchmarks Aleph Alpha reports, the honest hardware reality, how to get it, and how it fits with Hermes Agent.
If you want hands-on Hermes and AI agent courses plus my Agent OS setup, that's inside the AI Profit Boardroom for $69/mo.
What Is Kolibri AI?
Kolibri-1 is a large language model trained from scratch by Aleph Alpha. It isn't a fine-tune of someone else's model. Per Aleph Alpha's model card, it was trained on 768 NVIDIA B200 GPUs in Germany and Finland.
It's a mixture-of-experts model. That means the full model is big, but for each token it only switches on a small slice of itself. Per the model card, it has 384 routed experts plus 1 shared expert, and it uses 6 routed experts per token. So you get the knowledge of a 78B model with the per-token compute of something closer to 3.5B.
The weights are open, and the Apache 2.0 licence is about as permissive as it gets. You can download it, use it commercially and build on it.
Kolibri-1 Specs
Here are the numbers, all taken from Aleph Alpha's model card and launch post.
| Spec | Kolibri-1 |
|---|---|
| Developer | Aleph Alpha (Heidelberg, Germany) |
| Released | 3 October 2026 |
| Licence | Apache 2.0, open weights |
| Architecture | Mixture-of-experts, 50 layers |
| Total parameters | 78.1B |
| Active parameters | 3.46B per token |
| Experts | 384 routed + 1 shared, 6 routed per token |
| Context | 262,144 tokens native, extended to 1,048,576 (1M) |
| Languages | German and English only |
| Training data | 20T pre-training, 3.44T mid-training, 200B long-context (about 24T total) |
| Knowledge cutoff | 18 June 2026 |
| Reasoning effort | none / low / medium / high |
| Tool calling | Hermes-style |
| Weights | FP8, about 78 GB |
One detail on context. It goes up to 1M tokens, but Aleph Alpha recommends staying at or under 262K for complex tasks. So treat 262K as the working limit and 1M as the stretch.
On languages, per the model card, German made up about 21% of pre-training and English about 62%. It's not a 100-language model. It's built to be really good in two.
Why Kolibri Matters
Three reasons.
1. It's European
Most top open models come from the US or China. Kolibri-1 was trained from scratch in Germany and Finland by a German company. For European businesses that care where their AI comes from, that's a big deal.
2. German and English, properly
If you work in German, most models treat it as one language among many. Kolibri-1 was built around German and English specifically. For German-language content, support or documents, that focus could matter.
3. Compliance was designed in
Per Aleph Alpha, it was built with the EU AI Act, the GPAI Code of Practice and GDPR in mind. If your legal team asks hard questions about AI, a model designed around those rules is an easier conversation.
Benchmarks (Per Aleph Alpha)
These are the scores Aleph Alpha reports on the model card. I haven't verified them myself.
- GPQA Diamond: 84.3%
- MMLU-Pro: 80%
- AIME 2026: 96%
- SWE-Bench Verified: 66.4%
Aleph Alpha compares it against Qwen 3.6-35B, Nemotron 3 Super 120B, Mistral Small 4 119B, Gemma 4 26B, Qwen3-Next 80B and gpt-oss. Their claim is that it matches models with up to four times its active parameters on maths and coding.
Benchmarks are a starting point, not the final word. What matters is how a model does on your actual work. That's why I run models through Goldie Bench, which tests them on real tasks rather than just reading the leaderboard.
The Honest Hardware Reality
This is where I need to be straight with you.
Kolibri-1 is free to download. It is not a normal-laptop model.
Only 3.46B parameters are active per token, but all 78.1B still have to sit in memory. Per the model card, the weights are FP8 and about 78 GB. Aleph Alpha lists the minimum hardware as:
- 2x A100 80GB, or
- 2x H100, or
- 1x H200, B200 or B300
That's data-centre hardware. The most common pushback on my video was basically "78B needs around 100GB, nobody has that laptop." And yeah, that's fair. Most people watching aren't running that at home.
For context, I run a Mac Studio and I don't run much local AI. I didn't run Kolibri-1 locally for the video.
If you really want to try it on smaller hardware, check Hugging Face for community quantised builds that work in tools like LM Studio. Those aren't official, so test them yourself before you trust the output.
How to Get Kolibri AI
The official route:
- Go to the model card on Hugging Face: huggingface.co/Aleph-Alpha/Kolibri-1
- Download the weights (about 78 GB in FP8).
- Serve it with vLLM using Aleph Alpha's inference plugin. That's the official serving setup.
- Point your tools at the OpenAI-compatible API it exposes.
Because the API is OpenAI-compatible, most agent tools that let you set a custom endpoint should be able to talk to it.
Using Kolibri With Hermes Agent
Here's the part I find interesting. Per the model card, Kolibri-1 uses Hermes-style tool calling. That makes it a natural fit as the brain for Hermes Agent.
The idea is simple: if you have the hardware, serve Kolibri-1 with vLLM, then connect Hermes Agent to that OpenAI-compatible endpoint. You get a free, open-weights agent brain that you control, with a long context window and reasoning effort you can turn up or down.
For a broader look at which models work best behind Hermes, read my guide to the best Hermes Agent LLM.
Want help deciding whether a self-hosted model like this makes sense for your business? Book a free AI strategy session and we'll work it out.
My Take on Kolibri AI
Kolibri-1 is impressive on paper. A European model, trained from scratch, open weights under Apache 2.0, strong reported scores, and Hermes-style tool calling built in. For European companies with real GPUs and compliance needs, it's worth a serious look.
Who's it for? Here's how I'd think about it:
- Good fit: European teams with GPU servers, German-language work, strict compliance requirements, and anyone who wants an open-weights agent brain they fully control.
- Poor fit: anyone hoping to run it on a normal laptop, or anyone who needs languages beyond German and English.
For most of you, though, it's not going to run on your machine. That's just the reality of a 78 GB model.
So what do I actually use locally? My best local Hermes model so far on my Mac Studio is LFM 2.5 2.6B. It's very fast, because it was trained on Hermes Agent. It's tiny compared to Kolibri-1, but it actually runs on hardware people own.
If you want to go the local route, start with my guide to the best Ollama model for Hermes Agent, then follow the steps in how to set up Hermes with Ollama.
And if you want to see how I tie models, agents and profiles together into one system, read the Agent OS guide. Or join the AI Profit Boardroom for the full Hermes and AI agent courses plus Agent OS, $69/mo.
FAQ
What is Kolibri AI?
Kolibri AI is Kolibri-1, an open-weights mixture-of-experts language model from Aleph Alpha in Germany, released on 3 October 2026. It has 78.1B total parameters, 3.46B active per token, a 262K native context window, and it supports German and English. Some people spell it "Colibri".
Is Kolibri free?
Yes. The weights are free to download from Hugging Face under the Apache 2.0 licence, which allows commercial use. The catch is hardware. Running it needs serious GPUs, and those aren't free.
Can I run Kolibri on my laptop?
Realistically, no. Per Aleph Alpha, the FP8 weights are about 78 GB and the minimum is 2x A100 80GB, 2x H100, or 1x H200/B200/B300. Community quantised builds may show up on Hugging Face for tools like LM Studio, but they're unofficial, so test them yourself.
Does Kolibri work with Hermes Agent?
It should be a good match. Kolibri-1 uses Hermes-style tool calling, and its official vLLM setup exposes an OpenAI-compatible API you can point Hermes Agent at. You'll need the hardware to serve it first.
Not sure which model or setup fits what you're building? Grab a free AI strategy session and we'll map it out together.











