Hero Background

Power Your Software Testing with AI Agents and Cloud

The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.

AIAgent Testing

Aleph Alpha Kolibri-1: What to Test Before You Self-Host It

Aleph Alpha Kolibri-1 is a 78B open German-English model. See where its benchmarks stop, what they skip in German, and how to test tool calls and abstention.

Published on:

If you plan to put a German-language agent on Kolibri-1, the score you most need is missing. Aleph Alpha's tech report says only four of the eight benchmark categories behind the model's overall scores have German versions, and tool calling and hallucination are not among them.

Aleph Alpha released Kolibri-1 on 3 October 2026, the Day of German Reunification, as an open-weight model with 78B total parameters, 3B active, and a context window of up to 1M tokens under Apache 2.0, according to Aleph Alpha's launch post. It is built for regulated work in public administration, industry, and aerospace, and is small enough to run on-premise without sending internal data to a third-party inference service.

Its headline grounding claim is that it is trained to say "I don't know" when the answer is not in the context. Test that claim, and the model's German tool calling, on your own data before an agent built on Kolibri-1 reaches users.

TL;DR

Aleph Alpha Kolibri-1 is an open-weight German-English mixture-of-experts model with 78.1B total and 3.46B active parameters, a context window of up to 1,048,576 tokens, tool calling, and four reasoning effort levels. It exists so organizations can run German-language assistants and agents on their own hardware under Apache 2.0.

  • Self-hosting: Can you run Kolibri-1 on a single GPU? Yes, on one H200, B200, or B300. Kolibri-1's FP8 weights take about 78 GB, so A100 80 GB and H100 SXM5 cards need a pair. The model is served through vLLM with Aleph Alpha's inference plugin.
  • German tool calling: Does Kolibri-1 support tool calling in German? Yes, but no published benchmark measures it. Only four of the eight benchmark categories behind Kolibri-1's overall scores have German versions, and agentic tool use is not one of them.
  • Multi-turn tool calls: Is Kolibri-1 strong at multi-turn tool calls? No. Kolibri-1 scores 47.5 on the BFCL v4 multi-turn split, below the 58.1 of Qwen3.6 35B-A3B, while scoring 38.1 on Tau3-Bench banking against 10.6 for the same Qwen model.
  • Abstention: Does Kolibri-1 abstain when the evidence is missing? Mostly, yes. Kolibri-1 holds back on 85.6% of RGB questions whose documents lack the answer, but scores 34.0 on RGB's error-correction condition, where the documents contain false information.
  • Pre-deployment testing: Write every Kolibri-1 tool-calling and abstention case in German and in English, and run each one several times at the settings you will ship. TestMu AI Agent Testing scores those conversations for hallucination and custom abstention rules against a self-hosted endpoint.

What Is Aleph Alpha Kolibri-1?

Kolibri-1 is Aleph Alpha's mixture-of-experts reasoning model for German and English. The Kolibri-1 model card lists multi-step reasoning, retrieval-augmented generation, agentic tool calling, coding, and German- and English-language assistants as its best uses.

SpecificationKolibri-1
Total parameters78.1B
Active parameters per token3.46B
Experts384 routed experts plus 1 shared expert per layer, 6 routed experts per token
Attention50 layers: 40 use a 512-token sliding window, 10 attend to the full context
Context length262,144 tokens native, validated up to 1,048,576
Reasoning effortnone, low, medium, or high, set per request
Tool callingHermes-style calls, parsed into an OpenAI-compatible API
Pre-training data20T tokens, about 4.3T of them German
Knowledge cutoff18 June 2026, for both English and German
LicenseApache 2.0, covering the weights and configuration files

Kolibri-1 succeeds Kolibri Origin, a smaller model that Aleph Alpha built to validate its training pipeline and never released. The launch post attributes Kolibri-1's German ability to a bilingual tokenizer that splits German compounds along their word parts, organic German web text curated and rephrased in-house, and sparing use of translation.

On the Hacker News thread for the launch, which passed 600 points, commenters pressed on multi-turn tool calling and on the memory needed to self-host the model.

How Do You Self-Host Kolibri-1?

Kolibri-1 runs on vLLM through Aleph Alpha's aleph-alpha-inference package, which ships the Kolibri plugin and installs the vLLM version it supports; the container image ghcr.io/aleph-alpha/aleph-alpha-inference is the alternative. The model card gives this serve command, which turns on the reasoning and tool-call parsers:

pip install 'aleph-alpha-inference>=1'

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

The server exposes an OpenAI-compatible chat completions API, so existing clients, agent frameworks, and test harnesses can point at it unchanged. The model card and launch post set these limits for planning a deployment:

  • Memory - the FP8 weights take about 78 GB, so the minimum is two A100 80 GB or two H100 SXM5 cards, or a single H200, B200, or B300.
  • Context - 262,144 tokens is the native window and the recommended ceiling for latency-sensitive or complex work; serving up to 1,048,576 tokens needs extra --max-model-len and --hf-overrides flags.
  • Throughput - Aleph Alpha chose 78B over a 123B candidate because, on two H100s, the smaller model serves 18 concurrent 256k-token requests against 3 and decodes 28% faster.
  • Sampling - the recommended settings are temperature 1.0, top_p 0.97, and top_k 128.
  • Reasoning effort - set per request through chat_template_kwargs as none, low, medium, or high, or turned off with enable_thinking set to false.

Each of those settings changes the system you are testing. Aleph Alpha's published scores use reasoning effort high, and at the recommended temperature the same prompt can return different answers, so run every test case several times at the settings you plan to ship.

Where Does Kolibri-1 Lead and Trail on Aleph Alpha's Benchmarks?

Aleph Alpha ran every model through its own evaluation harness with Kolibri-1 at reasoning effort high, so treat these numbers as vendor-reported. Against Qwen3.6 35B-A3B, which also activates about 3B parameters per token, the model card's release table shows a clear split:

BenchmarkWhat it testsKolibri-1Qwen3.6 35B-A3B
AIME 2026 (German)Competition math asked in German90.084.4
GPQA Diamond (German)Graduate-level science questions in German81.380.6
Tau3-Bench (Banking)A multi-step customer-service agent working through tools38.110.6
HoneypotAgentic retrieval asked in German and English80.874.3
BFCL v4 (overall)Function calling across single, multi-turn, memory, and web-search tasks61.467.2
BFCL v4 (multi-turn)Function calls that depend on earlier turns47.558.1
SWE-Bench VerifiedFixing real GitHub issues as a coding agent66.473.8
RGB Negative (Abstention)Holding back when the documents lack the answer85.679.6
RGB Fact-Check (Error Correction)Answering correctly when the documents contain false information34.074.0
RGB Closed-BookAnswering from memory with no documents51.079.0
AA-Omniscience non-hallucination rateAbstaining instead of answering wrong on knowledge questions44.056.7

The tech report names Kolibri-1's weakest rows itself:

  • Answering from memory - RGB Closed-Book, where Kolibri-1 places last of the twelve models compared.
  • Knowledge accuracy - AA-Omniscience accuracy, the share of knowledge questions it answers correctly.
  • Correcting false documents - RGB Fact-Check, the error-correction condition.
  • Multi-turn tool calls - the multi-turn splits of the Berkeley Function Calling Leaderboard (BFCL).

Its strongest rows line up with what Aleph Alpha says it specialized the model for: German, math, and agentic work, here retrieval over documents and the multi-step banking task. For an agent, the split means a Kolibri-1 assistant that searches your documents works in its strongest published area, while one that holds a long conversation and calls several tools in sequence works in its weakest.

One caveat from the report: on BFCL v4, web search ran on a different backend from the public leaderboard and each model sampled at its own parameters, so those scores do not compare with leaderboard entries.

Which Kolibri-1 Capabilities Have No German Benchmark?

Kolibri-1's German overall score is 70.8 and its English one 75.5, but the two are not built from the same tests. The Kolibri tech report states that only four of the eight categories behind the overall scores have German benchmarks, so the German figure rests on knowledge, math, agentic retrieval, and industry RAG alone.

CategoryGerman score publishedBenchmarks in the category
KnowledgeYesGPQA Diamond, Humanity's Last Exam, MMLU-ProX
MathYesAIME 2025, AIME 2026
Agentic retrievalYesAgentic Wiki QA, the German counterpart to MuSiQue and HotpotQA
Industry RAGYesGerman Public Sector, Aerospace, Industrial Drive Technology
Agentic tool useNoBFCL, Tau2-Bench, Tau3-Bench, TerminalBench, BrowseComp
Grounding and hallucinationsNoRGB, AA-Omniscience, SQuAD, FRAMES, SealQA
CodeNoLiveCodeBench, HumanEval+, SWE-Bench Verified
Instruction followingNoIFBench

The German scores that do involve tools are all retrieval tasks, where the model searches and reads before it answers:

  • Agentic Wiki QA - 69.4 on multi-hop German questions answered by searching a document corpus with tools.
  • Industry RAG - 67.5 averaged across the three German customer proxies, against 89.7 across the two English ones. The two sets cover different domains, and the report calls two of the five proxies too small to show a trend, so the gap is not a clean language effect.
  • Honeypot - 80.8 on agentic retrieval questions drawn in both German and English, which the report keeps out of either language's average.

No German score covers a tool that changes something, such as booking an appointment, filing a ticket, or updating a record. The tech report also describes how German tool use entered training: English tool-calling conversations were translated into German "with tool schemas, calls and results left unchanged", alongside German conversations that search the live German Wikipedia.

Part of Kolibri-1's German tool training therefore paired German requests with tools defined in English. If your tool names, descriptions, or parameter documentation are German, test that configuration directly instead of assuming the English scores carry over; the retrieval results above are the closer guide for an agentic RAG assistant.

Does Kolibri-1 Abstain When the Evidence Is Missing?

The model card says Aleph Alpha trained Kolibri-1 with abstention data and with RL environments built on its Merlin-Arthur protocol, which shows the model each question with parts of the context hidden. Where the hidden parts remove the evidence, the model is trained to abstain.

The Merlin-Arthur paper treats the retrieval pipeline as an interactive proof system. The generator model, Arthur, receives context of unknown provenance: Merlin gives helpful evidence, while Morgana injects adversarial, misleading context, and Arthur learns to answer when the evidence supports the answer and to abstain when the evidence is insufficient.

Aleph Alpha's launch post adds that Arthur does not know which player built the context, so any guess on a redacted context counts as incorrect, even a lucky one, which removes the usual reward for guessing. It reports these results:

  • Holding back - on RGB questions whose documents do not contain the answer, Kolibri-1 holds back 85.6% of the time, against 73.9% for Kolibri Origin and 79.6% for Qwen3.6 35B-A3B.
  • Inventing nothing - on the same questions, it produces no falsehood 87.3% of the time, against 84.3% for Qwen3.6 35B-A3B.
  • AA-Omniscience - of the questions it does not answer correctly, it abstains or gives a partial answer on 44%, up from 15% for Kolibri Origin and below the 56.7% of Qwen3.6 35B-A3B.
  • M/A grounding score - 0.23, Aleph Alpha's own metric for how much of an answer provably came from the document, where Kolibri Origin scores 0.

The RGB benchmark also tests wrong evidence: in its counterfactual robustness setting, the supplied documents contain false information. The model card's release table shows Kolibri-1 behind there and on closed-book questions:

  • Error correction - 34.0 on RGB Fact-Check, against 74.0 for Qwen3.6 35B-A3B.
  • Closed-book answers - 51.0 on RGB Closed-Book, the lowest score in the table, against 79.0 for Qwen3.6 35B-A3B.

Read together, the rows say Kolibri-1 is good at noticing when evidence is absent, weaker at noticing when evidence is present but wrong, and knows less from memory than its peers. For a document assistant, that makes the quality of the retrieved documents the main risk: an outdated policy PDF in the index is likely to be answered from rather than questioned.

The model card sets the deployment boundary to match. Kolibri-1 is meant for systems "in which a person reviews the model's output before it is acted on", and for orchestration layers that call APIs "provided the calling system validates the results".

Note

Note: Kolibri-1's abstention scores come from English benchmarks. TestMu AI Agent Testing runs German and English scenarios against your own assistant and returns a Green, Yellow, or Red production-readiness verdict, so the deployment decision rests on your documents instead of the model card. Try TestMu AI free!

How Should You Test Kolibri-1 Tool Calls in German and English?

Write every tool-calling case twice, once in German and once in English, with the same expected tool and arguments, so any difference you find comes from the language and not from a different task. These are the variations I would cover first:

TestGerman and English variationPass condition
Request languageThe same request asked in German and in EnglishSame tool and arguments, with the reply in the user's language
Tool definition languageAn English schema against German tool names and descriptionsTool selection holds steady across both
Dates and numbers"4. Oktober 2026", "04.10.2026", "1.234,56 Euro"Arguments match the schema's format, such as 2026-10-04 and 1234.56
Umlauts and compounds"Müller", "Hauptstraße", "Bundessozialgericht"Passed through intact, unless your API expects a transliteration
Multi-turn changeThe user changes one parameter two turns later, such as "doch lieber Freitag"The next call carries the new value, not the old one
Tool errorsThe tool returns an error or an empty resultA sensible retry or a plain statement of the failure, never an invented result
No tool neededA question the model can answer directly, or no tool that fitsNo tool call at all
Reasoning effortThe same case at none and at highThe same tool choice, within your latency budget

Kolibri-1 was trained to recover from tool errors: according to the model card, erroneous tool calls and terminal actions were masked out during training, so it learned the recovery without learning the mistake. Run the tool-error row in both languages anyway, since multi-turn work, where recovery matters most, is where its published tool scores are weakest.

Grade the call and its effect, not the model's summary of it. An agent that tells the user "Termin gebucht" after the booking API returned an error has committed an agent action hallucination, and the transcript alone will not reveal it.

For Kolibri-1 agents that change state, TestMu AI Agent Assurance tests how agents actually behave across workflows, tools, and actions. It derives a test suite from your code or spec, invokes the agent for real, checks each criterion against what the run changed rather than what the agent reported, and lists separately what it could not verify. It is pre-alpha and installs through the rook CLI.

How Do You Test Kolibri-1 Abstention Before Production?

Build abstention cases from the documents your assistant will actually answer from, and cover wrong evidence and over-cautious refusals as well as the missing-evidence case Kolibri-1 was trained on. The structure follows standard RAG testing, with each case written in German and in English:

CaseHow to build itExpected behavior
Answer presentA question whose answer is in the supplied documentA correct answer drawn from the document
Evidence removedThe same question with the answering passage deletedA plain statement that the documents do not say, such as "Das geht aus den Unterlagen nicht hervor", with no guess
Evidence wrongThe same document with one date or amount changed to contradict a known factThe conflict flagged, or whatever policy you defined for conflicting sources
Answer paraphrasedThe answer present but worded differently, such as "Kündigungsfrist" in the question and "Frist für die Kündigung" in the textAn answer, not a refusal
Recently changed factA closed-book question whose answer moved, such as which law sets the German Impressum duty (§ 5 DDG since May 2024, § 5 TMG before)The current answer or a stated uncertainty, never the superseded law

Track two rates separately: how often the assistant abstains when the evidence is removed, and how often it refuses when the answer is present. A model tuned to abstain can push both up at once, and only the pair tells you whether the assistant became more careful or simply less helpful. For scoring the answers themselves, the methods in LLM hallucination detection apply to Kolibri-1 unchanged.

Agent Testing from TestMu AI runs this kind of suite against a chat endpoint, which suits a self-hosted Kolibri-1 server:

  • Scenarios from your documents - upload the documents your assistant answers from, and it generates 60 to 100+ test scenarios per workflow, in the language you choose at generation time.
  • Hallucination scoring - every conversation is scored on 9 metrics, including Hallucination Detection, Context Awareness, and Completeness.
  • Abstention as a rule - a custom validation criterion such as "If the answer is not in the provided documents, say so and do not guess" is scored pass or fail with a confidence level, alongside the standard metrics.
  • Private endpoints - chat endpoints connect over REST, WebSocket, or SSE, and agents behind a firewall connect through a secure tunnel without a public URL.
  • Pipeline runs - the Agent Testing CLI wraps each test message in your API's request format with --body-template and reads the reply with --response-path, so the same suite can gate a deployment.

Should You Build a German Agent on Kolibri-1?

Start with one workflow you already run in German, such as answering questions over internal policy documents, and run the paired tool-calling and abstention cases above at the reasoning effort you plan to ship. If Kolibri-1 holds back on removed evidence in German as reliably as in English and keeps its tool arguments intact in both languages, it is a credible self-hosted base for that workflow.

To turn those cases into a repeatable suite, follow getting started with Agent Testing and point it at your Kolibri-1 endpoint.

Author

...

Anubhav Singhmaar

Blogs: 41

  • Linkedin

Anubhav Singhmaar is an AI Product Manager at TestMu AI driving Kane CLI, the command-line tool that brings browser automation to the terminal, turning natural-language flows into runs in a real Chrome browser that return pass or fail with shareable proof. He owns the roadmap and prioritization and works with engineering to ship developer-facing features. Before TestMu AI, he spent over four years at Sprinklr owning enterprise voice AI across APAC and EMEA. A mechanical engineer turned product manager, he grounds guidance in real QA workflows.

Reviewer

...

Sirajuddin Khan

Reviewer

  • Linkedin

Sirajuddin Khan is Vice President of Product Management at TestMu AI (formerly LambdaTest), where he drives the company's agentic AI product strategy, building a suite of autonomous agents that includes Agentic Browsers and Agentic Visual Testing and shifting the unit of work from test execution to autonomous outcomes. One of the company's earliest product leaders, he has owned the roadmap for the high-performance execution cloud and grew the cross-browser testing products from early adoption to market leadership. He brings over a decade of experience across SaaS, B2B, and eCommerce, with earlier product roles at Wydr and ShopClues, where his catalog and search work cut delivery SLAs and lifted seller activity. Sirajuddin holds an MBA in Information Technology from Sikkim Manipal University and a B.Tech in Computer Science Engineering from Maharshi Dayanand University.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

Aleph Alpha Kolibri-1 FAQs

Did you find this page helpful?

More Related Blogs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests