Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Cloudflare Clef: Choosing and Testing a Decision Model
Cloudflare Clef: Choosing and Testing a Decision Model
Cloudflare Clef and Clef-flash are open decision models that return typed probabilities. See where each wins, where Clef-flash drops, and what to test first.
Published on:
Clef-flash, the smaller of Cloudflare's two new decision models, answers in a median 38.8 ms, against 524.1 ms for TypeSafe AI's Jev, according to the Clef model card. The same table puts Clef-flash at 66.8 macro-F1 on CLINC150+OOS, against 97.4 for the larger Clef.
Cloudflare released both models on 1 October 2026, hosted on Workers AI and published as open weights under Apache 2.0, in its Clef launch post. They answer typed questions about an input with probabilities instead of text, and Cloudflare says they are fully compatible with the Jev API, so a team already calling Jev can swap them in easily.
CLINC150+OOS mixes in requests that match none of the known intents, the input a routing agent must not force into a category, so picking between Clef and Clef-flash is a choice about which errors your agent can afford. If Jev is new to you, what is Jev covers the model Clef was built to be compatible with.
TL;DR
Cloudflare Clef is a family of open-weight decision models, the 27B Clef and the 9B Clef-flash, that return a probability for every allowed answer to typed yes/no, choice, and score questions instead of generating text. It exists so agents can make fast, bounded decisions inside a workflow, such as routing a ticket or flagging a domain.
- Clef-flash accuracy: Is Clef-flash as accurate as Clef? No, not on every task. Clef-flash matches Clef on tool-calling benchmarks such as BFCL, but scores 66.8 against 97.4 on CLINC150+OOS and 35.6 against 79.4 on RAGTruth hallucination detection.
- Latency: Clef-flash is the fastest of the three on Cloudflare's runs, with a 38.8 ms median latency. Clef's median is 209.3 ms and Jev's is 524.1 ms. Clef-flash's p95 latency rises to 122.4 ms.
- Jev compatibility: Does Clef work with Jev requests? Yes. Clef accepts the Jev and SystemOne request body, returns the same response shape, and adds image input, which Cloudflare says Jev lacks today.
- Reasoning: Does Clef beat Jev on reasoning-heavy questions? No. Jev scores 78.3 on GPQA Diamond against 48.0 for Clef. Jev also outscores Clef on MMLU-Pro, BBH, and When2Call.
- Pre-production testing: Test Clef on your own labeled decisions with out-of-scope inputs, confidence bands, over-long state, and reworded schemas. TestMu AI Agent Assurance then checks what an agent actually did with each Clef decision.
What Is Cloudflare Clef?
Clef is a decision model: given a state, such as a support message, a JSON record, or an image, and a schema of typed questions, it returns a probability for every allowed option of every question in a single forward pass. There is no free-form text and no output to parse. The model card defines three question types:
- noul - a true or false question; the answer is the probability of true.
- choice - named options with descriptions; the answer carries the chosen option, a confidence, and a probability for each option.
- score - ordered options indexed from 0; the answer carries a probability-weighted score, a confidence, a legend, and the probabilities.
The Workers AI Clef docs set the hosted limits: a 65,536-token context window, 1 to 64 questions per request, and up to four inline images, with remote image URLs rejected. Long text state is truncated to fit the model's token limit.
| Specification | Clef | Clef-flash |
|---|---|---|
| Backbone | Qwen3.8-27B, frozen | Qwen3.5-9B, frozen |
| Workers AI model ID | @cf/cloudflare/clef | @cf/cloudflare/clef-flash |
| Context window | 65,536 tokens | 65,536 tokens |
| Inputs (open weights) | Text, JSON, images, and video | Text, JSON, images, and video |
| Median latency | 209.3 ms | 38.8 ms |
| p95 latency | 238.6 ms | 122.4 ms |
| License | Apache 2.0 | Apache 2.0 |
Cloudflare's launch post gives one in-house use: its Threat Intelligence team gives Clef a domain through Cloudflare's Browser Run, which renders the site, and gets back category probabilities, such as fashion, ecommerce, or phishing. That loop took Clef 2.2 seconds, while gpt-oss-120b took 4.7 seconds in the same workflow and returned only two classifications.
How Clef Decides Without Generating Text
Clef runs its Qwen backbone in a prefill-only pass, then a small joint schema head scores every valid option in parallel, so the decision step is non-autoregressive and has no text to generate token by token. Both backbones stay frozen, and the post-training goes into the routing head and rank-256 low-rank adapters. The launch post describes the training choices behind it, each of which points to a test:
- Calibration - Cloudflare paired label-smoothed cross-entropy with a Brier loss to calibrate the probabilities, so confidence is meant to track accuracy. Measure that on your own data before you set an automation threshold.
- Field order - the synthetic training data permuted field orders, prompts, and schema structures, so shuffling questions should leave answers unchanged. It is a cheap invariance test.
- Ordinal answers - Reinforcement Learning for Calibrated Decisions (RLCD) grants partial credit for adjacent ordinal choices, so a near miss on a score question is penalized less than a far one in training. Decide whether an off-by-one severity is acceptable in your workflow.
- Input length - the open-weights encoder defaults to a max_length of 16,384 tokens, well below the hosted window, so a self-hosted Clef can decide on a shorter slice of the same input unless you raise it.
Clef vs Clef-flash vs Jev on the Decision Index
Cloudflare ran the Decision Index suite itself and publishes every row in the model card, so treat these as vendor-reported numbers. These are the rows closest to the decisions an agent makes:
| Benchmark | What it measures | Clef | Clef-flash | Jev |
|---|---|---|---|---|
| BFCL (case exact) | Choosing the right function and arguments | 98.5 | 98.8 | 95.8 |
| API-Bank (accuracy) | Calling APIs correctly | 91.9 | 93.1 | 88.2 |
| BANKING77 (macro-F1) | Intent classification across 77 banking intents | 94.2 | 90.9 | 79.7 |
| CLINC150+OOS (macro-F1) | Intent classification with out-of-scope requests mixed in | 97.4 | 66.8 | 89.3 |
| RAGTruth (hallucination F1) | Flagging hallucinated content in retrieval-augmented answers | 79.4 | 35.6 | 76.5 |
| When2Call MCQ (accuracy) | Deciding whether to call a tool, ask, or answer | 72.4 | 65.6 | 81.0 |
| GPQA Diamond (accuracy) | Graduate-level science reasoning | 48.0 | 51.0 | 78.3 |
| ForecastBench (Brier, lower is better) | How well forecast probabilities match outcomes | 13.9 | 10.6 | 17.4 |
Grouped by the kind of decision, the rows read like this:
- Tool selection - Clef-flash matches or beats Clef on BFCL and API-Bank, so its speed costs nothing on these rows.
- Guardrail-style decisions - Clef-flash drops 30.6 points on CLINC150+OOS and 43.8 points on RAGTruth against Clef, the two rows that test spotting an input that fits no category or an answer the evidence does not support.
- Reasoning-heavy questions - Jev leads both Clef models on GPQA Diamond, MMLU-Pro (82.7 against 65.9 for Clef), and BBH (92.9 against 73.7), and on When2Call.
The model card also scores four end-to-end workflows from Typesafe's eval suite, where the margins are narrower:
- Invoice processing - exact actions 64.7 for Clef, 57.1 for Clef-flash, and 61.8 for Jev.
- Customer service - exact actions 76.3, 77.0, and 76.0.
- Security incidents - exact actions 62.9, 61.7, and 61.7.
- Agent trace observability - primary action 68.5, 69.8, and 71.6, where Jev leads.
The best exact-action score is between 62.9 and 77.0 in each of the three workflows that report one, which is the strongest argument for checking what happens after the decision. RAGTruth scores a model as a judge of hallucinated text, the same job covered in LLM hallucination detection, so the Clef-flash score there matters if you planned to use it as a cheap groundedness check.
Which Clef Model Fits Which Decision?
Map each decision in your agent to the benchmark row that resembles it, and start from the model that leads there. This is how I would set the defaults:
| Decision in your agent | Start with | Evidence from the model card |
|---|---|---|
| Picking a tool or filling API arguments | Clef-flash | BFCL 98.8 and API-Bank 93.1 at a 38.8 ms median |
| Routing among a fixed, complete set of intents | Clef-flash, checked against Clef | BANKING77 90.9 against 94.2 |
| Routing where some requests fit no category | Clef | CLINC150+OOS 97.4 against 66.8 |
| Flagging unsupported or hallucinated content | Clef | RAGTruth 79.4 against 35.6 |
| Whether to call a tool at all or ask the user | Clef and Jev, side by side | When2Call 81.0 for Jev against 72.4 for Clef |
| Questions that need multi-step reasoning or expert knowledge | An LLM or Jev | GPQA Diamond 78.3 for Jev against 48.0 for Clef |
| Classifying screenshots, receipts, or other images | Clef or Clef-flash | Both carry a vision encoder; Cloudflare says Jev classifies text only |
A two-tier setup is a reasonable starting point: Clef-flash on every request, with anything below a confidence threshold sent to Clef or a human. Set that threshold from your own calibration results, not from the benchmark table.
How to Test Clef Before It Drives an Agent
Collect a few hundred real decisions your team has already labeled, such as routed tickets or reviewed domains, and run them through both Clef models. Then run these checks, each one aimed at a failure the benchmark table hints at:
| Test | How to run it | What a failure looks like |
|---|---|---|
| Per-class accuracy | Report recall and precision per option, not one overall score | A rare, costly option such as "urgent" or "phishing" missed while the average looks healthy |
| Out-of-scope inputs | Add inputs that fit none of your options, and give the schema an explicit "none of these" option | A confident assignment to a real option, the failure CLINC150+OOS measures |
| Confidence bands | Bucket decisions by probability and compare each bucket's accuracy with its average probability | Decisions above your automation threshold that are wrong more often than the threshold implies |
| Long state | Send inputs near and past the token limit with the deciding fact at the end | The answer flips once the deciding fact falls past the truncation point |
| Order invariance | Shuffle the order of questions and of options within each question | Different answers for the same state |
| Criteria wording | Reword option descriptions without changing their meaning | Probabilities that swing more than your tolerance |
| Jev-to-Clef swap | Replay recorded Jev requests against Clef and diff the answers | Disagreements on the decisions that trigger actions |
| Tail latency | Measure p95 and p99 from your own region at your real request rate | The hot path's timeout exceeded at p95 even though the median is fine |
| Fine-tuned checkpoints | Rerun the full suite on every checkpoint from Cloudflare's new RL fine-tuning service | Domain accuracy up while out-of-scope and general cases regress |
The fine-tuning row comes from Cloudflare's own warning that fine-tuning may give up some general-purpose performance in exchange for accuracy in a specific domain. The service starts as a hands-on engagement with its forward-deployed engineers, with a self-serve platform planned later.
Web classification adds a step before Clef: the page has to be fetched and rendered, and because Workers AI does not take image URLs, screenshots must be captured and inlined. TestMu AI Browser Cloud runs real Chrome sessions that render JavaScript-heavy pages, captures full-page or element screenshots, and reaches staging or internal pages through its built-in tunnel; the same loop for Jev is measured in Jev Browser Cloud.
Note: When Clef misclassifies a page, you need to see the page it was given. Every TestMu AI Browser Cloud session records video, console logs, network logs, and a step-by-step command replay, so you can check whether the page was blocked, half-loaded, or simply misread. Try TestMu AI free!
Testing What the Agent Does With a Clef Decision
A decision model's output is a probability, and the risk sits in what the agent does next: routing the ticket, blocking the domain, or skipping the escalation. A confidence score is the model's own report about the decision, so it cannot confirm that the action happened or was right; Jev typed output makes the same point for Jev.
TestMu AI Agent Assurance tests how agents actually behave across workflows, tools, and actions. For an agent that acts on Clef decisions, it works through these steps:
- Discovery - it reads the agent's codebase to work out what the agent does, and marks a tool as unknown rather than guessing when nothing declares it.
- Scenarios - it generates functional, non-functional, and adversarial scenarios, each with criteria that are graded one by one.
- Real runs - it invokes the agent the way a user would, through a command, an HTTP endpoint, or an MCP server.
- Evidence - each criterion is judged against files that changed, artifacts produced, and tool calls checked against the agent's declared tools, not against the agent's summary.
- Assurance gap - criteria it could not check are reported as Unable to Verify and kept out of the pass rate, so the report shows how much of the run was actually proven.
Agent Assurance is pre-alpha and installs through the rook CLI. In a pipeline, a finished run exits 0 whether its scenarios passed or failed, so gate on the per-criterion verdicts in the report rather than on the exit code.
Getting Started With Cloudflare Clef
Pick one reversible decision your agent already makes, such as ticket routing, and replay your labeled history through Clef-flash and Clef on Workers AI. Promote Clef-flash only where its per-class accuracy and confidence bands hold up against Clef on your data.
Then cover the actions that follow each decision with the Agent Assurance overview as your setup guide.
Author
Samyak Goyal is a Senior Member of Technical Staff at TestMu AI engineering Kane CLI, the command-line tool that runs browser automation from the terminal, where a flow described in natural language executes in a real Chrome browser and returns pass or fail with shareable proof. He is a backend engineer with 4+ years of experience, previously an SDE at Innovaccer, where he built APIs, introduced Kafka, and cut deployment from weeks to hours. Samyak also builds multi-agent systems, skill-orchestration frameworks, and a personal copilot that indexes 200+ microservice repositories.
Reviewer
Sirajuddin Khan is Vice President of Product Management at TestMu AI (formerly LambdaTest), where he drives the company's agentic AI product strategy, building a suite of autonomous agents that includes Agentic Browsers and Agentic Visual Testing and shifting the unit of work from test execution to autonomous outcomes. One of the company's earliest product leaders, he has owned the roadmap for the high-performance execution cloud and grew the cross-browser testing products from early adoption to market leadership. He brings over a decade of experience across SaaS, B2B, and eCommerce, with earlier product roles at Wydr and ShopClues, where his catalog and search work cut delivery SLAs and lifted seller activity. Sirajuddin holds an MBA in Information Technology from Sikkim Manipal University and a B.Tech in Computer Science Engineering from Maharshi Dayanand University.
Cloudflare Clef FAQs
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




