Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Claude Haiku 5.5: Pricing, Benchmarks, and Subagent Setup
Claude Haiku 5.5: Pricing, Benchmarks, and Subagent Setup
Claude Haiku 5.5 runs about 75% cheaper than Haiku 4.5. See its pricing, benchmarks, breaking API changes, Claude Code subagent setup, and how to test agents.
Published on:
A subagent that summarizes test logs or compacts a long session runs on every turn, so the model behind it can be billed thousands of times a day. Claude Haiku 5.5 cuts that bill, and the same switch makes some Haiku 4.5 requests fail outright: a thinking budget, a custom temperature, or an assistant prefill now returns a 400 error.
Anthropic released Claude Haiku 5.5 on October 7, 2026, and its launch post calls it the cheapest, fastest, and most capable small model the company has released. On average it costs around 75% less to run than Haiku 4.5, and Anthropic positions it for summaries, compactions, database queries, classification, and subagent work alongside Opus 5.5 and Sonnet 5.5.
Moving that work to a cheaper model usually means running more of it, with less human review per call. TestMu AI tests the coding agents and customer-facing chatbots built on models like this one, and provides and checks the real browsers that browser agents work in, before a model change reaches production.
TL;DR
Claude Haiku 5.5 is Anthropic's fastest and lowest-priced Claude model, released on October 7, 2026, for high-volume work such as summaries, classification, and subagent tasks. It costs $0.10 per million input tokens and $0.50 per million output tokens on prompts up to 100K tokens, and it has a 1M-token context window.
- Cost vs Haiku 4.5: Is Claude Haiku 5.5 cheaper to run than Haiku 4.5? Yes. Anthropic puts the average saving at around 75%. Token prices are 90% lower on prompts up to 100K tokens and 50% lower above that, partly offset by a tokenizer that counts about 30% more tokens for the same text.
- Effort setting: Does Claude Haiku 5.5 support the effort parameter? Yes. Claude Haiku 5.5 is the first Haiku model with effort levels, from low to max, and medium is the default on the Claude API and in Claude Code.
- Claude Code subagents: Can Claude Haiku 5.5 run as a Claude Code subagent? Yes. In Claude Code v2.1.293 or later, the haiku alias resolves to Haiku 5.5 on the Anthropic API, so a subagent defined with model: haiku runs on it.
- Haiku 4.5 requests: Do Haiku 4.5 requests work unchanged on Claude Haiku 5.5? Not always. Manual thinking budgets, non-default temperature or top_p values, any top_k value, and assistant prefill all return a 400 error on Haiku 5.5.
- Complex agentic coding: Does Claude Haiku 5.5 replace Sonnet 5.5 for complex agentic coding? No. Claude Haiku 5.5 scores 39.2% on Terminal-Bench 4.0 against 70.6% for Sonnet 5.5, and Anthropic recommends Sonnet 5.5 or Opus 5.5 for that work.
- Anthropic browser use tool: Does Claude Haiku 5.5 support Anthropic's browser use tool? Yes, on the Claude API and Google Cloud. The tool sends every action to a browser your application runs, such as a real Chrome session from TestMu AI Browser Cloud.
What Is Claude Haiku 5.5?
The Haiku 5.5 model page lists Claude Haiku 5.5 for high-volume, latency-sensitive tasks such as classification, extraction, and routing. Its comparison table puts Haiku 5.5 below Sonnet 5.5 and Opus 5.5 on price and rates its latency the fastest in the current Claude lineup. Its specifications:
- Model ID -
claude-haiku-5-5on the Claude API, Google Cloud, Microsoft Foundry, and Claude Platform on AWS, andanthropic.claude-haiku-5-5on Amazon Bedrock. - Context and output - a 1M-token context window and 128K max output tokens, or up to 300K output tokens on the Message Batches API with the
output-300k-2026-03-24beta header. - Thinking and effort - adaptive thinking is on by default, the effort parameter sets how much the model thinks, and the default effort is
medium. - Inputs and outputs - text and images in, text out, with a reliable knowledge cutoff of June 2026.
- Tokenizer - the same tokenizer as Claude 4.7 and later models, which counts the same text as more tokens than Haiku 4.5 did.
- Retirement - not sooner than October 7, 2027.
The Claude Opus 5.5 breakdown covers the Opus-tier model in the same family, including the API changes it introduced in September.
Claude Haiku 5.5 Pricing
Anthropic's launch post prices Haiku 5.5 in two tiers split at a 100K-token prompt, a size it says covers most requests to the previous Haiku model. Prices per million tokens:
| Token type | Haiku 5.5, prompts up to 100K | Haiku 5.5, prompts over 100K | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 | $2.00 |
| Output | $0.50 | $2.50 | $5.00 | $10.00 |
| Cache writes | $0.125 | $0.625 | $1.25 | $2.50 |
| Cache reads | $0.01 | $0.05 | $0.10 | $0.10 |
- The 75% average - per the launch post's footnote, Haiku 5.5 is priced 90% lower than Haiku 4.5 on requests up to 100K tokens and 50% lower above that. On Haiku 4.5, 90% of requests fell under 100K tokens, and the figure also accounts for the newer tokenizer using slightly more tokens per task.
- Batch and caching discounts - the model page lists a 50% discount on Batch API input and output, and 1-hour cache writes at $0.20 per million tokens on prompts up to 100K.
- Sonnet 5.5 cache reads - at the same launch, Anthropic halved Sonnet 5.5's cache read price from $0.20 to $0.10 per million tokens, which it says makes Sonnet 5.5 around 20% cheaper on most agentic work.
- API credits for subscribers - Claude Max 5x subscribers get $100 a month in Claude Platform API credits, Max 20x subscribers get $200, and Team subscribers get up to $500 pooled across users, usable on any Claude model.
Recount prompts before you rebudget. Anthropic's Haiku 5.5 migration guide says the same input text produces approximately 30% more tokens on Haiku 5.5 than on Haiku 4.5, so usage fields, token counts, max_tokens limits, and cost estimates measured on Haiku 4.5 all need redoing with model set to claude-haiku-5-5.
Claude Haiku 5.5 Benchmarks
Anthropic's launch post scores Haiku 5.5 against Haiku 4.5, OpenAI's GPT-6 Luna, and Sonnet 5.5, which it includes for reference. The Haiku 5.5 system card documents how each evaluation ran.
| Benchmark | Area | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|---|
| GDPval-AA v2.1 | Knowledge work | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 | Knowledge work | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1, offline subset | Computer use | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity's Last Exam, no tools | Multidisciplinary reasoning | 45.9% | 10.2% | Not reported | 56.9% |
| Humanity's Last Exam, with tools | Multidisciplinary reasoning | 57.4% | 18.7% | Not reported | 64.5% |
| Terminal-Bench 4.0 | Agentic coding | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 (Main) | Agentic coding | 46.4% | Not reported | 42.4% | 52.1% at xhigh effort |
| Chartography, no tools | Visual reasoning | 46.4% | 6.4% | 29.1% | 61.6% |
Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, so one model ID can trade cost for intelligence per request. The launch post charts OSWorld, GDPval-AA, and Humanity's Last Exam at each effort level from low to max, with cost per attempt on the horizontal axis. The scores point to where the model fits:
- Computer use - Haiku 5.5 scores 72.4% on the OSWorld 2.1 offline subset against 15.7% for Haiku 4.5 and 48.9% for GPT-6 Luna, and the launch post names browser use and live customer support as speed-sensitive tasks the model suits.
- Complex agentic coding - Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 against 39.2% for Haiku 5.5. Anthropic says Sonnet 5.5 and Opus 5.5 remain the better choices there, and that Haiku 5.5 suits narrowly scoped jobs such as compaction, summarization, and subagent work.
- Speed - Haiku 5.5 is Anthropic's fastest model at each model's standard speed, though it runs slower than the Opus models in Fast Mode.
Early-access customers quoted in the launch post reported these results:
- Asana - over a 30% reduction in task-completion latency and up to 2.5x faster inference per agent turn on its AI Teammates eval suite, compared with the model it uses today.
- HubSpot - 92.8% averaged over three runs on its simulated-portal CRM suite, the best score HubSpot has seen there, plus the fastest completion and lowest false positive rate on a stale-records audit task.
- AlphaSense - 0.84 against 0.76 for Haiku 4.5 across 400 queries to Ask in Document, a feature that handles about 8M calls a week.
- Box - 11 points higher than Haiku 4.5 at about half the latency in early testing.
Claude Haiku 5.5 as a Subagent in Claude Code
Claude Code needs v2.1.293 or later for Haiku 5.5. Its model configuration docs say the haiku alias resolves to Haiku 5.5 on the Anthropic API, but still resolves to Haiku 4.5 on Amazon Bedrock, Google Cloud's Agent Platform, Microsoft Foundry, and Claude Platform on AWS. To run it as the main model, use /model claude-haiku-5-5 or start Claude Code with claude --model claude-haiku-5-5.
The more common setup keeps Opus 5.5 or Sonnet 5.5 as the lead and hands narrow jobs to Haiku 5.5. A subagent file in .claude/agents/ sets the model and effort in its frontmatter:
---
name: test-log-summarizer
description: Summarizes test runner and CI output into failing tests, error lines, and likely flaky tests. Use after a test run finishes.
tools: Read, Grep, Glob
model: haiku
effort: low
---
You summarize test output. For each failing test, report the file, the
assertion that failed, and the first error line. Flag tests that failed
and then passed on retry as likely flaky. Do not edit any files.Claude Code picks a subagent's model in this order:
- The
modelparameter Claude passes for that specific invocation. - The
modelfield in the subagent's frontmatter, whereinheritselects the main conversation's model. - The
CLAUDE_CODE_SUBAGENT_MODELenvironment variable, which also sets the default for agent team teammates and workflow agents. - The main conversation's model.
The Claude Code subagents guide covers the remaining frontmatter fields and where subagent files live. Alex Wang of Rogo described the lead-and-subagent split in the launch post:
Anthropic's Prompting Claude Haiku 5.5 guide flags a subagent risk: at low and medium effort, the model sometimes reports a code change as done without running a check that exercises it. The guide supplies a system-prompt paragraph that tells coding agents to run the project's tests, type-checker, or build before reporting.
To check whether a Haiku 5.5 subagent did what it reported, Agent Assurance from TestMu AI tests how agents behave across workflows, tools, and actions. Its rook CLI derives a suite from your agent's code or spec, invokes the agent for real, and checks each criterion against what the run changed, such as files on disk and tool calls, instead of the agent's own summary. Criteria it could not verify are reported separately, so an unverified "done" never counts as a pass.
For a model switch, rook finds Claude Code agents defined in .claude/agents/, and a command profile such as claude -p "{{goal}}" invokes the agent the way a user would. Every run is diffed against the last, so scenarios that newly fail after you set model: haiku are reported apart from flaky ones. Agent Assurance is in pre-alpha, and rook installs with npm install -g @testmuai/rook.
Anthropic's Browser and Computer Use Tools on Claude Haiku 5.5
Haiku 5.5 supports Anthropic's browser use tool on the Claude API and Google Cloud. One browser_toolset_20260801 entry in tools gives Claude 27 member tools by default, such as navigate, read_page, left_click, and screenshot, and your application executes every call against its own browser automation. Nothing runs on Anthropic's side.
- New to the Haiku line - the migration guide notes that Haiku 4.5 does not support the browser use tool.
- Platform coverage - the browser use tool is not available on Amazon Bedrock, Claude Platform on AWS, Microsoft Foundry, or in Claude Managed Agents.
- Computer use toolset - on the Claude API and Google Cloud, Haiku 5.5 takes computer use only through
computer_toolset_20260801, and a request that declarescomputer_20250124returns a 400 error. - SDK support - Anthropic's launch post says the Python and TypeScript SDKs now support computer use and browser use in beta, and calls Haiku 5.5 especially well suited to both.
- Untrusted pages - the tool docs recommend a domain allowlist enforced at the network layer and re-checked after redirects, and page reads built from the accessibility tree or visible text rather than raw DOM, so hidden text never reaches Claude.
Your application has to supply that browser. TestMu AI Browser Cloud provides real Chrome sessions on demand for AI agents, with Playwright, Puppeteer, and Selenium adapters, a built-in tunnel to localhost and staging, and session transparency for debugging a failed run. The what is Browser Cloud doc covers session setup, and browser infrastructure for AI agents walks through scaling, debugging, and deploying those sessions.
To confirm what a browser agent changed in your web app, Kane CLI runs a natural-language objective in a real Chrome browser and returns a pass or fail with a replayable .evidence pack of per-step screenshots, DOM snapshots, and console and network logs. It attaches to an existing Chrome with --cdp-endpoint or to a remote browser with --ws-endpoint, and it installs with npm install -g @testmuai/kane-cli.
Breaking Changes From Claude Haiku 4.5
The migration guide lists these request changes for code that called Haiku 4.5:
| Haiku 4.5 request | Result on Haiku 5.5 | Fix |
|---|---|---|
thinking set to enabled with budget_tokens | 400 error | Send {"type": "adaptive"} and set output_config.effort |
temperature or top_p at a non-default value, or any top_k | 400 error | Omit all three and steer behavior through the prompt |
| Assistant prefill as the last turn | 400 error, even with thinking off | End with a user turn and use structured outputs for formatting |
computer_20250124 on the Claude API or Google Cloud | 400 error | Declare computer_toolset_20260801 |
Edited system, tools, or earlier messages with thinking blocks sent back | 400 error | Keep conversations append-only |
| Code that reads the first content block as the answer | The response can begin with a thinking block | Select content blocks by type |
A small max_tokens value | Can stop after thinking, before any text | Raise max_tokens or lower the effort level |
Forced tool_choice, any or a named tool | Accepted, but the call starts with no thinking block | Use auto and say in the prompt when to call the tool |
- Automated migration - in Claude Code,
/claude-api migrate this project to claude-haiku-5-5runs the bundled Claude API skill, which swaps the model ID, applies the breaking changes, and returns a checklist of items to verify by hand. - Thinking display - Haiku 5.5 returns each
thinkingblock with an emptythinkingfield by default. Setdisplaytosummarizedto receive summarized thinking, as Haiku 4.5 returned. - Turning thinking off - on the API,
thinking: {"type": "disabled"}works atlow,medium, andhigheffort and returns a 400 error atxhighandmax. In Claude Code, thinking can't be turned off on Haiku 5.5. - Priority Tier - not supported on Haiku 5.5, so organizations with a Haiku 4.5 Priority Tier commitment need to plan capacity separately.
- Stored conversations - Haiku 5.5 thinking blocks work only in the account that produced them or a linked account. Replayed through any other account, they are dropped before the model sees them, and the request still succeeds.
How to Test Agents Before You Switch to Claude Haiku 5.5
Anthropic expects existing Haiku 4.5 prompts to perform well without changes, and its prompting guide still documents several places where Haiku 5.5 behaves differently. Before moving production traffic, run your AI agent testing suite against these behaviors:
- Compare effort levels - run the same evals at two or three levels. The prompting guide says that in long agent prompts,
lowis more likely to skip a search, stop early, or skip a check, and that moving fromlowtomediumroughly halved early stopping while more than doubling output tokens per attempt. - Catch unverified completions - for coding agents, compare what the agent reports against the tests, builds, and files the run actually touched.
- Check for empty replies - at
xhigheffort in multi-turn chats, the model sometimes writes its whole answer in its thinking and ends the turn with no visible text. - Test search behavior - give the model today's date when it has a search tool. The guide reports that the date grounded answers in recent results, and that a blanket instruction to always search triggered searches on half of the prompts that needed none.
- Pressure-test chatbot rules - Anthropic suggests a system-prompt line that keeps rules in force when a user argues, gives a sympathetic reason, claims an approved exception, or keeps asking. Script those exact patterns as test conversations.
- Route mid-turn messages correctly - Haiku 5.5 is trained to resist prompt injection through tool results, so user text placed inside a
tool_resultblock can be ignored as untrusted. Deliver it as a user turn instead. - Handle refusals in the client - Haiku 5.5 has no server-side fallback, and resending a declined request usually returns another refusal.
For chat and voice agents moving to Haiku 5.5, Agent Testing from TestMu AI generates 60 to 100+ scenarios from uploaded docs and PRDs or connected Jira and Confluence, covering edge cases, adversarial inputs, and persona-specific conversations. It scores each conversation on 9 quality metrics, including hallucination, completeness, and context awareness, and returns a Green, Yellow, or Red production-readiness verdict. One suite run on Haiku 4.5 and again on Haiku 5.5 shows whether the cheaper model holds the same quality bar.
Note: Test the chatbots and voice agents you move to Claude Haiku 5.5 with TestMu AI Agent Testing. Try TestMu AI free!
Safeguards and Refusals on Claude Haiku 5.5
Anthropic's launch post reports far fewer instances of misaligned behavior than on Haiku 4.5, and a lower willingness to cooperate with misuse. The safeguards also change what an agent can be asked to do:
- Cybersecurity - more restrictive than Haiku 4.5's safeguards but less restrictive than those on other recent models. They permit a wider range of defensive tasks than Sonnet 5.5's safeguards and still block penetration testing.
- Biology - the same safeguards as Sonnet 5, Sonnet 5.5, and Opus 5, which allow research biology questions and restrict requests judged likely to cause harm.
- Verification programs - organizations doing wider-ranging work can apply to the Life Sciences Verification Program and the Cyber Verification Program.
- Refusal categories - a declined request returns
stop_reasonset torefusal, withstop_details.categoryset tocyber,frontier_llm,bio, orgeneral_harms. Benign work can also triggercyberandgeneral_harms.
A refusal arrives as a regular response rather than an error, so a pipeline that reads content without checking the stop reason can pass an empty result downstream. Teams that run security scans through an agent should test that path directly, and prompt injection testing covers the attack side of the same tool-result boundary.
Getting Started With Claude Haiku 5.5
Start with one high-volume job, such as summarizing test logs or classifying support tickets. Switch its model to claude-haiku-5-5, remove sampling parameters and prefill, and run your existing evals at low and medium effort next to the same evals on Haiku 4.5.
For a chat or voice agent, Agent Testing scores both runs on the same metrics, and the getting started with Agent Testing guide covers connecting your agent.
Author
Chaitanya Sharma is an AI Product Manager at TestMu AI (formerly LambdaTest), where he builds agentic AI capabilities focused on computer vision and multi-modality, moving testing beyond static script execution toward autonomous, agent-driven workflows. Before TestMu AI he shipped 135+ features at Sprinklr for a no-code community and website builder used by Fortune 500 enterprises including Dell, Samsung, and Polestar. At Policybazaar he led the zero-to-one launch of a digital lending and insurance marketplace embedded in Bahrain's dominant payments app, building a risk-intelligence engine that compressed loan-approval times by 80%. He explored machine learning and NLP through research at the University of Cambridge, and holds a B.Tech from Delhi Technological University.
Reviewer
Sirajuddin Khan is Vice President of Product Management at TestMu AI (formerly LambdaTest), where he drives the company's agentic AI product strategy, building a suite of autonomous agents that includes Agentic Browsers and Agentic Visual Testing and shifting the unit of work from test execution to autonomous outcomes. One of the company's earliest product leaders, he has owned the roadmap for the high-performance execution cloud and grew the cross-browser testing products from early adoption to market leadership. He brings over a decade of experience across SaaS, B2B, and eCommerce, with earlier product roles at Wydr and ShopClues, where his catalog and search work cut delivery SLAs and lifted seller activity. Sirajuddin holds an MBA in Information Technology from Sikkim Manipal University and a B.Tech in Computer Science Engineering from Maharshi Dayanand University.
Claude Haiku 5.5 FAQs
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests





