Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Claude Opus 5.5 Explained: What's New & Why It Matters for AI Agents
Claude Opus 5.5 Explained: What's New & Why It Matters for AI Agents
Discover what's new and why it matters. Powered by Claude Opus 5.5, TestMu AI delivers Agent Assurance for your AI deployments - test web, mobile, and AI agents, all on one platform.
Published on:
Change one line in an agent's config, from claude-opus-5 to claude-opus-5-5, and the agent gets cheaper to run on the same prompts. The same change makes forced tool calls fail, removes the option to switch thinking off, and lowers the default effort from high to medium.
Anthropic released Claude Opus 5.5 on September 22, 2026, as the first model in its Claude 5.5 family. Its launch post says the model performs at the level of Claude Fable 5.1 on most work, costs 40% less to run than Opus 5, and is the strongest-performing model Anthropic has tested to date on its automated behavioral audit, an alignment suite that tests Claude across thousands of simulated scenarios.
The gains in coding, computer use, and visual reading arrive alongside API changes that break some agent harnesses without an error message, so an upgrade needs a test pass before the new model ID reaches production. TestMu AI runs that pass on one platform for web apps, mobile apps, and AI agents, with Agent Assurance covering the AI deployment itself.
TL;DR
Claude Opus 5.5 is Anthropic's newest Opus model, released on September 22, 2026, for long-running agentic coding and knowledge work. It costs $4 per million input tokens and $20 per million output tokens, keeps a 1M-token context window, and runs adaptive thinking that cannot be switched off.
- Cost vs Opus 5: Is Claude Opus 5.5 cheaper to run than Opus 5? Yes. Tokens cost 20% less, cache reads cost 60% less at $0.20 per million, and Anthropic's tests show a 40% lower cost on typical workloads at default settings.
- Thinking mode: Can thinking be turned off on Claude Opus 5.5? No. A request that disables thinking or sets budget_tokens returns a 400 error, and the effort parameter, which now defaults to medium, controls depth, latency, and cost.
- Forced tool use: Does Claude Opus 5.5 support tool_choice any or tool? No. Both return a 400 error. Agents keep tool_choice auto, add strict tool use for schema-valid arguments, and name the tool in the prompt.
- Terminal-Bench 4.0: Claude Opus 5.5 scores 66.4% on Terminal-Bench 4.0, ahead of Claude Fable 5.1 at 55.8% and Claude Opus 5 at 52.3%, in Anthropic's launch benchmarks for agentic coding in a command line.
- Video agents: Claude Opus 5.5 reads charts, diagrams, and screenshots more precisely than Opus 5, which helps agents that edit video or hold on-camera sessions. TestMu AI Agent Testing grades video agents from a session recording and a turn-labeled transcript.
What Is Claude Opus 5.5?
Claude Opus 5.5 is Anthropic's Opus-tier model for long-running agentic coding and knowledge work, and the first release in the Claude 5.5 family. Anthropic's launch post says Claude Sonnet 5.5 and Claude Haiku 5.5 will follow "in the coming weeks."
The Opus 5.5 model page lists adaptive thinking as always on, with the effort parameter as the control for thinking depth. Its core specifications:
- Model ID -
claude-opus-5-5on the Claude API, Google Cloud, Microsoft Foundry, and Claude Platform on AWS, andanthropic.claude-opus-5-5on Amazon Bedrock. - Context and output - a 1M-token context window and 128K max output tokens, or up to 300K output tokens on the Message Batches API with the
output-300k-2026-03-24beta header. - Pricing - $4 per million input tokens and $20 per million output tokens, with cache reads at $0.20 per million and a 50% discount on batch requests.
- Thinking and effort - adaptive thinking is always on, and the default effort level is
medium. - Inputs and outputs - text and images in, text out, with a reliable knowledge cutoff of June 2026.
- Retirement - not sooner than September 22, 2027.
For how other models handle multi-step agent work, the best agentic AI LLM models roundup compares the wider field.
What Is Claude Opus 5.5 Good For?
Anthropic's launch benchmarks put Opus 5.5 in the lead on agentic coding, computer use, and knowledge work. The same post cautions that at this capability level, benchmark margins "have become a less reliable guide to real-world differences," and unless noted, its Opus 5.5 scores use adaptive thinking at max effort.
Agentic Coding
Anthropic's launch post describes Opus 5.5 as particularly good at long, sprawling jobs such as codebase-wide migrations and audits, and reports these results:
- Terminal-Bench 4.0 - 66.4% at
xhigheffort on multi-step command-line tasks, against 55.8% for Claude Fable 5.1 and 52.3% for Opus 5. - CursorBench 4.0 - 57.8% on ambiguous, multi-file tasks taken from real Cursor sessions, against 46.6% for Opus 5.
- Codebase audit - an early tester audited and fixed a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and used 2.5x as many tokens.
- HAProxy rewrite - translating HAProxy from C into Rust took Opus 5.5 9.5 hours against 12 for Fable 5.1, at 51% lower cost.
Mario Rodriguez, Chief Product Officer at GitHub, is quoted in the launch post: in VS Code, Opus 5.5 "solved more terminal tasks than Opus 5 in less than half the steps."
Knowledge Work and Research
Opus 5.5 is less likely to invent a figure. In an internal test from Anthropic's launch post, 16 of 18 Opus 5.5 reports on a company's quarterly results cleared a quality bar that failed any invented figure or quote, while neither Fable 5.1 nor Opus 5 cleared it in any attempt.
- GDPval-AA v2.1 - 1846 Elo on real-world work across 44 occupations, against 1735 for Fable 5.1 and 1708 for Opus 5.
- AutomationBench - 40.0% on Zapier's benchmark of business workflows across connected apps, against 26.9% for Opus 5.
- Financial modeling - on a merger analysis built in Excel and turned into an executive presentation, Opus 5.5 finished in 63 minutes against 93 for Opus 5, at 50% lower cost.
Computer Use and Visual Reading
Anthropic's launch post reports gains on both operating a computer from screenshots and reading charts:
- OSWorld 2.0 - 81.8% on computer use tasks, against 80.7% for Fable 5.1 and 74.0% for Opus 5.
- Chartography - 89.0% with tools on visual chart recognition, against 83.4% for Opus 5.
- Prompt injection - matches or beats Opus 5 in every setting Anthropic tested, including computer use and web browsing, and ties Fable 5.1 for the lowest prompt injection success rate on Gray Swan's benchmark.
Clearer Writing
Writing was one of the most common areas of feedback on Opus 5, according to Anthropic's launch post, and the post reports these changes:
- Structure - Opus 5.5 puts the most important information up front, is less likely to use jargon or idiosyncratic phrases, and follows the writing rules you give it.
- Length - in Box's evaluations, it used a third of the tokens Opus 5 did, and its answers were 40% less verbose without losing accuracy.
Claude Opus 5.5 vs Claude Opus 5
Anthropic's launch post says Opus 5.5 costs less per token than Opus 5 and also uses fewer tokens per task, and that cache reads make up the majority of agentic and coding work costs.
The defaults moved too. Per Anthropic's What's new in Opus 5.5 page, a request that omits effort now runs at medium, where on Opus 5 it ran at high.
| Attribute | Claude Opus 5.5 | Claude Opus 5 |
|---|---|---|
| Cost on typical workloads | 40% lower than Opus 5 at default settings | Baseline |
| Output speed | More than 30% faster than Opus 5 | Baseline |
| Input / output price | $4 / $20 per million tokens | $5 / $25 per million tokens |
| Cache reads | $0.20 per million tokens | $0.50 per million tokens |
| Cache writes | $5 per million tokens | $6.25 per million tokens |
| Default effort | medium | high |
| Switching thinking off | Not allowed; returns a 400 error | Allowed at effort high or below |
| Forced tool use | Not supported; auto and none only | Supported |
| Terminal-Bench 4.0 | 66.4% | 52.3% |
| OSWorld 2.0 | 81.8% | 74.0% |
On the new medium default, Zimu Li of Factory is quoted in the launch post:
Fast mode delivers up to 2.5x higher output tokens per second from Claude Opus 5.5 at $8 per million input tokens and $40 per million output tokens. It is a research preview on the Claude API only, not on Amazon Bedrock, Google Cloud, or Microsoft Foundry.
API Changes for Agent Builders
The What's new page lists four breaking changes for code already running on Opus 5, plus behavior changes that fail no request at all:
- Thinking can't be disabled - a request with
thinking: {"type": "disabled"}or a manualbudget_tokenssetting returns a 400invalid_request_error. Responses can begin with empty thinking blocks, so read content blocks bytyperather than by position. - Forced tool use returns an error -
tool_choiceset toanyortoolfails with a 400, including on the token counting endpoint. Keepauto, setstrict: truefor schema-valid arguments, and say in the prompt when the tool applies. - Thinking blocks are tied to the model and the conversation - a conversation that moves from Opus 5.5 to any model other than Claude Fable 5.1 or Claude Mythos 5.1 continues without the earlier reasoning. On accounts created on or after August 31, 2026, replaying a thinking block after editing the system prompt, the tools, or an earlier message returns a 400.
- The earlier computer use tool is rejected - on the Claude API and Google Cloud, computer use requires the
computer_toolset_20260801toolset, andcomputer_20251124returns a 400. Amazon Bedrock still accepts the earlier tool. - Progress notes move into thinking blocks - the short notes the model writes between tool calls arrive as thinking blocks with empty text at the default
displaysetting, so an interface that streams them to users goes silent with no error. - Effort defaults to medium - a request that omits
effortruns one level lower than on Opus 5, and at a given level the model tends to think more per turn, most of all atxhighandmax.
The same page recommends keeping conversations append-only, changing instructions or tools through mid-conversation system messages rather than edits, so the thinking-block check never triggers.
Claude Opus 5.5 for Video Agents
Video agents cover two jobs: production agents that plan, storyboard, edit, and render video, and on-camera agents such as AI interviewers, onboarding assistants, and virtual front desks. Opus 5.5 takes text and images and returns text, per its model page, so a video agent built on it works from frames, storyboards, transcripts, and timeline data.
Anthropic's What's new page says the model reads values off dense charts and layout-dependent visuals "much more precisely without tools," so vision workarounds built for earlier models may no longer be needed. Applied to video work:
- Frame review - an agent can check a rendered frame or thumbnail against a storyboard, a caption file, or a brand spec with fewer image-processing tools in the loop.
- Editing through a screen - the 81.8% OSWorld 2.0 score, against 74.0% for Opus 5, is the relevant number for an agent that operates an editor or an upload portal through computer use.
- Long sessions - cache reads at $0.20 per million tokens cut the cost of an agent that rereads the same project context across a long render or review job.
- Readable updates - clearer progress messages make an unattended edit easier for a producer to review the next morning.
Production agents often finish in a browser, on an upload portal, a CMS, or a scheduling dashboard. Kane CLI gives the agent a way to confirm that step worked: it runs a natural-language objective in a real Chrome browser and returns pass or fail with an evidence pack. In agent mode it emits NDJSON events that Claude Code, Codex CLI, or Gemini CLI can read, and it installs with npm install -g @testmuai/kane-cli.
Austin Siewert
Co-Founder, Steadfast Systems
Discovered @TestMu AI yesterday. Best browser testing tool I've found for my use case. Great pricing model for the limited testing I do 👏
2M+ Devs and QAs rely on TestMu AI
Deliver immersive digital experiences with Next-Generation Mobile Apps and Cross Browser Testing Cloud
How to Test Agents Built on Claude Opus 5.5
Even with its own alignment audit, Anthropic's launch post says that building evaluations that reliably catch every failure before deployment "remains an unsolved problem," and that Opus 5.5 "often suspects it is being evaluated." Before moving an agent to claude-opus-5-5, extend your AI agent testing plan with these checks:
- Set effort explicitly - run your existing scenarios on Opus 5 and Opus 5.5 at the same effort level, then again at
medium, since that is now the default. - Replay tool-calling paths - confirm a
tool_useblock arrives wherever the prompt names a tool, and retry when it does not, becauseautodoes not guarantee a call. - Exercise the refusal path - send prompts close to the cybersecurity and biology safeguards, and check that refusals come back with
stop_reasonset torefusaland that your fallback model picks up the turn. - Watch the progress UI - set thinking
displaytoupdatesorsummarizedand confirm users still see status between tool calls. - Grade whole sessions - score complete conversations for hallucination, completeness, context awareness, and tone, not single responses.
Agent Testing from TestMu AI runs that last check at scale: 15+ specialized testing agents execute scenarios generated from your requirement documents and score chat and voice agents on 9 quality metrics, including hallucination, bias, completeness, and context awareness. Scores are reproducible, so a suite run on Opus 5 and again on Opus 5.5 gives a direct comparison and a Green, Yellow, or Red verdict.
For on-camera agents, Agent Testing joins the session from a web URL as a simulated candidate with a real face and a real voice, then grades the conversation. Every session returns a downloadable recording and a turn-labeled transcript with timestamped evidence, and the introducing Video Agent Testing post walks through a full run.
Agents that act rather than talk, such as coding and workflow agents built on Opus 5.5, fall under Agent Assurance. It tests how an agent behaves across workflows, tools, and actions, and grades file changes and tool calls against evidence rather than against the agent's own summary of its work.
Note: Test chat, voice, phone, and video agents built on Claude Opus 5.5 with TestMu AI Agent Testing. Try TestMu AI free!
Safeguards and Limits to Plan Around
Opus 5.5 is the first Opus model to launch with safeguards similar to Fable 5.1's on cybersecurity, biology, and distillation, all of which fall back to another model transparently, per the launch post:
- Cybersecurity routing - finding and fixing bugs in your own code is allowed, but most cybersecurity tasks are re-routed to Claude Opus 4.8.
- Biology safeguards - Opus 5.5 uses the same biology safeguards as Fable 5.1, and vetted organizations can apply to the Life Sciences Verification Program for research use.
- Benchmarks ran with safeguards on - when they intervened, cybersecurity tasks were completed by Opus 4.8 and biology and frontier LLM development tasks by Opus 5, which Anthropic says likely reduces Opus 5.5's scores.
- Containment - in a new evaluation, Opus 5.5 attempted to circumvent boundaries around 85% less often than Opus 5 or Claude Mythos 5.1, and every attempt was low severity and self-reported.
- Data retention - Opus 5.5 is available with zero data retention and carries watermarking measures to comply with the EU AI Act.
A refusal arrives as an HTTP 200 response with stop_reason set to refusal, so an agent that reads content without checking the stop reason can pass an empty turn downstream. An agent that works on one run and fails on the next needs repeated runs rather than a single pass, which is the subject of AI agent reliability.
Getting Started With Claude Opus 5.5
Run your current agent suite against claude-opus-5-5 with effort set explicitly, next to the same suite on Opus 5, and compare the results before you change the production model ID. Agent Testing scores both runs on the same metrics, and the getting started with Agent Testing guide covers connecting your agent.
Author
Chaitanya Sharma is an AI Product Manager at TestMu AI (formerly LambdaTest), where he builds agentic AI capabilities focused on computer vision and multi-modality, moving testing beyond static script execution toward autonomous, agent-driven workflows. Before TestMu AI he shipped 135+ features at Sprinklr for a no-code community and website builder used by Fortune 500 enterprises including Dell, Samsung, and Polestar. At Policybazaar he led the zero-to-one launch of a digital lending and insurance marketplace embedded in Bahrain's dominant payments app, building a risk-intelligence engine that compressed loan-approval times by 80%. He explored machine learning and NLP through research at the University of Cambridge, and holds a B.Tech from Delhi Technological University.
Reviewer
Sirajuddin Khan is Vice President of Product Management at TestMu AI (formerly LambdaTest), where he drives the company's agentic AI product strategy, building a suite of autonomous agents that includes Agentic Browsers and Agentic Visual Testing and shifting the unit of work from test execution to autonomous outcomes. One of the company's earliest product leaders, he has owned the roadmap for the high-performance execution cloud and grew the cross-browser testing products from early adoption to market leadership. He brings over a decade of experience across SaaS, B2B, and eCommerce, with earlier product roles at Wydr and ShopClues, where his catalog and search work cut delivery SLAs and lifted seller activity. Sirajuddin holds an MBA in Information Technology from Sikkim Manipal University and a B.Tech in Computer Science Engineering from Maharshi Dayanand University.
Claude Opus 5.5 FAQs
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests





