Hero Background

Power Your Software Testing with AI Agents and Cloud

The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.

AIAgent Testing

Claude Opus 5.5 Explained: What's New & Why It Matters for AI Agents

Discover what's new and why it matters. Powered by Claude Opus 5.5, TestMu AI delivers Agent Assurance for your AI deployments - test web, mobile, and AI agents, all on one platform.

Published on:

Change one line in an agent's config, from claude-opus-5 to claude-opus-5-5, and the agent gets cheaper to run on the same prompts. The same change makes forced tool calls fail, removes the option to switch thinking off, and lowers the default effort from high to medium.

Anthropic released Claude Opus 5.5 on September 22, 2026, as the first model in its Claude 5.5 family. Its launch post says the model performs at the level of Claude Fable 5.1 on most work, costs 40% less to run than Opus 5, and is the strongest-performing model Anthropic has tested to date on its automated behavioral audit, an alignment suite that tests Claude across thousands of simulated scenarios.

The gains in coding, computer use, and visual reading arrive alongside API changes that break some agent harnesses without an error message, so an upgrade needs a test pass before the new model ID reaches production. TestMu AI runs that pass on one platform for web apps, mobile apps, and AI agents, with Agent Assurance covering the AI deployment itself.

TL;DR

Claude Opus 5.5 is Anthropic's newest Opus model, released on September 22, 2026, for long-running agentic coding and knowledge work. It costs $4 per million input tokens and $20 per million output tokens, keeps a 1M-token context window, and runs adaptive thinking that cannot be switched off.

  • Cost vs Opus 5: Is Claude Opus 5.5 cheaper to run than Opus 5? Yes. Tokens cost 20% less, cache reads cost 60% less at $0.20 per million, and Anthropic's tests show a 40% lower cost on typical workloads at default settings.
  • Thinking mode: Can thinking be turned off on Claude Opus 5.5? No. A request that disables thinking or sets budget_tokens returns a 400 error, and the effort parameter, which now defaults to medium, controls depth, latency, and cost.
  • Forced tool use: Does Claude Opus 5.5 support tool_choice any or tool? No. Both return a 400 error. Agents keep tool_choice auto, add strict tool use for schema-valid arguments, and name the tool in the prompt.
  • Terminal-Bench 4.0: Claude Opus 5.5 scores 66.4% on Terminal-Bench 4.0, ahead of Claude Fable 5.1 at 55.8% and Claude Opus 5 at 52.3%, in Anthropic's launch benchmarks for agentic coding in a command line.
  • Video agents: Claude Opus 5.5 reads charts, diagrams, and screenshots more precisely than Opus 5, which helps agents that edit video or hold on-camera sessions. TestMu AI Agent Testing grades video agents from a session recording and a turn-labeled transcript.

What Is Claude Opus 5.5?

Claude Opus 5.5 is Anthropic's Opus-tier model for long-running agentic coding and knowledge work, and the first release in the Claude 5.5 family. Anthropic's launch post says Claude Sonnet 5.5 and Claude Haiku 5.5 will follow "in the coming weeks."

The Opus 5.5 model page lists adaptive thinking as always on, with the effort parameter as the control for thinking depth. Its core specifications:

  • Model ID - claude-opus-5-5 on the Claude API, Google Cloud, Microsoft Foundry, and Claude Platform on AWS, and anthropic.claude-opus-5-5 on Amazon Bedrock.
  • Context and output - a 1M-token context window and 128K max output tokens, or up to 300K output tokens on the Message Batches API with the output-300k-2026-03-24 beta header.
  • Pricing - $4 per million input tokens and $20 per million output tokens, with cache reads at $0.20 per million and a 50% discount on batch requests.
  • Thinking and effort - adaptive thinking is always on, and the default effort level is medium.
  • Inputs and outputs - text and images in, text out, with a reliable knowledge cutoff of June 2026.
  • Retirement - not sooner than September 22, 2027.

For how other models handle multi-step agent work, the best agentic AI LLM models roundup compares the wider field.

What Is Claude Opus 5.5 Good For?

Anthropic's launch benchmarks put Opus 5.5 in the lead on agentic coding, computer use, and knowledge work. The same post cautions that at this capability level, benchmark margins "have become a less reliable guide to real-world differences," and unless noted, its Opus 5.5 scores use adaptive thinking at max effort.

Agentic Coding

Anthropic's launch post describes Opus 5.5 as particularly good at long, sprawling jobs such as codebase-wide migrations and audits, and reports these results:

  • Terminal-Bench 4.0 - 66.4% at xhigh effort on multi-step command-line tasks, against 55.8% for Claude Fable 5.1 and 52.3% for Opus 5.
  • CursorBench 4.0 - 57.8% on ambiguous, multi-file tasks taken from real Cursor sessions, against 46.6% for Opus 5.
  • Codebase audit - an early tester audited and fixed a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and used 2.5x as many tokens.
  • HAProxy rewrite - translating HAProxy from C into Rust took Opus 5.5 9.5 hours against 12 for Fable 5.1, at 51% lower cost.

Mario Rodriguez, Chief Product Officer at GitHub, is quoted in the launch post: in VS Code, Opus 5.5 "solved more terminal tasks than Opus 5 in less than half the steps."

Knowledge Work and Research

Opus 5.5 is less likely to invent a figure. In an internal test from Anthropic's launch post, 16 of 18 Opus 5.5 reports on a company's quarterly results cleared a quality bar that failed any invented figure or quote, while neither Fable 5.1 nor Opus 5 cleared it in any attempt.

  • GDPval-AA v2.1 - 1846 Elo on real-world work across 44 occupations, against 1735 for Fable 5.1 and 1708 for Opus 5.
  • AutomationBench - 40.0% on Zapier's benchmark of business workflows across connected apps, against 26.9% for Opus 5.
  • Financial modeling - on a merger analysis built in Excel and turned into an executive presentation, Opus 5.5 finished in 63 minutes against 93 for Opus 5, at 50% lower cost.

Computer Use and Visual Reading

Anthropic's launch post reports gains on both operating a computer from screenshots and reading charts:

  • OSWorld 2.0 - 81.8% on computer use tasks, against 80.7% for Fable 5.1 and 74.0% for Opus 5.
  • Chartography - 89.0% with tools on visual chart recognition, against 83.4% for Opus 5.
  • Prompt injection - matches or beats Opus 5 in every setting Anthropic tested, including computer use and web browsing, and ties Fable 5.1 for the lowest prompt injection success rate on Gray Swan's benchmark.

Clearer Writing

Writing was one of the most common areas of feedback on Opus 5, according to Anthropic's launch post, and the post reports these changes:

  • Structure - Opus 5.5 puts the most important information up front, is less likely to use jargon or idiosyncratic phrases, and follows the writing rules you give it.
  • Length - in Box's evaluations, it used a third of the tokens Opus 5 did, and its answers were 40% less verbose without losing accuracy.

Claude Opus 5.5 vs Claude Opus 5

Anthropic's launch post says Opus 5.5 costs less per token than Opus 5 and also uses fewer tokens per task, and that cache reads make up the majority of agentic and coding work costs.

The defaults moved too. Per Anthropic's What's new in Opus 5.5 page, a request that omits effort now runs at medium, where on Opus 5 it ran at high.

AttributeClaude Opus 5.5Claude Opus 5
Cost on typical workloads40% lower than Opus 5 at default settingsBaseline
Output speedMore than 30% faster than Opus 5Baseline
Input / output price$4 / $20 per million tokens$5 / $25 per million tokens
Cache reads$0.20 per million tokens$0.50 per million tokens
Cache writes$5 per million tokens$6.25 per million tokens
Default effortmediumhigh
Switching thinking offNot allowed; returns a 400 errorAllowed at effort high or below
Forced tool useNot supported; auto and none onlySupported
Terminal-Bench 4.066.4%52.3%
OSWorld 2.081.8%74.0%

On the new medium default, Zimu Li of Factory is quoted in the launch post:

Comma

Fast mode delivers up to 2.5x higher output tokens per second from Claude Opus 5.5 at $8 per million input tokens and $40 per million output tokens. It is a research preview on the Claude API only, not on Amazon Bedrock, Google Cloud, or Microsoft Foundry.

API Changes for Agent Builders

The What's new page lists four breaking changes for code already running on Opus 5, plus behavior changes that fail no request at all:

  • Thinking can't be disabled - a request with thinking: {"type": "disabled"} or a manual budget_tokens setting returns a 400 invalid_request_error. Responses can begin with empty thinking blocks, so read content blocks by type rather than by position.
  • Forced tool use returns an error - tool_choice set to any or tool fails with a 400, including on the token counting endpoint. Keep auto, set strict: true for schema-valid arguments, and say in the prompt when the tool applies.
  • Thinking blocks are tied to the model and the conversation - a conversation that moves from Opus 5.5 to any model other than Claude Fable 5.1 or Claude Mythos 5.1 continues without the earlier reasoning. On accounts created on or after August 31, 2026, replaying a thinking block after editing the system prompt, the tools, or an earlier message returns a 400.
  • The earlier computer use tool is rejected - on the Claude API and Google Cloud, computer use requires the computer_toolset_20260801 toolset, and computer_20251124 returns a 400. Amazon Bedrock still accepts the earlier tool.
  • Progress notes move into thinking blocks - the short notes the model writes between tool calls arrive as thinking blocks with empty text at the default display setting, so an interface that streams them to users goes silent with no error.
  • Effort defaults to medium - a request that omits effort runs one level lower than on Opus 5, and at a given level the model tends to think more per turn, most of all at xhigh and max.

The same page recommends keeping conversations append-only, changing instructions or tools through mid-conversation system messages rather than edits, so the thinking-block check never triggers.

Claude Opus 5.5 for Video Agents

Video agents cover two jobs: production agents that plan, storyboard, edit, and render video, and on-camera agents such as AI interviewers, onboarding assistants, and virtual front desks. Opus 5.5 takes text and images and returns text, per its model page, so a video agent built on it works from frames, storyboards, transcripts, and timeline data.

Anthropic's What's new page says the model reads values off dense charts and layout-dependent visuals "much more precisely without tools," so vision workarounds built for earlier models may no longer be needed. Applied to video work:

  • Frame review - an agent can check a rendered frame or thumbnail against a storyboard, a caption file, or a brand spec with fewer image-processing tools in the loop.
  • Editing through a screen - the 81.8% OSWorld 2.0 score, against 74.0% for Opus 5, is the relevant number for an agent that operates an editor or an upload portal through computer use.
  • Long sessions - cache reads at $0.20 per million tokens cut the cost of an agent that rereads the same project context across a long render or review job.
  • Readable updates - clearer progress messages make an unattended edit easier for a producer to review the next morning.
Youtube thumbnail

Production agents often finish in a browser, on an upload portal, a CMS, or a scheduling dashboard. Kane CLI gives the agent a way to confirm that step worked: it runs a natural-language objective in a real Chrome browser and returns pass or fail with an evidence pack. In agent mode it emits NDJSON events that Claude Code, Codex CLI, or Gemini CLI can read, and it installs with npm install -g @testmuai/kane-cli.

Austin Siewert

Austin Siewert

Co-Founder, Steadfast Systems

Discovered @TestMu AI yesterday. Best browser testing tool I've found for my use case. Great pricing model for the limited testing I do 👏

2M+ Devs and QAs rely on TestMu AI

Deliver immersive digital experiences with Next-Generation Mobile Apps and Cross Browser Testing Cloud

How to Test Agents Built on Claude Opus 5.5

Even with its own alignment audit, Anthropic's launch post says that building evaluations that reliably catch every failure before deployment "remains an unsolved problem," and that Opus 5.5 "often suspects it is being evaluated." Before moving an agent to claude-opus-5-5, extend your AI agent testing plan with these checks:

  • Set effort explicitly - run your existing scenarios on Opus 5 and Opus 5.5 at the same effort level, then again at medium, since that is now the default.
  • Replay tool-calling paths - confirm a tool_use block arrives wherever the prompt names a tool, and retry when it does not, because auto does not guarantee a call.
  • Exercise the refusal path - send prompts close to the cybersecurity and biology safeguards, and check that refusals come back with stop_reason set to refusal and that your fallback model picks up the turn.
  • Watch the progress UI - set thinking display to updates or summarized and confirm users still see status between tool calls.
  • Grade whole sessions - score complete conversations for hallucination, completeness, context awareness, and tone, not single responses.

Agent Testing from TestMu AI runs that last check at scale: 15+ specialized testing agents execute scenarios generated from your requirement documents and score chat and voice agents on 9 quality metrics, including hallucination, bias, completeness, and context awareness. Scores are reproducible, so a suite run on Opus 5 and again on Opus 5.5 gives a direct comparison and a Green, Yellow, or Red verdict.

For on-camera agents, Agent Testing joins the session from a web URL as a simulated candidate with a real face and a real voice, then grades the conversation. Every session returns a downloadable recording and a turn-labeled transcript with timestamped evidence, and the introducing Video Agent Testing post walks through a full run.

Agents that act rather than talk, such as coding and workflow agents built on Opus 5.5, fall under Agent Assurance. It tests how an agent behaves across workflows, tools, and actions, and grades file changes and tool calls against evidence rather than against the agent's own summary of its work.

Note

Note: Test chat, voice, phone, and video agents built on Claude Opus 5.5 with TestMu AI Agent Testing. Try TestMu AI free!

Safeguards and Limits to Plan Around

Opus 5.5 is the first Opus model to launch with safeguards similar to Fable 5.1's on cybersecurity, biology, and distillation, all of which fall back to another model transparently, per the launch post:

  • Cybersecurity routing - finding and fixing bugs in your own code is allowed, but most cybersecurity tasks are re-routed to Claude Opus 4.8.
  • Biology safeguards - Opus 5.5 uses the same biology safeguards as Fable 5.1, and vetted organizations can apply to the Life Sciences Verification Program for research use.
  • Benchmarks ran with safeguards on - when they intervened, cybersecurity tasks were completed by Opus 4.8 and biology and frontier LLM development tasks by Opus 5, which Anthropic says likely reduces Opus 5.5's scores.
  • Containment - in a new evaluation, Opus 5.5 attempted to circumvent boundaries around 85% less often than Opus 5 or Claude Mythos 5.1, and every attempt was low severity and self-reported.
  • Data retention - Opus 5.5 is available with zero data retention and carries watermarking measures to comply with the EU AI Act.

A refusal arrives as an HTTP 200 response with stop_reason set to refusal, so an agent that reads content without checking the stop reason can pass an empty turn downstream. An agent that works on one run and fails on the next needs repeated runs rather than a single pass, which is the subject of AI agent reliability.

Getting Started With Claude Opus 5.5

Run your current agent suite against claude-opus-5-5 with effort set explicitly, next to the same suite on Opus 5, and compare the results before you change the production model ID. Agent Testing scores both runs on the same metrics, and the getting started with Agent Testing guide covers connecting your agent.

Author

...

Chaitanya Sharma

Blogs: 15

  • Linkedin

Chaitanya Sharma is an AI Product Manager at TestMu AI (formerly LambdaTest), where he builds agentic AI capabilities focused on computer vision and multi-modality, moving testing beyond static script execution toward autonomous, agent-driven workflows. Before TestMu AI he shipped 135+ features at Sprinklr for a no-code community and website builder used by Fortune 500 enterprises including Dell, Samsung, and Polestar. At Policybazaar he led the zero-to-one launch of a digital lending and insurance marketplace embedded in Bahrain's dominant payments app, building a risk-intelligence engine that compressed loan-approval times by 80%. He explored machine learning and NLP through research at the University of Cambridge, and holds a B.Tech from Delhi Technological University.

Reviewer

...

Sirajuddin Khan

Reviewer

  • Linkedin

Sirajuddin Khan is Vice President of Product Management at TestMu AI (formerly LambdaTest), where he drives the company's agentic AI product strategy, building a suite of autonomous agents that includes Agentic Browsers and Agentic Visual Testing and shifting the unit of work from test execution to autonomous outcomes. One of the company's earliest product leaders, he has owned the roadmap for the high-performance execution cloud and grew the cross-browser testing products from early adoption to market leadership. He brings over a decade of experience across SaaS, B2B, and eCommerce, with earlier product roles at Wydr and ShopClues, where his catalog and search work cut delivery SLAs and lifted seller activity. Sirajuddin holds an MBA in Information Technology from Sikkim Manipal University and a B.Tech in Computer Science Engineering from Maharshi Dayanand University.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

Claude Opus 5.5 FAQs

Did you find this page helpful?

More Related Blogs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests