Hero Background

Next-Gen App & Browser Testing Cloud

Trusted by 2 Mn+ QAs & Devs to accelerate their release cycles

Next-Gen App & Browser Testing Cloud
AIAgent TestingCoding

Claude Code Subagents: Setup and Use Cases

Context isolation, the complete frontmatter reference, where definitions live, and how to verify what a subagent actually shipped.

Author

Mythili Raju

Author

Author

Shahzeb Hoda

Reviewer

Published on: August 27, 2026

Anthropic's own engineering write-up on building a multi-agent research system states it plainly: "agents typically use about 4x more tokens than chat interactions, and multi-agent systems use about 15x more tokens than chats."

That is not a cost of subagents working badly. It is the cost of the isolation working correctly, spending tokens in a window that never touches the one you are reading.

This covers what a subagent actually is, the complete frontmatter reference, where definitions live, and how to prove that what one shipped is real, not just a summary that sounds finished.

TL;DR

A Claude Code subagent is a named, isolated instance with its own context window, tool access, model, and permission mode. The main session delegates a task to it, the subagent works alone, and only its final result returns; the file reads, searches, and dead ends behind that result never enter the primary conversation.

  • Context isolation: the entire reason the feature exists, not a side effect. A 200K-token window fills with tool output long before it fills with useful conversation.
  • 16 frontmatter fields: name and description are required; tools, model, permissionMode, maxTurns, skills, mcpServers, hooks, memory, background, effort, isolation, color, and initialPrompt are optional.
  • Concurrency and nesting limits: 20 subagents at once by default, 3 layers of nesting by default, both configurable by environment variable.
  • Verification: a subagent's default tools confirm code compiles, not that a page renders. TestMu AI's Kane CLI gives a subagent a real Chrome browser and a pass or fail it can act on.

What a Subagent Actually Is

A subagent is a Claude instance the main session can delegate to. It gets its own context window, its own tool access list, its own model, and its own permission mode, and none of that state is shared with the session that called it.

The isolation is the whole point. A long session on a real codebase fills its context with file reads, search results, and repeated tool output long before the conversation itself gets long. Once that window is mostly noise, response quality drops, not because the model changed but because the signal-to-noise ratio collapsed. A subagent absorbs the noisy part of a task in its own window and hands back only what the caller actually needs.

  • Own context window - the exploration a subagent does never counts against the main session's budget.
  • Own tool access - a subagent can be scoped to fewer tools than the session that spawned it, never more.
  • Own model - a subagent can run on a cheaper or more capable model than the main session, independently.
  • Own permission mode - a subagent can be locked to read-only exploration while the main session keeps full access.

That isolation guarantees a clean handoff, but it says nothing about whether the work behind the handoff was correct. The verification section below covers how TestMu AI's Kane CLI closes that specific gap. The pattern also generalizes past one vendor's CLI, and our guide to multi-agent AI systems covers the architectures this borrows from.

Built-In vs Custom Subagents

Claude Code ships with built-in subagents you rarely invoke by name; Claude routes to them on its own. Anything more specific than that is a custom subagent you define yourself.

Built-in subagentDefault modelWhat it is for
ExploreInherits session (capped at Opus)Read-only search and codebase understanding
PlanInherits sessionGathers context before Claude presents a strategy in plan mode
general-purposeInherits sessionTasks that need both exploration and modification together

A custom subagent is a Markdown file with YAML frontmatter that you write and check into your project, or keep personally under your home directory. It behaves identically to a built-in one once defined; Claude does not distinguish the two when deciding whether to delegate.

Explore's model default is worth double-checking against a fresh copy of the docs before you rely on it. As of v2.1.198, Explore inherits the main conversation's model rather than always running on Haiku, per Anthropic's own Claude Code subagents documentation. A session on a higher tier runs Explore on Opus; a session on Sonnet or Haiku runs Explore on that same model. Guides written before that release still describe a fixed Haiku default.

The Full Frontmatter Reference

Most write-ups on this topic stop at four or five fields. Anthropic's Claude Code subagents documentation currently defines 16, and two of them (fable as a model choice, and the xhigh and max effort levels) are recent enough that older guides on this exact keyword do not mention them.

FieldRequiredWhat it does
nameYesUnique identifier, lowercase with hyphens only, no colons
descriptionYesThe routing trigger. Claude matches requests against this text to decide when to delegate
toolsNoAllowlist of available tools. Inherits every tool the session has if omitted
disallowedToolsNoDenylist, for removing specific tools without hand-listing everything else
modelNosonnet, opus, haiku, fable, a full model ID, or inherit (the default)
permissionModeNodefault, acceptEdits, auto, dontAsk, bypassPermissions, or plan
maxTurnsNoCaps agentic turns, useful for tasks that can spiral (test-fixing loops especially)
skillsNoSkills preloaded into the subagent's context at spawn
mcpServersNoMCP servers made available to the subagent specifically
hooksNoLifecycle hooks scoped to this subagent's own run
memoryNoPersistent memory scope: user, project, or local
backgroundNoSet true to keep the subagent running in the background
effortNolow, medium, high, xhigh, or max reasoning depth
isolationNoSet to worktree for an isolated git worktree per subagent
colorNoDisplay color in the task list, cosmetic only
initialPromptNoAuto-submitted first turn when the definition runs as a main session, not a subagent

An accessibility-focused example, scoped read-only so it can never touch a file it is reviewing:

---
name: accessibility-auditor
description: Reviews rendered pages for WCAG issues. Use proactively after
             any change to markup, styling, or a form flow.
tools: Read, Grep, Glob, Bash
disallowedTools: Edit, Write
model: sonnet
effort: high
permissionMode: plan
---

You are an accessibility auditor. You do not fix issues, you find and report them.

When invoked:
1. Identify which component or page changed
2. Check for missing alt text, unlabeled form fields, insufficient color
   contrast, and keyboard-trap patterns
3. Reference the relevant WCAG 2.1 success criterion for each finding
4. Report severity as CRITICAL, SERIOUS, MODERATE, or MINOR

Never modify a file. Report findings only.

The description field carries more weight than its length suggests. Claude matches incoming requests against that text to decide whether to delegate automatically, so a vague description gets skipped even for tasks it should own. TestMu AI's own accessibility testing platform runs the equivalent WCAG checks against a real rendered page rather than static markup, which is the gap the verification section below covers for subagents generally.

Where Definitions Live

Five scopes, and Claude Code resolves a naming collision in a strict order rather than picking arbitrarily.

PriorityScopeLocation
1 (highest)Managed settingsDeployed by organization admins
2CLI flagPassed with --agents when launching Claude Code
3Project.claude/agents/, checked into version control
4User~/.claude/agents/, available across every project
5 (lowest)Pluginagents/ directories bundled inside installed plugins

The practical split: project-scoped agents are the team's shared specialists, the reviewer and the auditor everyone's session sees the same way. User-scoped agents are personal, brought to every project you open, tuned to how you individually work.

How Claude Decides to Delegate

Four ways, ranging from fully automatic to a guarantee that leaves nothing to Claude's judgment.

  • Automatic delegation - Claude matches your request against every visible subagent's description field and delegates without being asked.
  • Natural language naming - mention the subagent by name in your prompt; Claude still decides whether to actually delegate.
  • @-mention - type @ and pick from the typeahead list, which guarantees that exact subagent runs.
  • Session-wide default - the --agent CLI flag, or an agent key in .claude/settings.json, makes one subagent's definition run as the entire session.

Automatic delegation is the one worth writing carefully. It reads like a routing rule, not a label: a description of "reviews code" delegates unreliably, while "reviews code for security vulnerabilities. Use proactively after writing authentication, authorization, or data-handling code" gives Claude an actual trigger to match against.

Note

Note: TestMu AI's Kane CLI gives a subagent something a file-editing tool never can: a real Chrome browser and a pass or fail it can trust. Try TestMu AI free!

Can a Subagent Verify Its Own Work

Not with what it gets by default. Read, Edit, Bash, and Grep confirm that code compiles and unit tests pass, none of which is the same claim as a button rendering or a page loading for a real user.

Kane CLI from TestMu AI is built for exactly that call: give it a plain-English objective, and it drives a real Chrome browser and returns a machine-readable verdict. A subagent scoped for verification can shell out to it directly. Here is a real one, saved as a project-level definition:

---
name: browser-verifier
description: Verifies a shipped UI change in a real Chrome browser using Kane
             CLI. Use proactively after any change that touches rendered UI,
             before marking the task complete.
tools: Bash, Read
model: sonnet
permissionMode: default
---

You are a browser verification specialist. You do not read source code to
decide whether a UI change works; you drive a real browser and check what
it renders.

When invoked:
1. Identify the user-facing flow the change affects
2. Write a plain-English Kane CLI objective describing the expected outcome
3. Run it: kane-cli run --agent --headless "<objective>" --url <target-url>
4. Read the terminal run_end event. Report the status field verbatim
5. On status: failed, quote the reason field and stop

Never report a change as verified without a run_end event showing
status: "passed".

The isolation property from the first section is not theoretical. To confirm it, we spawned a real subagent against this repository with a small, verifiable task rather than describe the mechanism secondhand: audit which LearningHubCta variant each of seven blog files published today actually uses. The subagent read seven files and returned only this:

Results (2 CTAs per file for opencode-vs-claude-code and gemini-cli-testing;
1 each for the rest):

- claude-code-hooks/index.js -> KaneCLICertificationCTA
- opencode-vs-claude-code/index.js -> FlakyTestCTA, NextGenerationTestExecutionCTA
- add-testing-to-ai-coding-tools.../index.js -> InfrastructureDoesntBreakCTA
- gemini-cli-testing/index.js -> TestUpToSeventyFasterCTA, KaneCLICertificationCTA
- quality-gate-for-ai-pull-requests/index.js -> ThreeCombinationsCTA
- coding-agent-plugins.../index.js -> GartnerReportCTA
- agentic-test-run-cost-vs-ci/index.js -> WebinarCTA

Distinct variants across all 7 files: 7 unique names, 8 total usages
(KaneCLICertificationCTA repeats once).

Every file it opened, every grep it ran to find the pattern, stayed inside that subagent's own context. The main session never saw them, only the seven-line answer above. That is the mechanism this entire article is about, not a description of it.

The same isolation applies to a browser-verifier subagent, with one difference: instead of grep output, its final result is a Kane CLI run_end event carrying a status field of passed or failed, the exact contract the definition above is written to trust. Kane CLI ships as an agent mode built for this: newline-delimited JSON on stdout, a terminal run_end event with the full result, and standard exit codes (0 passed, 1 failed, 2 environment error, 3 timeout) a hook or a subagent's own logic can branch on. Install it with a single npm command; the Kane CLI introduction documentation covers authentication and the full command reference.

Get Kane CLI certified for free with TestMu AI

Subagents vs Agent Teams

Both delegate work out of the main session. Who talks to whom, and what that costs, is where they split.

DimensionSubagentsAgent teams
Reports toThe caller, as a summarized final resultNobody by default; teammates message each other directly
Your accessThrough the main session onlyYou can open and message any teammate directly
Runs in CI or headless modeYes, this is the normal caseNo, teams require an interactive session
Token costClose to a standard sessionRoughly 7x a standard session, per Anthropic's cost documentation
Best forA self-contained task where only the result mattersWork whose value comes from two workers disagreeing

A subagent type defined once under .claude/agents/ can run as either: spawned normally it is a subagent, spawned as a teammate it becomes one, with the same YAML frontmatter underneath. Our full breakdown of Claude Code agent teams covers enabling teams, the settings that control them, and the exact token multiplier in more depth than fits here.

What Goes Wrong

Four failure modes account for most of the frustration reported with subagents, and none of them are bugs.

  • Silent under-delegation - a description written as a label rather than a trigger phrase gets matched inconsistently, so the same request sometimes delegates and sometimes does not.
  • The token multiplier compounds fast - Anthropic's own figure is roughly 4x tokens for a single agent's delegated work and roughly 15x for a full multi-agent system, so five subagents fired for a task one could handle is not a free parallelization.
  • Concurrency and depth caps are hard limits - 20 concurrent subagents and 3 levels of nesting by default. A workflow assuming more silently queues or fails rather than scaling past them.
  • Fresh context on every spawn - a subagent starts with none of the main session's accumulated understanding, so a task that needs that context runs slower inside a subagent than it would have run directly.

The rule that follows from all four: delegate work that is genuinely self-contained and verbose, not work that needs the same back-and-forth a direct conversation would need anyway.

Note

Note: A subagent's Bash tool proves the code runs. It does not prove the page loads. TestMu AI's Kane CLI closes that gap in one command. Wire the same check into a Stop hook

Which Subagents to Build First

Three cover most of the value before you need a fourth.

  • A read-only explorer scoped to Read, Grep, and Glob, for mapping an unfamiliar part of the codebase without any risk of it editing something.
  • A browser-verifier wired to Kane CLI, invoked proactively after any change that touches rendered UI, using the definition earlier in this article as a starting point.
  • A scoped implementer limited to the files one feature actually touches, with maxTurns set low enough that a stuck loop fails fast instead of burning the session.

Write the description field for each one as a routing rule before anything else, since that single field decides whether automatic delegation ever actually fires.

Context isolation solves the problem this article opened with; it does not solve whether the isolated work was correct. For the flows that split across multiple agents, our guide to multi-agent testing covers verifying the group's output rather than one subagent at a time, and whether coding agents can test their own code covers the structural reason that blind spot exists in the first place. Start with the Kane CLI documentation to wire a browser-verifier subagent into your own project today.

Author

...

Mythili Raju

Blogs: 51

  • Twitter
  • Linkedin

Mythili is a Community Contributor at TestMu AI with 3+ years of experience in software testing and marketing. She holds certifications in Automation Testing, KaneAI, Selenium, Appium, Playwright, and Cypress. At TestMu AI, she leads go-to-market (GTM) strategies, collaborates on feature launches, and creates SEO optimized content that bridges technical depth with business relevance. A graduate of St. Joseph’s University, Bangalore, Mythili has authored 35+ blogs and learning hubs on AI-driven test automation and quality engineering. Her work focuses on making complex QA topics accessible while aligning content strategy with product and business goals.

Reviewer

...

Shahzeb Hoda

Reviewer

  • Linkedin

Shahzeb Hoda is the Associate Director of Marketing and a Community Contributor at TestMu AI, leading strategic initiatives in developer marketing, content, and community growth. With 10+ years of experience in quality engineering, software testing, automation testing, and e-learning, he has authored and reviewed 70+ technical articles on software testing and automation. Shahzeb holds an M.Tech in Computer Science from BIT, Mesra, and is certified in Selenium, Cypress, Playwright, Appium, and KaneAI. He brings deep expertise in CI/CD pipeline automation, cross-browser testing, AI-driven testing practices, and framework documentation. On LinkedIn, he is followed by 3,700+ engineers, developers, DevOps professionals, tech leaders, and enthusiasts.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

Claude Code Subagents FAQs

Did you find this page helpful?

More Related Blogs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests