Claude Code Agent Teams: Lead, Teammates and Limits
Updated 7 skills6 stepsOrigin: Anthropic
Created by Anthropic - https://www.anthropic.com/
Overview
Claude Code Agent Teams is an experimental feature of Anthropic's coding agent that runs several Claude Code sessions as one coordinated team. The Claude Code agent teams documentation describes it as a way to coordinate multiple Claude Code instances working together, with shared tasks, inter-agent messaging and centralized management. The Claude Code glossary puts it more structurally: multiple independent sessions coordinated by a team lead, with a shared task list and peer-to-peer messaging. The point is to split a job that is too broad for one context window into pieces that separate sessions can work on at the same time, while one session keeps the overall picture.
The architecture has three parts. One session acts as the team lead, which, according to Anthropic's agent teams docs, coordinates work, assigns tasks and synthesizes results. Teammates are independent sessions, each operating in its own context window, and they can message each other directly instead of routing every question through the lead. The shared task list is the coordination surface: the lead writes tasks into it, teammates claim and complete them, and dependent tasks wait until their prerequisites are done. You can talk to the lead or to any individual teammate, which the glossary names as a key difference from subagents.
flowchart TD
U[User] --> L[Team lead session]
L --> T[Shared task list]
T --> A[Teammate A with own context]
T --> B[Teammate B with own context]
T --> C[Teammate C with own context]
U -.-> B
Anthropic introduced agent teams on February 5, 2026, in the Claude Opus 4.6 announcement, labelling them a research preview and saying they are best for tasks that split into independent, read-heavy work such as codebase reviews. A companion engineering post, Building a C compiler with a team of parallel Claudes, described a more ambitious use: multiple Claude instances working in parallel on a shared codebase without active human intervention, applied to building a C compiler. The current documentation has since turned the announcement into operational guidance on leads, teammates, shared tasks, messaging, activation and limits. The two launch framings differ in emphasis. The announcement stresses read-heavy review work, while the compiler write-up shows parallel agents writing code together, so treat write-heavy use as possible but harder to get right.
The feature is experimental and disabled by default. You enable it by setting CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 in your environment or settings.json, and without that variable Claude Code does not set up a team, write team directories, or spawn or propose teammates, per the agent teams docs. A June 2026 practitioner comparison notes the variable applies on v2.1.32 or later and lists further constraints, such as a fixed lead and permissions inherited when a teammate is created. The documented limitations shape how you plan a session:
| Limitation | Practical effect |
|---|---|
| No session resume | In-process teammates are not restored by /resume or /rewind, and the lead may message teammates that no longer exist (docs) |
| One team per session | You cannot run a second team or share a team across sessions (docs) |
| No nested teams | Only the lead manages the team; teammates cannot spawn teammates (docs) |
| Slow shutdown | Teammates finish their current request or tool call before stopping (docs) |
| Task status lag | Teammates sometimes fail to mark tasks complete, which blocks dependent tasks (docs) |
The published evidence is thin and points mostly at cost. Anthropic states that teams use significantly more tokens than a single session because every teammate has its own context window, and that token costs scale linearly with the number of teammates (agent teams docs). It also flags coordination overhead, possible conflicts and diminishing returns. A Reddit post reporting 52 controlled benchmarks claims agent teams cost 73-124% more than sequential execution with zero quality gain, blaming each agent loading the full codebase context; that is practitioner evidence, not a peer-reviewed or Anthropic evaluation. On adoption, Anthropic's analysis of roughly 400,000 Claude Code sessions from about 235,000 people between October 2025 and April 2026 does not report what share used agent teams, so no published figure shows how widely the feature is used.
The closest alternative is subagents. The glossary describes subagents as running within a single session and reporting only to the parent, while teammates have separate context windows and can be reached directly; the parallel agents overview frames agent teams as multiple coordinated sessions managed by a lead. Use subagents for focused helpers that return a summary, and a team when workers need to talk to each other about overlapping questions. The skill page on choosing agent teams versus subagents walks through that decision. Teams that track agent work next to human work can mirror the lead's task list in a shared workspace such as Hamster Studio, so the plan stays visible outside the terminal.
Core Principles
The lead owns the plan
One session acts as team lead, and Anthropic's docs give it three jobs: coordinating work, assigning tasks and synthesizing results. Only the lead can manage the team, and teammates cannot spawn their own teammates. That makes the lead's decomposition the single biggest lever on quality. If the task list is vague, every teammate inherits the vagueness.
Separate context windows are both the feature and the bill
Each teammate runs in its own context window, which lets a team hold more of a large codebase in view than one session can, as the glossary describes. The same property is why Anthropic warns that token costs scale linearly with teammate count. Every teammate you add pays to load its own context. Add one only when it will read or change something the others are not already covering.
Split along independent, read-heavy seams
The launch announcement says agent teams are best for tasks that split into independent, read-heavy work, such as codebase reviews. A practitioner guide reports the corollary: teams struggle with write-heavy tasks where multiple agents modify the same files. Review, research and investigation parallelize cleanly because agents do not overwrite each other's output. When writing is unavoidable, give each teammate its own files.
Message the session that holds the context
Teammates can message each other directly, and you can interact with any of them rather than only the lead, per the glossary. Send a question to the teammate whose task holds the answer when it concerns that task's details, such as why a test fails in its module. Send it to the lead when it changes scope, priorities or a decision that crosses tasks, because the lead assigns work and synthesizes results. Routing everything through the lead wastes its context, and routing cross-task decisions around it leaves the plan out of date.
Treat it as experimental
Agent teams are experimental and disabled by default, according to the agent teams docs. Documented gaps include teammates that are not restored on resume, lagging task status and slow shutdown. Plan sessions you can finish in one sitting and verify the task list yourself before trusting it. Build workflows that still hold up if the feature changes.
Prove the speedup before scaling
Anthropic flags diminishing returns: adding teammates does not speed work up proportionally (docs). A community benchmark report went further, claiming agent teams cost 73-124% more than sequential runs with zero quality gain, though that is unreviewed practitioner evidence. Run the same task once with a single session and once with a small team before making teams your default. Keep the team only if it finishes faster or catches things the single session missed.
Steps
-
Check that the task splits Before enabling anything, write the objective in one sentence and list the pieces it breaks into. Mark which pieces mostly read and which write, and which touch the same files. If most pieces share files or wait on one decision, stay with a single session or subagents. If you can name several independent slices, a team is worth trying, and choosing agent teams versus subagents covers the full decision.
-
Enable the feature Set CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 in your shell or settings.json; without it Claude Code will not spawn or propose teammates, per the agent teams docs. Confirm your Claude Code version is recent enough, since a practitioner comparison cites v2.1.32 or later. Start a fresh session so the team is set up at session start. If the lead never offers teammates, check the variable first.
-
Plan the task list with the lead Ask the lead for a planning pass that produces tasks with clear inputs, outputs and dependencies before any teammate starts. Review the list yourself and push back on tasks that are vague or overlap. If the lead creates too few tasks, ask it explicitly to split the work into smaller pieces. The task decomposition skill covers how to test each subtask.
A good sign is that you could hand any single task to a stranger and they would know when it is done.
-
Define teammate roles and file ownership Give each teammate one focus, a set of files or modules it owns, and a list of areas it must leave alone. Keep review roles separate from implementation so the check stays independent. Where teammates must interoperate, agree interfaces or stubs before parallel work begins. The roles and scopes skill has a role catalog and sizing guidance.
-
Run, monitor and route messages Let teammates claim tasks and watch the shared task list and message traffic rather than each terminal. Answer task-specific questions by messaging the teammate directly, and send scope changes to the lead. Look for tasks that seem finished but are still open, because task status can lag and block dependents. See managing shared task state and coordinating inter-agent communication for the details.
-
Synthesize, validate and shut down When tasks complete, have the lead combine changes, resolve conflicts and run checks that span modules. Review the integrated result yourself before accepting it, because the lead's summary is not a test run. Shut teammates down deliberately and allow time, since they finish their current request or tool call first (docs). Record the tokens spent and whether the team beat a single session, and see reviewing and synthesizing teammate outputs for the acceptance review.
When to Use
- A codebase review or audit that splits by module, because the launch guidance names independent, read-heavy work as the best fit and reviewers do not edit shared files.
- Debugging with several competing hypotheses, since teammates can pursue separate theories in their own context windows and message each other when their evidence overlaps.
- A feature that divides cleanly into backend, frontend and tests living in separate directories, so each teammate owns distinct files and the lead integrates at the end.
- Research across a repository too large to reason about in one context window, because each teammate loads only its slice and reports back.
- Work where workers genuinely need to ask each other questions mid-task, which subagents cannot do because they report only to the parent session.
When Not to Use
- Small or strictly sequential changes a single session can finish, because each teammate adds token cost and coordination overhead with no parallelism to gain.
- Write-heavy refactors where several agents would edit the same files, since teammates share one working tree and concurrent edits collide.
- Long-running work you expect to pause and resume later, because in-process teammates are not restored after /resume or /rewind.
- Budget-constrained tasks where token spend matters more than wall-clock time, since costs scale with the number of teammates.
- Tasks that hinge on one design decision nobody has made yet, because parallel workers will each guess differently and force rework.
Skills in this method
Each skill is a self-contained write-up your agent can run. Install the ones you need; nothing here is a bundle.
Claude Code Agent Teams vs Subagents: Choosing Well
A short written decision naming the execution mode (single session, subagents or agent team), the reason, the expected cost, and the condition that would make you switch.
npx skills add gethamster/skills --skill choosing-agent-teams-versus-subagents --agent claude-code --yesClaude Code Agent Teams Inter-Agent Communication
Teammates exchange the information they need at their shared boundaries, stay in sync with the task list, and do not stall or duplicate work.
npx skills add gethamster/skills --skill coordinating-inter-agent-communication --agent claude-code --yesClaude Code Agent Teams Task Decomposition, Step by Step
A reviewed task list in which every task has owned files, explicit inputs, a checkable output and mapped dependencies, sized to the team.
npx skills add gethamster/skills --skill decomposing-tasks-for-agent-teams --agent claude-code --yesClaude Code Agent Teams Designing Roles, Step by Step
A written team roster in which every teammate has one focus, a named owned area, explicit exclusions and a defined output, with review kept separate from implementation.
npx skills add gethamster/skills --skill designing-agent-roles-and-scopes --agent claude-code --yesClaude Code Agent Teams Shared Task List Management
A shared task list that accurately shows what is pending, blocked, claimed and done, so dependent tasks start as soon as their prerequisites are truly finished.
npx skills add gethamster/skills --skill managing-shared-task-state --agent claude-code --yesRunning Claude Code Agent Teams Parallel Tasks Safely
A task list where every concurrent task owns distinct files, every dependency is sequenced, and every spawned teammate has enough independent work to justify its cost.
npx skills add gethamster/skills --skill parallelizing-independent-work-across-sessions --agent claude-code --yesClaude Code Agent Teams Reviewing Synthesizing Results
A single integrated, validated change with every conflict resolved and a clear accept, rework or reject decision recorded for each task.
npx skills add gethamster/skills --skill reviewing-and-synthesizing-teammate-outputs --agent claude-code --yesFAQ
Are Claude Code Agent Teams generally available?
No. Anthropic introduced them as a research preview in the Claude Opus 4.6 announcement, and the docs still describe them as experimental and disabled by default. You turn them on with the CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS environment variable. Expect behavior and limits to change.
How do agent teams differ from subagents?
Subagents run inside a single session and report only to the parent, while teammates are separate sessions with their own context windows, according to the Claude Code glossary. Teammates share a task list and can message each other and you directly. Subagents are cheaper and simpler for focused lookups. Teams fit work where the workers need to coordinate among themselves.
Do agent teams cost more than a single session?
Yes. Anthropic states that token costs scale linearly with the number of teammates because each has its own context window (docs). A community benchmark post claims teams cost 73-124% more than sequential runs with no quality gain, though it is not peer reviewed. Measure on your own tasks before committing.
Should I message a teammate or the lead?
You can message any teammate directly, as the glossary notes. Go to the teammate when the question is about its own task, such as a failing test in its module. Go to the lead when the answer changes scope, priorities or anything another teammate depends on, so the task list stays accurate.
What happens to my team if I resume a session?
In-process teammates are not restored by /resume or /rewind, per the agent teams docs. After resuming, the lead may try to message teammates that no longer exist. Tell it to spawn new teammates for unfinished tasks. Plan team sessions so they can finish without a pause where possible.
Is there evidence that agent teams produce better code?
Not yet in any independent, controlled form. The only benchmark in circulation is a practitioner Reddit report claiming zero quality gain at higher cost. Anthropic's usage analysis of roughly 400,000 sessions does not break out agent teams, so adoption is also unmeasured. Treat quality gains as something to verify on your own work.
Related methods
gstack Framework: Garry Tan's Claude Code Skill Pack
The gstack framework is Garry Tan's open-source Claude Code skill pack that runs a sprint as slash commands: features, setup, gstack vs other frameworks.
Context Engineering for Claude AI: The 6 Pillars
How the six-pillar context engineering framework for Claude AI works, where it came from, what the evidence shows, and when to use it.
What Is Jev Engineering? A Guide for Agent Builders
Jev Engineering splits AI agents into an LLM that writes, a decision model that decides and code that acts. Learn its origin, evidence and limits.
What Is the State–Questions–Action–Verify Loop?
The State–Questions–Action–Verify Loop: build state, ask typed questions, act, then verify with evidence. Origins, evidence, limits and steps.
Download the Claude Code Agent Teams: Lead, Teammates and Limits pack
One zip with the whole method, to read offline or drop into a repository:
METHOD.md, this write-up in full- 7 skill folders, each with its
SKILL.mdand the references it ships - MIT licensed, the same files the commands above install
Source: gethamster/skills on GitHub, MIT licensed.
Install the skills
Claude Code Agent Teams: Lead, Teammates and Limits is a write-up of how the method works, so there is nothing to install for the method itself. Its skills are what your agent runs, and each one installs separately. The section above carries the Claude Code command for every skill, and each skill's own page carries the commands for Cursor, Codex, and Antigravity.
Or browse the skills and pick interactively:
npx skills add gethamster/skillsSource: gethamster/skills on GitHub, MIT licensed.