Guides

Claude Code Subagents: Haiku vs Sonnet vs Opus, Benchmarked

Scorecards for Haiku 5.5, Sonnet 5.5 and Opus 5.5: 14, 4 and 0 of 18 agents, with quality, perfect scores, median time and an effort heatmap for each

If every Claude Code subagent you spawn runs on the model of your main session, most of your delegated work runs on a model it does not need. We ran Haiku 5.5, Sonnet 5.5 and Opus 5.5 through the same 23 kinds of delegated work, from web research and CSV analysis to n8n workflow edits and client emails, and graded every output against answers fixed before the run.

Over 458 graded runs in three rounds, 14 of the 18 agent definitions that came out of it moved to Haiku 5.5. Four stay on Sonnet 5.5, and each of them writes judgment prose: a client email, a build spec, a translation, a fact check. None needs Opus 5.5. The savings come from two places and multiply. Haiku 5.5 lists at $0.10 per million input tokens against $2 for Sonnet 5.5, a twentieth of the price, and in round 1 a narrow tool list cut tokens about 3 to 5 times on each of the three models.

Claude Code delegation benchmark, 2026-09-29 to 2026-10-08
Where 18 agent definitions landed
458graded runs over three rounds
14 of 18agent definitions now on Haiku 5.5
20xlower input price for Haiku 5.5 than Sonnet 5.5
3 to 5xfewer tokens with a narrow tool list, on every model
Haiku 5.5: 14Sonnet 5.5: 4Opus 5.5: 0

How the benchmark worked

We built the benchmark while setting up Claude Code delegation for a client, and ran it between 2026-09-29 and 2026-10-08. Each task used the prompt we would send in real work, through the agent type we would use for it. Ground truth was fixed before the runs: planted bugs, decoys, a missing title as a hallucination trap, and counts computed by script. The 458 runs count Haiku 5.5 only; the 27 runs of the retired Haiku 4.5 are left out.

File and live-system tasks were graded by script or by re-reading the live object, meaning hidden tests, a diff against untouched fixtures, and each n8n workflow fetched back from the instance. Agent self-reports did not count. Five prose tasks (a debug write-up, a status update, a client email, a spec and a translation) were scored against a written rubric. Every prompt, output and hand-back report is saved, so any grade can be re-read.

How the 458 runs split
Three rounds, then the whole suite again on Haiku 5.5
Round 1, 2026-09-29
79runs
Three models, 23 task types
23 Opus 5.5, 28 Sonnet 5.5 and 28 Haiku 5.5 runs, with repeats where a pick hinged on it.
Round 2
99runs
Effort grid
Eight task types at low, medium and high effort, plus one run of each production definition.
Round 3
280runs
Five new fixtures per task type
Each through its production definition, with the runner-up setting where a decision hinged on it.
Rerun, 2026-10-08
167of the 458
Haiku 5.5, full suite
Every Sonnet 5.5 cell in all three rounds gets a Haiku 5.5 cell beside it.

Round 1: Opus 5.5 scored perfect, and 17 task types were a three-way tie

In round 1, Opus 5.5 scored 1.0 on all 23 task types. Sonnet 5.5 and Haiku 5.5 each averaged 0.98 with 18 perfect scores. Seventeen task types finished at 1.0 for all three models, so the contest came down to six.

ModelQualityPerfectMedian timeInput price
Opus 5.51.0023 of 2338.5 s$4.00/M
Sonnet 5.50.9818 of 2319.6 s$2.00/M
Haiku 5.50.9818 of 2323.3 s$0.10/M

Scores are the mean where a task type ran more than once. The Haiku 5.5 column comes from the 2026-10-08 rerun, whose parent session loaded more tools, which adds to its time per run. Opus 5.5 took about twice as long per run as Sonnet 5.5 and costs twice as much per input token; Haiku 5.5 costs a twentieth of Sonnet 5.5 for prompts up to 100,000 tokens ($0.50 per million above that).

Task typeOpus 5.5Sonnet 5.5Haiku 5.5
n8n debug from an execution error1.000.930.93
Code review, planted bugs1.001.000.95
Memory audit, planted defects1.000.901.00
Status update in a set house style1.000.950.95
n8n build spec1.000.950.90
EN to PL translation1.000.900.90

On round 1 alone, Opus 5.5 wins the hard work. Rounds 2 and 3 tested whether the lead survives repeats, effort settings and new inputs.

Round 2: the Opus 5.5 lead did not survive repeats and effort settings

Round 2 took eight task types from the original round 1 run, when the third model was still Haiku 4.5: the seven that split then, plus bulk edit as a control. Against the Haiku 5.5 column above, CSV analysis and the client email tied and code review split, so the two sets differ. Each ran at low, medium and high effort on each model. It aimed at the four task types behind the original Opus 5.5 pick: n8n debugging, the memory audit, the translation and the build spec. Sonnet 5.5 scored 1.0 on the n8n debug at every effort once a fixture defect was fixed (the error log contradicted one planted bug), was clean on the memory audit, and matched Opus 5.5 on the translation at high effort. The build spec fell last. On the bare prompts in the heatmap below, every model scored 0.85 to 0.92, and over 7 Opus and 9 Sonnet runs the two sat within 0.05 of each other. Through the build-spec-writer definition, which adds a coverage checklist, both scored 1.0 on the same fixture.

Round 2 and the 2026-10-08 rerun
Quality by model and effort on the eight round 2 task types
Sonnet 5.5
Opus 5.5
Haiku 5.5
Low
Med
High
Low
Med
High
Low
Med
High
n8n debug
1.00
1.00
1.00
1.00
1.00
1.00
1.00
0.80
1.00
CSV analysis, 864 rows
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
Memory audit
1.00
1.00
1.00
0.93
1.00
1.00
1.00
1.00
1.00
Status update
0.98
0.95
1.00
1.00
1.00
1.00
1.00
0.95
1.00
Client email
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
n8n build spec
0.85
0.90
0.85
0.92
0.88
0.90
0.85
0.90
0.90
EN to PL translation
0.60
0.87
1.00
0.65
1.00
1.00
0.80
1.00
0.95
Bulk edit, 30 files
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
0.95 or better0.80 to 0.94below 0.80
Score from 0 to 1 against ground truth fixed before the run. One run per cell, three where a pick hinged on it (the cell shows the mean).

Low effort held on everything mechanical or scripted. It failed outright only on the Polish translation, where Sonnet 5.5 and Opus 5.5 both broke case grammar around the English terms the text had to keep, on every low run. Opus 5.5 at low also skipped a dead index line in the memory audit once in three runs. High effort cost Sonnet 5.5 21% more tokens and doubled the wall time of Opus 5.5, for the same scores on six of the eight task types.

Haiku 5.5 ran the same grid on 2026-10-08. It averaged 0.956 at low and at medium and 0.981 at high, level with Sonnet 5.5 at high. At low effort it lost points only on the translation (0.80) and the build spec (0.85).

Speed against quality
Haiku 5.5 at low effort is the fastest setting that scores above 0.95
Mean quality against median seconds per run for nine model and effort settings0.920.940.960.981.0010s15s20s25s30s35sMedian seconds per run (faster is left)Mean qualitySonnet 5.5, low effort: quality 0.929, median 15.5 slowSonnet 5.5, medium effort: quality 0.965, median 20.5 smedSonnet 5.5, high effort: quality 0.981, median 23.9 shighOpus 5.5, low effort: quality 0.938, median 23.3 slowOpus 5.5, medium effort: quality 0.984, median 27.5 smedOpus 5.5, high effort: quality 0.988, median 32.3 shighHaiku 5.5, low effort: quality 0.956, median 17.3 slowHaiku 5.5, medium effort: quality 0.956, median 23.3 smedHaiku 5.5, high effort: quality 0.981, median 29.4 shigh
Haiku 5.5Sonnet 5.5Opus 5.5
Mean over the eight effort-grid task types. Haiku 5.5 ran on 2026-10-08 through a parent session that loaded more tools, which slows it and adds tokens for reasons unrelated to the model.

At list input prices, a Haiku 5.5 run in the effort grid cost about one cent, against 12 to 15 cents for Sonnet 5.5 and 27 to 33 cents for Opus 5.5. Haiku 5.5 ran through the heavier parent session, so its token counts, and its cost, are if anything overstated. The table counts input tokens only, the measure the benchmark used to rank models; output tokens add to every column.

ModelCost, low effortMediumHighTokens per run
Haiku 5.5$0.009$0.009$0.01087k to 100k
Sonnet 5.5$0.125$0.147$0.13862k to 74k
Opus 5.5$0.274$0.296$0.32768k to 82k

Round 3: five new inputs for every task type

A score on one input can be luck. Round 3 gave every task type five new fixtures (new workflows, transcripts, notes, diffs, question sets and payloads) and ran each through its production agent definition, with the runner-up setting alongside on the ten cells where a decision hinged on it. Seventeen of the 23 picks scored 1.0 on all five variants.

Six slipped: the wrap once, the code review once, and the four judgment tasks (debugging, status update, spec and translation). Those four slipped the same way at the runner-up setting, so the setting was not the cause. Four judgment definitions that had been set to medium effort tied with low across all five variants, and now run at low.

Haiku 5.5 then took the same five variants through the same definitions with only the model switched. It tied or beat the Sonnet 5.5 pick on 19 task types and trailed on four.

Task typeSonnet 5.5 pickHaiku 5.5Where Haiku 5.5 lost points
Client email1.000.86Unsupported "sole blocker" wording on two variants, and one next step left out
EN to PL translation0.890.82Glosses and stiff phrasing; one variant put a Polish case ending on a kept English noun
Code review0.990.91One run missed an un-awaited call and reported a decoy as a bug (scored 0.55)
Fact check1.000.95A claim called unsupported in the prose but left off the list of wrong claims

It edged the Sonnet 5.5 pick on n8n debugging (0.94 against 0.92) and the status update (0.98 against 0.96), and scored 1.0 on every variant of 13 mechanical definitions. Through the production definitions its median run took 15 seconds, against 16 for Sonnet 5.5. The client email, the translation and the fact check stay on Sonnet 5.5 for the losses above. The build spec stays on Sonnet 5.5 as a judgment call: over the five variants Sonnet 5.5 and Opus 5.5 averaged 0.89 and Haiku 5.5 0.88, a gap inside one grader judgment.

A narrow tool list cuts tokens on every model

In round 1, no general-purpose subagent ran below 69,000 tokens on Opus 5.5 or Sonnet 5.5, or below 86,000 on Haiku 5.5, even for a one-line answer, because they load every tool and the project's CLAUDE.md. The four restricted agents (wrap writer, code investigator, builder and reviewer) used 14,500 to 28,000 on each of the three models, about 3 to 5 times fewer. The production definitions range wider, from 12,000 to 64,000 tokens; the three above 28,000 load connectors, a skill or web search.

Tokens and cost per run
Restricted agents use a third to a fifth of the tokens
Round 1, on each model
Opus 5.5, general-purpose70k
Opus 5.5, restricted15 to 24k
Sonnet 5.5, general-purpose69k
Sonnet 5.5, restricted15 to 23k
Haiku 5.5, general-purpose86k
Haiku 5.5, restricted16 to 28k
Production definitions, one run eachTokensEst. cost
html-fragment-writer12k$0.0012
status-writer13k$0.0013
transcript-extractor13k$0.0013
web-researcher14k$0.0014
fact-checker14k$0.0280
pl-translator14k$0.0282
memory-auditor17k$0.0017
n8n-inspector18k$0.0018
feature-builder19k$0.0019
web-extractor20k$0.0020
vault-researcher26k$0.0026
n8n-editor28k$0.0028
connector-reader38k$0.0038
client-email-writer45k$0.0901
build-spec-writer64k$0.1270
Round 1: the lowest general-purpose run on each model against the range of the four restricted agents (wrap writer, code investigator, builder, reviewer). Haiku 5.5 ran on 2026-10-08 through a heavier parent session. Production definitions: one graded run each on its own task. Estimated cost is tokens times the list input price of the model the definition runs on now ($0.10 per million for Haiku 5.5, $2 for Sonnet 5.5); output tokens are not included. Grey bars run on Sonnet 5.5.

Each definition is a Markdown file in the .claude/agents/ folder of your home directory or of a project, and its frontmatter fixes the three things the benchmark measured: model, effort and tools.

One agent definition
.claude/agents/pl-translator.md
--- name: pl-translator description: EN to PL game text with tokens kept... model: sonnet effort: high tools: Read, Write ---
modelSonnet 5.5, because on five new fixtures in round 3 Haiku 5.5 scored 0.82 against 0.89.
effortHigh scored 1.0 on all three runs; medium slipped on two of three (a typo, a dropped verb).
toolsTwo tools. Its production run used 14k tokens, against 69k or more for any general-purpose run in round 1.

The tool list does most of the work. A translator needs Read and Write. A CSV analyst gets Bash so it scripts its counts instead of estimating them. An n8n inspector gets read-only n8n tools and nothing that can change a workflow. To put a monthly figure on your own agent volumes, the AI agent cost estimator takes runs per day and tokens per run.

The 18 agent definitions

Every definition below has at least one graded run on its own task at its own effort, and every one ran on five new fixtures in round 3. Filter by model to see which jobs stayed on Sonnet 5.5.

The production roster after 2026-10-08
18 agent definitions, each with its model, effort and tool list
Web
web-researcherWeb lookup and fresh web research4 tools: WebSearch, WebFetch, Bash, ReadHaiku 5.5Low
web-extractorWeb fetch and exact extraction, 60 items4 tools: WebFetch, Bash, Read, WriteHaiku 5.5Low
n8n
n8n-inspectorn8n workflow inspection12 tools: Read + n8n (11 read tools)Haiku 5.5Low
n8n-editorn8n workflow edit and save13 tools: Read + n8n (12 tools)Haiku 5.5Low
n8n-debuggern8n debug from an execution error10 tools: Read, Grep, Glob + n8n (7 read tools)Haiku 5.5Low
build-spec-writern8n build spec5 tools: Read, Write, Skill, WebSearch, WebFetchSonnet 5.5Medium
Code and data
feature-builderMulti-file feature with self-test6 tools: Read, Edit, Write, Glob, Grep, BashHaiku 5.5Low
csv-analystCSV analysis, 864 rows3 tools: Bash, Read, WriteHaiku 5.5Low
html-fragment-writerHTML fragment to a contract2 tools: Read, WriteHaiku 5.5Low
Session ops
vault-researcherQ&A over a notes vault, with citations3 tools: Read, Grep, GlobHaiku 5.5Low
memory-auditorMemory audit, planted defects3 tools: Read, Grep, GlobHaiku 5.5Low
bulk-editorBulk edit, 30 files6 tools: Read, Edit, Write, Glob, Grep, BashHaiku 5.5Low
Writing
status-writerStatus update in a set house style2 tools: Read, WriteHaiku 5.5High
client-email-writerClient email in brand voice3 tools: Read, Write, SkillSonnet 5.5Low
fact-checkerFact check against a source2 tools: Read, GrepSonnet 5.5Low
pl-translatorEN to PL translation2 tools: Read, WriteSonnet 5.5High
transcript-extractorMeeting transcript to actions and decisions1 tool: ReadHaiku 5.5Low
Connectors
connector-readerRead-only lookups in connected tools17 tools: Read + 3 connected servicesHaiku 5.5Low

Four more task types run on existing agents outside this list: wrap apply on a wrap writer, and code locate, surgical fix and code review on investigator, builder and reviewer agents called with model: sonnet.

How Haiku 5.5 fails, and how Haiku 4.5 did

Haiku 4.5 ran round 1 first. It estimated counts instead of computing them (three different wrong answers to one CSV question), invented n8n nodes and a client commitment, broke a house rule against em dashes in 3 of its 27 runs, and took up to 146 seconds on the CSV task. Its figures are retired.

Haiku 5.5 scripts the same CSV answers in 17 seconds, invented no nodes and carried every fact into the client email. Three of its runs flagged fixture defects nobody had asked about, among them weekday names in a transcript that did not match the 2026 calendar. What it still loses is judgment in prose, which is why the four writing-heavy definitions stay on Sonnet 5.5.

Set up your own Claude Code subagents this way

  1. Write one agent definition per task type, with a model, an effort level and a tool list no wider than the job.
  2. Start each definition on Haiku 5.5 at low effort.
  3. Move a definition to Sonnet 5.5 only when a graded run shows a loss in judgment, and try a higher effort first.
  4. Write the checklist into the definition before you reach for Opus 5.5. A coverage checklist brought Sonnet 5.5 level with Opus 5.5 on the build spec.
  5. Grade against answers fixed before the run, by script wherever the answer can be computed.
  6. Re-test on new inputs. Five new fixtures per task type moved four settings down and none up.

What this benchmark does not prove

Every grade came from a Claude model. A second, blind Sonnet 5.5 grader re-scored all 48 prose outputs without knowing which model wrote them, and 44 of 48 landed within 0.1 of the original grade (mean gap 0.03), with no pick moving. No grader from outside the Claude family was available.

Rounds 1 and 2 hold one to three runs per cell on a single fixture. Haiku 5.5 ran the full suite on one day through a parent session that loaded more tools than on 2026-09-29, so its token figures do not compare with the other columns. The task mix is one client's delegated work, so a team whose subagents mostly write prose would see more of it land on Sonnet 5.5. Agents in production covers the checks that sit around a model once it ships.

Frequently asked questions

Which model should Claude Code subagents use?

In this benchmark, Haiku 5.5 at low effort handled the mechanical and scripted work, which covered 14 of 18 agent definitions, and Sonnet 5.5 handled judgment prose such as client emails, specs, translation and fact checks. No task type needed Opus 5.5, and Haiku 5.5 lists at a twentieth of the Sonnet 5.5 input price.

Does the effort level matter for subagents?

On a few task types only. Low effort held on every mechanical task. The Polish translation needed high effort on Sonnet 5.5 to keep its case grammar, and high effort cost Sonnet 5.5 21% more tokens for the same scores on six of eight task types.

Why do general-purpose subagents use so many tokens?

They load every available tool and the project's CLAUDE.md on every call. In round 1 none ran below 69,000 tokens, while the four restricted agents used 14,500 to 28,000 on each model.

How do you set the model for a Claude Code subagent?

Add model, effort and tools to the frontmatter of the agent's Markdown file in .claude/agents/, either in your home directory or inside a project. The main session then delegates to that agent by name.

When is Opus 5.5 worth it for a subagent?

It was not worth it anywhere in this suite once the agent definition carried a checklist. Opus 5.5 led on the first run, and the lead did not survive repeats, effort settings and new fixtures.

If you run agents on n8n or delegate code work to Claude Code, Ovidius builds agent workflows on n8n and keeps them running under AI managed services. Book a discovery call to scope yours.

Put your first workflow on the board

Answer a few questions about your team and what you want automated, then Jason scopes it with you on the call.

Owen, Maciej, Jason, Oskar and Ben