
If every Claude Code subagent you spawn runs on the model of your main session, most of your delegated work runs on a model it does not need. We ran Haiku 5.5, Sonnet 5.5 and Opus 5.5 through the same 23 kinds of delegated work, from web research and CSV analysis to n8n workflow edits and client emails, and graded every output against answers fixed before the run.
Over 458 graded runs in three rounds, 14 of the 18 agent definitions that came out of it moved to Haiku 5.5. Four stay on Sonnet 5.5, and each of them writes judgment prose: a client email, a build spec, a translation, a fact check. None needs Opus 5.5. The savings come from two places and multiply. Haiku 5.5 lists at $0.10 per million input tokens against $2 for Sonnet 5.5, a twentieth of the price, and in round 1 a narrow tool list cut tokens about 3 to 5 times on each of the three models.
We built the benchmark while setting up Claude Code delegation for a client, and ran it between 2026-09-29 and 2026-10-08. Each task used the prompt we would send in real work, through the agent type we would use for it. Ground truth was fixed before the runs: planted bugs, decoys, a missing title as a hallucination trap, and counts computed by script. The 458 runs count Haiku 5.5 only; the 27 runs of the retired Haiku 4.5 are left out.
File and live-system tasks were graded by script or by re-reading the live object, meaning hidden tests, a diff against untouched fixtures, and each n8n workflow fetched back from the instance. Agent self-reports did not count. Five prose tasks (a debug write-up, a status update, a client email, a spec and a translation) were scored against a written rubric. Every prompt, output and hand-back report is saved, so any grade can be re-read.
In round 1, Opus 5.5 scored 1.0 on all 23 task types. Sonnet 5.5 and Haiku 5.5 each averaged 0.98 with 18 perfect scores. Seventeen task types finished at 1.0 for all three models, so the contest came down to six.
| Model | Quality | Perfect | Median time | Input price |
|---|---|---|---|---|
| Opus 5.5 | 1.00 | 23 of 23 | 38.5 s | $4.00/M |
| Sonnet 5.5 | 0.98 | 18 of 23 | 19.6 s | $2.00/M |
| Haiku 5.5 | 0.98 | 18 of 23 | 23.3 s | $0.10/M |
Scores are the mean where a task type ran more than once. The Haiku 5.5 column comes from the 2026-10-08 rerun, whose parent session loaded more tools, which adds to its time per run. Opus 5.5 took about twice as long per run as Sonnet 5.5 and costs twice as much per input token; Haiku 5.5 costs a twentieth of Sonnet 5.5 for prompts up to 100,000 tokens ($0.50 per million above that).
| Task type | Opus 5.5 | Sonnet 5.5 | Haiku 5.5 |
|---|---|---|---|
| n8n debug from an execution error | 1.00 | 0.93 | 0.93 |
| Code review, planted bugs | 1.00 | 1.00 | 0.95 |
| Memory audit, planted defects | 1.00 | 0.90 | 1.00 |
| Status update in a set house style | 1.00 | 0.95 | 0.95 |
| n8n build spec | 1.00 | 0.95 | 0.90 |
| EN to PL translation | 1.00 | 0.90 | 0.90 |
On round 1 alone, Opus 5.5 wins the hard work. Rounds 2 and 3 tested whether the lead survives repeats, effort settings and new inputs.
Round 2 took eight task types from the original round 1 run, when the third model was still Haiku 4.5: the seven that split then, plus bulk edit as a control. Against the Haiku 5.5 column above, CSV analysis and the client email tied and code review split, so the two sets differ. Each ran at low, medium and high effort on each model. It aimed at the four task types behind the original Opus 5.5 pick: n8n debugging, the memory audit, the translation and the build spec. Sonnet 5.5 scored 1.0 on the n8n debug at every effort once a fixture defect was fixed (the error log contradicted one planted bug), was clean on the memory audit, and matched Opus 5.5 on the translation at high effort. The build spec fell last. On the bare prompts in the heatmap below, every model scored 0.85 to 0.92, and over 7 Opus and 9 Sonnet runs the two sat within 0.05 of each other. Through the build-spec-writer definition, which adds a coverage checklist, both scored 1.0 on the same fixture.
Low effort held on everything mechanical or scripted. It failed outright only on the Polish translation, where Sonnet 5.5 and Opus 5.5 both broke case grammar around the English terms the text had to keep, on every low run. Opus 5.5 at low also skipped a dead index line in the memory audit once in three runs. High effort cost Sonnet 5.5 21% more tokens and doubled the wall time of Opus 5.5, for the same scores on six of the eight task types.
Haiku 5.5 ran the same grid on 2026-10-08. It averaged 0.956 at low and at medium and 0.981 at high, level with Sonnet 5.5 at high. At low effort it lost points only on the translation (0.80) and the build spec (0.85).
At list input prices, a Haiku 5.5 run in the effort grid cost about one cent, against 12 to 15 cents for Sonnet 5.5 and 27 to 33 cents for Opus 5.5. Haiku 5.5 ran through the heavier parent session, so its token counts, and its cost, are if anything overstated. The table counts input tokens only, the measure the benchmark used to rank models; output tokens add to every column.
| Model | Cost, low effort | Medium | High | Tokens per run |
|---|---|---|---|---|
| Haiku 5.5 | $0.009 | $0.009 | $0.010 | 87k to 100k |
| Sonnet 5.5 | $0.125 | $0.147 | $0.138 | 62k to 74k |
| Opus 5.5 | $0.274 | $0.296 | $0.327 | 68k to 82k |
A score on one input can be luck. Round 3 gave every task type five new fixtures (new workflows, transcripts, notes, diffs, question sets and payloads) and ran each through its production agent definition, with the runner-up setting alongside on the ten cells where a decision hinged on it. Seventeen of the 23 picks scored 1.0 on all five variants.
Six slipped: the wrap once, the code review once, and the four judgment tasks (debugging, status update, spec and translation). Those four slipped the same way at the runner-up setting, so the setting was not the cause. Four judgment definitions that had been set to medium effort tied with low across all five variants, and now run at low.
Haiku 5.5 then took the same five variants through the same definitions with only the model switched. It tied or beat the Sonnet 5.5 pick on 19 task types and trailed on four.
| Task type | Sonnet 5.5 pick | Haiku 5.5 | Where Haiku 5.5 lost points |
|---|---|---|---|
| Client email | 1.00 | 0.86 | Unsupported "sole blocker" wording on two variants, and one next step left out |
| EN to PL translation | 0.89 | 0.82 | Glosses and stiff phrasing; one variant put a Polish case ending on a kept English noun |
| Code review | 0.99 | 0.91 | One run missed an un-awaited call and reported a decoy as a bug (scored 0.55) |
| Fact check | 1.00 | 0.95 | A claim called unsupported in the prose but left off the list of wrong claims |
It edged the Sonnet 5.5 pick on n8n debugging (0.94 against 0.92) and the status update (0.98 against 0.96), and scored 1.0 on every variant of 13 mechanical definitions. Through the production definitions its median run took 15 seconds, against 16 for Sonnet 5.5. The client email, the translation and the fact check stay on Sonnet 5.5 for the losses above. The build spec stays on Sonnet 5.5 as a judgment call: over the five variants Sonnet 5.5 and Opus 5.5 averaged 0.89 and Haiku 5.5 0.88, a gap inside one grader judgment.
In round 1, no general-purpose subagent ran below 69,000 tokens on Opus 5.5 or Sonnet 5.5, or below 86,000 on Haiku 5.5, even for a one-line answer, because they load every tool and the project's CLAUDE.md. The four restricted agents (wrap writer, code investigator, builder and reviewer) used 14,500 to 28,000 on each of the three models, about 3 to 5 times fewer. The production definitions range wider, from 12,000 to 64,000 tokens; the three above 28,000 load connectors, a skill or web search.
Each definition is a Markdown file in the .claude/agents/ folder of your home directory or of a project, and its frontmatter fixes the three things the benchmark measured: model, effort and tools.
The tool list does most of the work. A translator needs Read and Write. A CSV analyst gets Bash so it scripts its counts instead of estimating them. An n8n inspector gets read-only n8n tools and nothing that can change a workflow. To put a monthly figure on your own agent volumes, the AI agent cost estimator takes runs per day and tokens per run.
Every definition below has at least one graded run on its own task at its own effort, and every one ran on five new fixtures in round 3. Filter by model to see which jobs stayed on Sonnet 5.5.
Four more task types run on existing agents outside this list: wrap apply on a wrap writer, and code locate, surgical fix and code review on investigator, builder and reviewer agents called with model: sonnet.
Haiku 4.5 ran round 1 first. It estimated counts instead of computing them (three different wrong answers to one CSV question), invented n8n nodes and a client commitment, broke a house rule against em dashes in 3 of its 27 runs, and took up to 146 seconds on the CSV task. Its figures are retired.
Haiku 5.5 scripts the same CSV answers in 17 seconds, invented no nodes and carried every fact into the client email. Three of its runs flagged fixture defects nobody had asked about, among them weekday names in a transcript that did not match the 2026 calendar. What it still loses is judgment in prose, which is why the four writing-heavy definitions stay on Sonnet 5.5.
Every grade came from a Claude model. A second, blind Sonnet 5.5 grader re-scored all 48 prose outputs without knowing which model wrote them, and 44 of 48 landed within 0.1 of the original grade (mean gap 0.03), with no pick moving. No grader from outside the Claude family was available.
Rounds 1 and 2 hold one to three runs per cell on a single fixture. Haiku 5.5 ran the full suite on one day through a parent session that loaded more tools than on 2026-09-29, so its token figures do not compare with the other columns. The task mix is one client's delegated work, so a team whose subagents mostly write prose would see more of it land on Sonnet 5.5. Agents in production covers the checks that sit around a model once it ships.
In this benchmark, Haiku 5.5 at low effort handled the mechanical and scripted work, which covered 14 of 18 agent definitions, and Sonnet 5.5 handled judgment prose such as client emails, specs, translation and fact checks. No task type needed Opus 5.5, and Haiku 5.5 lists at a twentieth of the Sonnet 5.5 input price.
On a few task types only. Low effort held on every mechanical task. The Polish translation needed high effort on Sonnet 5.5 to keep its case grammar, and high effort cost Sonnet 5.5 21% more tokens for the same scores on six of eight task types.
They load every available tool and the project's CLAUDE.md on every call. In round 1 none ran below 69,000 tokens, while the four restricted agents used 14,500 to 28,000 on each model.
Add model, effort and tools to the frontmatter of the agent's Markdown file in .claude/agents/, either in your home directory or inside a project. The main session then delegates to that agent by name.
It was not worth it anywhere in this suite once the agent definition carried a checklist. Opus 5.5 led on the first run, and the lead did not survive repeats, effort settings and new fixtures.
If you run agents on n8n or delegate code work to Claude Code, Ovidius builds agent workflows on n8n and keeps them running under AI managed services. Book a discovery call to scope yours.