Engineering notesBy Daniel ChenLast verified 17 min read

Do Agent Skills pay for themselves? A paired test of SKILL.md on real tasks

A SKILL.md file is cheap to write, and it costs context on every run where it loads. This note measures what it buys on work that depends on house rules, with Claude Opus 5.5 and Sonnet 5.5 in Claude Code. On release tasks, Opus followed every rule in 6 of 24 runs without a skill and in 24 of 24 with a short one; Sonnet went from 10 of 30 to 30 of 30.

  • agents
  • skills
  • evals
  • claude-code

If you give a coding agent a SKILL.md that describes how your team does a job, does it do that job better, and is the extra context worth paying for? The answer below comes from a paired experiment: the same 12 tasks in the same repository, run with no skill, a short skill, a long skill carrying the same rules, and a skill the agent wrote for itself, with every run scored by a deterministic checker. Claude Sonnet 5.5 ran the full grid. Claude Opus 5.5, the model Claude Code uses by default on the API and most plans, ran the release tasks with no skill and with the short skill. Claude Haiku 5.5, the small model some teams use for automated and background jobs, also ran the full grid; its results are reported for completeness, below the main ones.

TL;DR

  • Release tasks: without a skill, Claude Opus 5.5 followed every house rule in 6 of 24 runs and Claude Sonnet 5.5 in 10 of 30. With a short (about 270-word) skill: 24 of 24 and 30 of 30.
  • Migration tasks: no measurable effect. Sonnet passed 30 of 30 runs without a skill, and Opus 12 of 12 in a smaller exploratory check, because the conventions were visible in the existing files.
  • Cost: the short skill made runs cheaper, not dearer. On Sonnet, 14% less per run at API prices. On Opus, Claude Code's own estimate fell 25% per run; those runs used a Claude Pro subscription, so that is an estimate, not an API bill, and not directly comparable with the Sonnet figure. A skill five to seven times longer, with the same rules, passed just as often on Sonnet and cost 7% more than no skill.
  • Self-written skills hurt: the skill Sonnet wrote by studying the repository passed 0 of 30 release runs, mainly because it turned a habit seen in the history (tagging releases by hand) into an instruction and added "ask the user first" steps that stall a headless run. Opus was not tested with self-written or long skills.
  • Loading was not the problem: the relevant skill loaded in every skill run on every model, and in a separate test Sonnet loaded a skill for 20 of 20 prompts that needed one and for none of 20 that did not.
  • Haiku 5.5, for completeness: the small model showed the same pattern: 5 of 30 release runs passed without a skill and 29 of 30 with the short one, at 31% lower cost per run.

Tested with

Agent
Claude Code 2.1.293, headless (claude -p), project settings only, web tools off
Models
claude-opus-5-5: 6 release tasks, no skill and compact skill, 4 repeats, Claude Pro subscription. claude-sonnet-5-5 and claude-haiku-5-5: all 12 tasks, four conditions, 5 repeats, API
Runs
48 Opus runs and 480 Sonnet and Haiku runs (12 tasks × 4 conditions × 5 repeats × 2 models), plus 47 exploratory Opus runs on the API and 80 trigger-test runs
Run date
8 October 2026
Spend
US$23.17 in API usage for everything, including set-aside batches; the Opus subscription runs were covered by the Pro plan (Claude Code estimate: US$7.35)

Why measure this at all

An Agent Skill is a folder with a SKILL.md file: YAML front matter with a name and a description, then Markdown instructions. The agent sees only the name and description of every installed skill at startup, roughly 100 tokens each, and reads the body when it decides the skill is relevant. The specification recommends keeping that body under 5,000 tokens and 500 lines (Agent Skills specification).

Anthropic's own authoring guidance says to build evaluations before writing extensive documentation and to measure the agent's performance without the skill first. It also notes that "there is not currently a built-in way to run these evaluations" (Skill authoring best practices). So most teams who add skills are guessing.

The largest public measurement is SkillsBench: 87 tasks, 18 model and harness configurations, curated skills raising the average pass rate from 33.9% to 50.5%. Software engineering was one of the smaller gains (+11.6 points), compact skills did better than comprehensive ones, and skills the agent wrote for itself scored below no skills at all on all three configurations tested (Li et al., SkillsBench, 2026). Those are averages over other people's tasks. The question here is narrower and closer to day-to-day work: in one repository with house rules, what does a skill change, and what does it cost?

The setup

One repository, two kinds of job

The test repository, tillroll, is a small Python order service written for this experiment: a SQLite schema with four migrations, a manifest of migration hashes, a generated schema snapshot, a Keep-a-Changelog-style CHANGELOG.md and one metadata file per release. It has the kind of rules every team accumulates. Some can be read off the existing files, some cannot.

  • Six migration tasks, each phrased as a ticket ("Ticket OPS-437: orders can now carry a discount amount. Add a column for it"). House rules: file naming, a three-line header, reversible up and down sections, a manifest entry with the file's SHA-256, a regenerated schema snapshot, money as integer pence in a column ending _minor, index names like ix_orders__customer_id and unique indexes named ux_….
  • Six release tasks ("Cut the next release of tillroll"). House rules: the version bump follows the changelog sections (Changed means minor, Removed means major), entries move word for word, sections go in a fixed order with Security last, release candidates are written 1.5.0rc1, a release/X.Y.Z.toml file lists the new migrations, and nobody tags by hand because CI creates the tag after review.

The prompts never mention the conventions. Two release rules (CI owns tags; Changed means minor) cannot be inferred from the repository at all, and the history contains a v1.4.2 tag that suggests the opposite of the first. That is deliberate: it is the kind of knowledge a skill exists to carry, and the results should be read with it in mind.

Four conditions

Skill conditions. Both skills (migrations and releases) are installed in every skill condition, so the agent has to pick the relevant one.
ConditionWhat is installedSize per skill
No skillNothing–
CompactFast path, rules, out of scope268–291 words
ComprehensiveThe same rules and descriptions, plus background, worked examples, troubleshooting and FAQ1,458–2,195 words
Self-writtenWritten by the same model after studying the repository, one run per skill516–977 words

The self-written skills were produced with a plain prompt: study the repository and write an Agent Skill for adding a migration (or cutting a release) "following this project's conventions", based only on what is in the repository. This mirrors what many people do in practice: ask the agent to write its own skill.

Models and runs

Sonnet 5.5 and Haiku 5.5 ran every task in every condition, five repeats each, on the API. Opus 5.5 ran the six release tasks with no skill and with the compact skill, four repeats each, using a Claude Pro subscription signed in to Claude Code instead of an API key: the same prompts, skills, flags and checkers, one run at a time. Opus was not run on the migration tasks, the long skill or the self-written skills in that batch. An earlier exploratory check on the API covered Opus on all 12 tasks with no skill and the compact skill, two repeats each; the harness's spending guard stopped it one run short, at 47 of 48. Fable 5.1 was planned alongside Opus but did not run: on this Claude Pro account it draws on paid usage credits rather than the plan, and the first Fable run was refused before any tokens were used ("You're out of usage credits"), so no Fable results are reported.

How runs were isolated and scored

Each run gets a fresh copy of the repository with its own git history, a throwaway home directory with a placeholder git identity, and Claude Code started with --setting-sources project, so nothing from the host machine leaks in. Web tools are disabled; permissions are bypassed inside the copy. Every task has a checker with 10 to 14 checks (for a migration, for example: the file applies, does what was asked, reverts exactly, the manifest hash matches, the snapshot was regenerated). A run passes only if every check passes. A reference solution passes all 12 tasks and an untouched repository fails all of them.

Cost is computed from the token counts the API returns, multiplied by Anthropic's published prices on the run date (Anthropic pricing). Claude Code's own total_cost_usd is documented as a client-side estimate that can differ from the bill (Run Claude Code programmatically); it is recorded too, and it is the only cost figure available for the Opus subscription runs, which are reported as estimates.

The question, conditions, metrics and analysis were written down before the runs. Pass rates carry 95% Wilson intervals. Differences against no skill are computed per task and then averaged, with a 95% bootstrap interval over tasks, because the task, not the run, is the honest unit: five repeats of the same task are not five independent samples.

Results

Bar chart of the share of runs passing every check, for Opus 5.5, Sonnet 5.5 and Haiku 5.5. Release tasks: no skill, Opus 25%, Sonnet 33%, Haiku 17%; compact skill, Opus 100%, Sonnet 100%, Haiku 97%; comprehensive skill, Sonnet 100%, Haiku 100%; self-written skill, Sonnet 0%, Haiku 0%. Migration tasks, Sonnet and Haiku: no skill 100% and 97%; compact 100% and 93%; comprehensive 100% and 93%; self-written 83% and 73%. Opus was run on release tasks with no skill and the compact skill only.
Share of runs that passed every check. Opus 5.5: release tasks, no skill and compact only, 24 runs per bar (6 tasks × 4 repeats, Claude Pro subscription). Sonnet 5.5 and Haiku 5.5: 30 runs per bar (6 tasks × 5 repeats, API). Whiskers are 95% Wilson intervals.
Runs passing every check. Sonnet and Haiku: 30 runs per task family and condition; "All" pools 60 runs with a 95% Wilson interval. The last column is the change against no skill in percentage points, averaged over the tasks, with a 95% bootstrap interval over tasks. *Opus 5.5: release tasks only, 24 runs per row (Claude Pro subscription), so its change is averaged over the six release tasks.
ModelConditionRelease tasksMigration tasksAllvs no skill, points
Opus 5.5*No skill6/24not run––
Compact24/24not run–+75 (+42 to +100), release only
Sonnet 5.5No skill10/3030/3067% (54%–77%)–
Compact30/3030/30100% (94%–100%)+33 (+8 to +58)
Comprehensive30/3030/30100% (94%–100%)+33 (+8 to +58)
Self-written0/3025/3042% (30%–54%)−25 (−50 to 0)
Haiku 5.5No skill5/3029/3057% (44%–68%)–
Compact29/3028/3095% (86%–98%)+38 (+10 to +67)
Comprehensive30/3028/3097% (89%–99%)+40 (+10 to +68)
Self-written0/3022/3037% (26%–49%)−20 (−43 to 0)

Three things stand out.

Skills mattered where the rules were invisible. On releases, the compact skill took Opus from 6 to 24 passing runs out of 24 and Sonnet from 10 to 30 out of 30. Averaged over the six release tasks, that is +75 points for Opus (95% interval +42 to +100) and +67 for Sonnet (+33 to +100). Sonnet's gain is capped by a ceiling: it already passed two of the six release tasks every time without help. Opus is covered in more detail below.

Where the rules were visible, there was nothing to gain. Without a skill, Sonnet passed 30 of 30 migration runs, and Opus 12 of 12 in the exploratory check. The agents copied the header format, index naming and money convention from the four existing migrations. The skill neither helped nor got in the way.

Compact and comprehensive were indistinguishable on pass rate. 60 of 60 against 60 of 60 on Sonnet. This experiment cannot separate them on correctness. They did differ on cost. (Opus was not run with the comprehensive skill.)

Haiku 5.5. The small model, which some teams use for automated and background jobs rather than day-to-day development, followed the same pattern: 5 to 29 of 30 release runs with the compact skill (+80 points, +47 to +100), 29 of 30 migration runs without a skill, where its one- or two-run differences are well inside the noise, and 57 of 60 against 58 of 60 for compact and comprehensive.

Opus 5.5 on the release tasks

Opus 5.5 is the model Claude Code uses by default on the API and on most plans (Model configuration), so it is the one most readers will run. On the release tasks it did no better than Sonnet without help: 6 of 24 runs followed every house rule, 25% (12%–45%), against 10 of 30 for Sonnet. With the compact skill it passed 24 of 24, 100% (86%–100%). Averaged over the six tasks that is +75 points (95% interval +42 to +100), with 5 of 6 tasks improved and none worse.

Every failed Opus run broke exactly one rule: 14 of the 18 created a git tag, and the other 4 were the security release, with Security left above Fixed. Without a skill Opus passed the minor release in 4 of 4 runs and the major release in 2 of 4, and no run of the other four tasks. A stronger model does not learn a rule nobody wrote down.

The skill also made Opus cheaper and quicker. Claude Code's own estimate fell from 17.5 to 13.1 cents per run (25% less), output from 2,084 to 1,234 tokens, and the median run from 22.4 to 12.7 seconds, with runs one at a time. These runs used a Claude Pro subscription, so the figures are estimates, not a bill: the 48 runs add up to US$7.35, which the plan covered. The estimates are higher than the same token counts would cost at published API prices, so compare them with each other, not with the Sonnet and Haiku costs below. An earlier exploratory check on the API, with Opus on all 12 tasks and two repeats (47 runs; the spending guard stopped one short), points the same way at published prices: 3/11 release runs passed without the skill and 12/12 with it, migrations passed 12/12 either way, and mean cost fell from 11.2 to 7.9 cents per run (29% less).

Per-task passes (out of 5)
Passing runs per task, model and condition: out of 4 for Opus 5.5 (release tasks only, no skill and compact), out of 5 for Sonnet 5.5 and Haiku 5.5.
TaskOpus: No skill (of 4)Opus: Compact (of 4)Sonnet: No skillSonnet: CompactSonnet: ComprehensiveSonnet: Self-writtenHaiku: No skillHaiku: CompactHaiku: ComprehensiveHaiku: Self-written
mig-01-phone––55554555
mig-02-channel––55555555
mig-03-refunds––55555534
mig-04-created-index––55555555
mig-05-sku––55505550
mig-06-discount––55555353
rel-01-patch0405500550
rel-02-minor4455505550
rel-03-major2455500550
rel-04-changed0405500550
rel-05-rc0405500550
rel-06-security0405500450

What goes wrong without a skill

Horizontal bar chart of the rules broken in no-skill release runs. Creating a git tag by hand was the most common: 58% of Opus 5.5 runs, 50% of Sonnet 5.5 runs and 50% of Haiku 5.5 runs. Next came changelog sections out of house order (17%, 20% and 17%), then smaller shares for reworded entries, the changelog heading, compare links and the release file, none of which Opus broke.
Share of no-skill release runs that failed each check: 24 Opus 5.5 runs, 30 each for Sonnet 5.5 and Haiku 5.5. A run can fail several.

Without a skill, 14 of 24 Opus release runs and 15 of 30 Sonnet runs created a git tag. That is a reasonable thing to do in most repositories and the wrong thing in this one. The rest were rules an agent can only guess: Security left above Fixed because that is where the entry already was (every no-skill run of that task, on all three models), a release candidate cut without moving the changelog (Sonnet, 3 of 5 runs, two of them also without the release file), and on Haiku, changelog entries reworded while being moved (5 of 5 runs of the major-release task).

The failures were not sloppy work. In most of them the agent did a careful, defensible release that broke one rule it had no way of knowing. That is the case for a skill: it carries the decisions that are not in the code.

Cost, tokens and time

Grouped bar chart of cost in US cents at published API prices. Sonnet 5.5, mean cost per run: no skill 4.27, compact 3.69, comprehensive 4.58, self-written 4.77; cost per passing run 6.41, 3.69, 4.58, 11.44. Haiku 5.5, mean cost per run: 0.49, 0.34, 0.41, 0.49; cost per passing run 0.87, 0.36, 0.43, 1.33.
API runs: mean cost per run and cost per passing run (total cost divided by passing runs), in US cents at published prices. Opus 5.5 ran on a subscription; its estimated costs are in the Opus section above.
API runs, means per run (60 runs per row). Cost in US cents at published prices. Input tokens include cache writes and reads; most input is cached. Wall time is the median, with six runs in parallel on a shared machine. Opus 5.5 ran on a subscription and is reported separately above.
ModelConditionCost per runCost per passInput tokensOutput tokensTurnsWall s
Sonnet 5.5No skill4.276.4185k1,3465.013.8
Compact3.693.6978k1,0125.511.0
Comprehensive4.584.5882k1,0565.610.4
Self-written4.7711.4498k1,5356.316.8
Haiku 5.5No skill0.490.87159k3,78011.219.6
Compact0.340.36108k2,3819.711.3
Comprehensive0.410.43124k2,6369.912.7
Self-written0.491.33161k3,51111.417.9

Loading a skill costs tokens, but in this experiment a compact skill paid for itself within the same run. Per task, it cut Sonnet's cost by 0.58 cents a run (95% interval 0.40 to 0.75), 14%, mostly through shorter output (1,012 tokens per run instead of 1,346). Opus showed the same direction on Claude Code's estimates (25% less per run, above). On Haiku the agent explored less, using about 108,000 input tokens per run instead of 160,000 and 9.7 turns instead of 11.2, and cost fell by 0.15 cents a run (0.12 to 0.19), 31%.

The comprehensive skill was a different trade. On Sonnet it cost 0.31 cents more per run than no skill (0.15 to 0.45), about 7%, for the same pass rate as the compact one. On Haiku it was still cheaper than no skill, by less. Five to seven times the words bought nothing measurable here.

Cost per passing run is the number that matters if failures have to be redone: on Sonnet, 3.7 cents with the compact skill against 6.4 cents without one. Absolute amounts are small because the tasks are small; the ratios are the useful part.

Wall time moved in the same direction (Sonnet's median fell from 13.8 to 11.0 seconds with the compact skill, Haiku's from 19.6 to 11.3), but treat timings as rough: API runs were six at a time on a shared machine.

Self-written skills: confident and wrong

Sonnet and Haiku each wrote a tidy, plausible skill (Opus was not tested in this condition). Both sets failed every release run: 0 of 30 on Sonnet and 0 of 30 on Haiku, against 10 and 5 with no skill. Reading the skills and transcripts shows three mechanisms.

  • Habits became rules. The repository history contains a v1.4.2 tag, so both self-written release skills told the agent to commit and tag. Without a skill, agents tagged in about half the release runs; with their own skill, they tagged in all 21 Sonnet runs that got as far as releasing and in all 30 release runs on Haiku.
  • Gaps became stalls. Sonnet's skill said to ask the user to confirm a major bump and to "propose the version and confirm". In a headless run nobody answers, so on the major-version and Changed tasks Sonnet stopped to ask before changing anything in 9 of 10 runs. Haiku's skill said a major bump should be used only if the user asks, so in 5 of 5 runs it released a Removed entry as 1.5.0 instead of 2.0.0.
  • Missing knowledge stayed missing. The repository has no unique index, so neither self-written migration skill knew about the ux_ prefix; both models named the unique SKU index ix_… in 5 of 5 runs. Without any skill, both passed that task in every run.

Ignoring the tag rule entirely does not rescue them. With that one check removed, self-written skills passed 16 of 30 release runs on Sonnet (no skill: 22) and 20 of 30 on Haiku (no skill: 19), while the compact skill still passed 30 and 29. This matches the SkillsBench audit, which found self-generated packs that "lock in confidently wrong assumptions" (SkillsBench, appendix D.6.1). A model cannot write down a rule it was never told, and it may write down the opposite.

Did the skill load when it should?

In all 360 Sonnet and Haiku skill runs and all 24 Opus skill runs, the agent loaded a skill, and it was always the relevant one. Separately, 40 prompts were run with read-only tools and a four-turn limit: 20 that should load one of the compact skills ("Products need a weight_grams field. Make the DB change") and 20 that should not ("Add type hints to the functions in orders.py"), on Sonnet and Haiku. Sonnet and Haiku both loaded the right kind of skill for 20 of 20 prompts that needed one and for none of the 20 that did not. With two short, clearly scoped descriptions, discovery was not the bottleneck. That may not hold with dozens of overlapping skills installed.

What I would do with this

  1. Write skills for decisions, not for procedures the code already shows. The gain came entirely from rules that were absent from, or contradicted by, the repository. If an agent can copy the convention from the file next door, a skill adds little.
  2. Keep them short. A one-screen skill with a fast path and the non-obvious rules did as well as a long one and cost less to run. The long one's background and examples bought nothing measurable.
  3. Do not ship a skill the agent wrote for itself without review. Use the draft as a starting point, then delete anything inferred from history that you do not actually want repeated, add the rules it could not know, and remove "ask the user" steps if the skill will run unattended.
  4. Say what not to do. The single most valuable line in the release skill was "Do not create git tags or commits: CI tags the release after review."
  5. Test with and without, on your own tasks. A dozen tasks with a scripted checker were enough to give a clear answer here.

A template with the same structure as the compact skills:

---
name: your-skill-name
description: What the skill does, in one sentence, in the words
  people actually type ("cut a release", "add a migration").
  Use when <situations>. Not for <nearest wrong use>.
---

# <Task> in <repository or team>

## Fast path

1. <The first concrete step, with the exact file path or command.>
2. <The next step. Put exact formats in a code block the agent can copy.>
3. <How to check the result, e.g. the test or validation command.>

## Rules that cannot be inferred from the code

- <A rule an agent would not guess, e.g. "CI creates tags".>
- <A format rule with an example.>

## Out of scope

<Where the skill stops, and what to do instead of guessing.>

Limitations

  • One small, synthetic repository. tillroll was written for this test. Real repositories are larger and messier, which probably makes skills more useful for navigation and less clear-cut to measure.
  • The release rules were chosen to be invisible. The effect size on releases reflects that choice. The migration result, a ceiling with no effect, is the other side of the same coin. Most real work sits somewhere in between.
  • Twelve tasks. Differences are averaged over six tasks per family. With so few tasks, task-level tests have little power: the permutation p-value for the compact skill is 0.06 on Opus (six release tasks) and 0.12 on Sonnet (all 12 tasks), although Opus improved on 5 of the 5 tasks that could improve and Sonnet on 4 of 4, the strongest possible results; on Haiku it is 0.04. Run-level intervals in the tables are narrower than the task-level ones and should be read as optimistic.
  • Opus on release tasks only. Claude Code defaults to Opus 5.5 on the Anthropic API and on most plans (Model configuration). The full grid used Sonnet 5.5 and Haiku 5.5 to fit the API budget; Opus ran the release tasks with no skill and the compact skill, four repeats each, on a Claude Pro subscription, so its costs are Claude Code's estimates. Opus was not run on the migration tasks (beyond the small exploratory check), the long skill or self-written skills, and Fable 5.1 was not run at all.
  • Claude Code and Claude models only. Other agents discover and apply skills differently; SkillsBench found the harness changes the benefit materially.
  • The comprehensive skills include worked examples. Length and the presence of examples are confounded, so "long" here means "long and example-heavy".
  • Self-written skills were generated once per model. A different draft could be better or worse; the mechanisms above are what this pair of drafts did.
  • The trigger test watched only the first four turns with read-only tools, and used two clearly separated skills.
  • Headless runs. Interactive use, where a person can answer "should I tag this?", would change the self-written result in particular.
  • Batches were set aside. Three earlier batches were discarded before the reported runs, each for an environment fault found by reading transcripts: the first for two errors in the test repository, the second for a stale copy the checker compared against, the third for a sandbox without a git identity, which stopped agents from committing. All are documented in the pre-registration and included in the spend. The no-identity batch (240 Haiku runs) shows the same pattern for the curated skills (no skill 32 of 60, compact 56, comprehensive 59); the self-written skills did better there (39 of 60) because commits failed and the agent usually stopped before tagging.

Methods note

The harness is a small set of Python scripts. For each run it makes a fresh copy of the test repository, starts Claude Code 2.1.293 headless with a throwaway home directory and --setting-sources project, and scores the result with that task's checker. It keeps a spend ledger and refuses to start a run that could take the total over a set budget. Python 3.13 and Node.js 22 were used. The 480 Sonnet and Haiku runs cost about US$11.43 at published prices. The Opus subscription runs used a variant of the harness that removes every API key from the child environment, runs one at a time and stops at the first usage-limit or sign-in error.

Bare mode (--bare) is the recommended mode for scripted and SDK calls, but it skips skill discovery (Run Claude Code programmatically), so the harness uses project settings and a temporary home directory instead. The same design works on any repository: a dozen tasks of your own, a scripted checker for each, and runs with and without the skill.

How this was made

Research, the test harness and the experiment runs were done with AI assistance; every number comes from the harness's scripts and the raw run data. Daniel Chen reviewed the design, the results and the text, and is responsible for what it says. Nothing here was paid for or reviewed by Anthropic or any other vendor.

Sources

  1. Agent Skills specification, agentskills.io. agentskills.io/specification (read 8 October 2026).
  2. Anthropic, "Skill authoring best practices". platform.claude.com (read 8 October 2026).
  3. X. Li, Y. Liu, W. Chen, B. You et al., "SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks", arXiv:2602.12670. arxiv.org/abs/2602.12670.
  4. Anthropic, "Run Claude Code programmatically". code.claude.com (read 8 October 2026).
  5. Anthropic, "Model configuration". code.claude.com (read 8 October 2026).
  6. Anthropic, "Pricing". platform.claude.com (prices as of 8 October 2026).
  7. E. B. Wilson, "Probable inference, the law of succession, and statistical inference", Journal of the American Statistical Association 22 (1927). doi.org/10.1080/01621459.1927.10502953.

Changelog

  • : first published.
  • : added 48 Opus 5.5 release-task runs (Claude Pro subscription) and rewrote the summary, charts and findings to lead with Opus 5.5 and Sonnet 5.5.
All engineering notes