Opus 5 vs Opus 5.5: what changed when we put both to work
Opus 5.5 replaces Opus 5 in every role we measured, and in two of the three it changes the effort we run.
That is the one-line result of putting both models through the three roles a coding agent plays for us: writing implementation plans from an issue, reviewing plans written by another model, and reviewing finished implementations against their plan. Everything below comes from those runs, not from the alias change.
When Claude Code's opus alias started resolving to Opus 5.5, every pin in our tooling that named it moved to the new model overnight, at the effort levels we had chosen for Opus 5. I did not want to learn the difference from production incidents, so we measured it. This post is how we compared them, what moved, what did not, and which version and effort I now use for each role.
How we compared them
We use Claude models in three roles: writing implementation plans from an issue, reviewing plans written by another model, and reviewing finished implementations against their plan. For each role we already had a measurement method from earlier model comparisons, so Opus 5.5 was dropped into the same harness Opus 5 had been measured in, on the same inputs, with the same graders.
Planning. Three real, already-implemented tasks from our backlog: a gating rule in a delivery tool that needed a parser and a refusal designed, a test-harness binding that needed a precise reading of an existing suite, and an architecture doctrine that needed judgement more than code. Each model received the issue text and an identical prompt asking for a design specification and an implementation plan under seven evidence gates: a census of every real instance of the formats it touches, a walk through the rules the change could trip, a failing test observed before every piece of code (with a named mutant where the property already holds), a quotation next to every line citation, an assumption table, and a self-audit. Each run worked alone in a sanitized clone of the repository at the same commit, with the real implementation of the task, the network and every other run walled off. Opus 5 ran at effort xhigh; Opus 5.5 ran at xhigh on all three tasks and at high on the design-heavy one.
Grading was mechanical first (every citation resolves, nothing edited outside the plan directory, our repository checks pass on the documents), then adversarial: a different vendor's model, GPT-5.6 Sol at effort xhigh, audited each plan read-only, verifying every citation, identifier and stated value on disk, mapping each acceptance criterion of the issue to a task, and classifying each finding as blocking or not. A separate reviewer then verified every finding against the documents and the tree, and findings that were only about the offline setting were removed. Each plan then got one amendment round by the model that wrote it, answering the audit, followed by a second audit. The number we compare is substantive blocking defects per plan after verification, alongside acceptance coverage, citation accuracy, cost at list price and wall time.
Reviewing plans. Three fresh plans written by a third model (GPT-6 Astra at effort medium) were audited by Opus 5 and by Opus 5.5, both at xhigh, with the same audit prompt. The same plans had already been audited by GPT-5.6 Sol at xhigh and by two Grok versions at high, which gives a scale to read the Opus counts against.
Reviewing implementations. Two merged implementations were checked out at their merge commit and given one further commit planting thirteen defects of the kinds real reviews here have reported: an identity check that lost its boundary, an assertion weakened to "is defined", a deleted negative test, an edit outside the allowed files, a wrong keyword in a runbook row, a commented-out assertion, a rule applied to every path instead of the first. Every plant keeps the test suites green. Opus 5.5 audited both at low, medium, high and xhigh; Opus 5 at high was re-run in the same session as a baseline. Every finding was graded by hand as a planted defect, a real extra defect, a hallucination, a duplicate or a note.
One caveat governs everything below. Every cell is one run. Re-running the same draft on the same task has moved its defect count by four or five, so differences of one or two per plan are noise, and only the results that repeat across tasks are worth acting on.
As a planner: same substance, better hygiene, and an amendment that finally converges
On the three plans, Opus 5 at xhigh had averaged 8.0 substantive defects per plan with 43% of acceptance criteria covered. Opus 5.5 at xhigh averaged 9.3 defects and 64% coverage. The defect difference is inside the noise; the coverage gain is not. Where the tasks separated, the pattern was Opus 5's in reverse: the new model wrote its cleanest plan on the doctrine task (7 defects against Opus 5's 12) and its worst on the design task (11 against 6), where it keyed the identity of a repeated finding on overlapping line anchors, a choice that misses the drift case the task exists to catch. The design edge we had seen from Opus 5 did not reappear, and we no longer assume it.
Opus at xhigh, per plan |
Opus 5 draft | Opus 5 after one round | Opus 5.5 draft | Opus 5.5 after one round |
|---|---|---|---|---|
| Substantive defects | 8.0 | 10.0 | 9.3 | 5.3 |
| Acceptance criteria covered | 43% | 57% | 64% | 86% |
| Wrong citations | 4 of 118 | 4 of 179 | 1 of 270 | 2 of 318 |
| Unquoted or drifting citations left for the auditor | 37.7 | 15.7 | 0 | 0 |
| Cost, list price (USD) | 25.87 | 44.92 | 24.22 | 36.27 |
| Wall time (min) | 43 | 79 | 64 | 106 |
Hygiene moved a long way. Opus 5.5 drafts had one wrong citation in 270 checked, against Opus 5's four in 118, and every draft ran the repository's own citation check before finishing, where Opus 5 drafts had left dozens of unquoted or drifting line references for the auditor to find. The test-first discipline moved too: one Opus 5.5 draft ran its mutants before the first green run, which no earlier Opus or Fable draft had done, and the test-first findings on the other two were narrower than the write-green-then-mutate habit every earlier draft had shown.
The headline is the amendment round. Given one audit and asked to repair or refuse each finding, Opus 5 had made its plans worse, from 8.0 to 10.0 defects per plan, because it answered every finding about the offline setting by adding set-up tasks that the second audit then took apart. Opus 5.5 did add the same kind of material, and the second audit did charge it, but it also repaired or removed the design and executability defects instead of rearranging them: it deleted an issue-creation transaction it could not specify, dropped a clause it had added to a doctrine, and rebuilt one task as thirty-four one-action steps with each test observed red before its code. The three plans went from 11, 10 and 7 defects to 7, 5 and 4, with coverage rising to 86%. Until now only one model, Grok 4.6 on its own drafts, had converged under this loop.
Cost per draft was level at list price for 45% more turns, and drafts took half as long again. On the one task we also ran at effort high, Opus 5.5 matched its own xhigh draft on defects and coverage at two-thirds of the cost, and paid for it in citations: four wrong in 110 on the draft and seven in 124 after amendment, against one and zero at xhigh.
As a plan reviewer: stricter than Opus 5, still not a substitute for a second vendor
Auditing the three GPT-6 Astra plans at xhigh, Opus 5.5 raised 2.3 verified blocking findings per plan and ten non-blocking ones, every one of which survived verification as real or overstated. Opus 5 at xhigh raised 1.0 blocking and eight non-blocking, in more time (13 to 19 minutes against 8 to 17) and at higher cost ($5.26 to $7.61 per audit against $3.40 to $4.91). Both were exact on citations, one wrong each in about 600 checked.
| Auditor and effort | Verified blocking findings per plan | Plans passed with no blocking finding |
|---|---|---|
GPT-5.6 Sol, xhigh |
6.3 | 0 of 3 |
Opus 5.5, xhigh |
2.3 | 1 of 3 |
Grok 4.6, high |
2.0 | 1 of 3 |
Opus 5, xhigh |
1.0 | 2 of 3 |
Grok 4.7, high |
0.7 | 2 of 3 |
Read against the other auditors on the same plans, Opus 5.5 sits between the Grok versions and the cross-vendor auditor. What Opus 5.5 raised that the GPT auditor did not were consequences of the change, such as a stop condition no task could reach, a contradiction between a new rule and an existing taxonomy, and index entries left pointing at a renamed section. What it missed were contract items: a single parameterised test bundling six behaviours, a pull-request body without its mandatory sections, missing checkpoint fields. On one of the three plans it raised no blocking finding at all where the GPT auditor found four, exactly as both Grok versions had. A same-family auditor, however strong, still passes plans a different vendor rejects.
As an implementation reviewer: full recall at every effort, and low is the value rung
Every Opus 5.5 audit found all thirteen planted defects on both subjects and invented nothing, at every effort. So did Opus 5 at high. What separated them was time and money.
| Auditor and effort | Planted defects found | Minutes, both subjects | Cost, both subjects (USD list) |
|---|---|---|---|
Opus 5.5, low |
13 of 13 | 6 | 1.31 |
Opus 5.5, medium |
13 of 13 | 7 | 1.87 |
Opus 5.5, high |
13 of 13 | 9 | 2.26 |
Opus 5.5, xhigh |
13 of 13 | 12 | 3.85 |
Opus 5, high |
13 of 13 | 14 | 5.53 |
Grok 4.7, low |
13 of 13 | 8 | 0.46 |
At the same rung the new model costs 41% of the old one and uses 42% fewer turns for the same result. Opus 5.5 at low is now the fastest full-recall implementation audit we have measured in any model family, though on list price it remains about three times Grok 4.7 at low.
These subjects are a floor, not a ceiling. Thirteen defects that a careful reader of the diff against the plan can find show that every rung clears the floor; they cannot separate the rungs on subtler defects.
Smaller differences we noticed along the way
- Opus 5.5 repairs by removing as readily as by adding. Opus 5 answered audits by growing the plan; Opus 5.5 grew its plans too, but its refusals were rare and each cited the rule that made the finding wrong.
- Given a repository with reviewer agents defined, one Opus 5.5 draft dispatched a read-only review of its own documents before finishing and repaired twelve defects that pass found. The cost is inside its figures, and the behaviour is what the production setting would produce.
- Both versions carry the same ceiling on design under ambiguity. Neither wrote a plan the cross-vendor auditor would accept without a round, and the audit-and-amend loop stays necessary whoever drafts.
Conclusion and recommendations
For every role we measured, Opus 5.5 replaces Opus 5, and in two of the three it changes the effort we run.
- Planning. Opus 5.5 at
xhigh, always inside an audit-and-amend loop with a different vendor's auditor: the draft alone is no better than Opus 5's on substance, but one round with the model amending its own plan is now the best endpoint a Claude planner has reached for us (5.3 defects and 86% coverage per plan). Efforthighis a defensible economy on bounded tasks if a citation check runs before the audit; do not use it where design must be decided. Do not send design escalations to Opus 5.5 on the strength of Opus 5's record; that edge has not been shown by the new model. - Reviewing plans. Opus 5.5 at
xhighas the second opinion beside a cross-vendor auditor of record, which for us remains GPT-5.6 Sol atxhigh. It finds three times what Grok 4.7 athighfinds, at one to three times the price, and its non-blocking list is long and real. It is not a sole gate. - Reviewing implementations. Opus 5.5 at
low. Full recall, no hallucination, the suites run, about three minutes and 65 cents per audit. Thexhighfloor we had given Opus audit sessions buys nothing here and costs three times as much.
Where Grok fits, from the same method run on Grok 4.6 and 4.7 the week before: Grok 4.7 at xhigh is the drafter and amender (5.0 defects and 71% coverage per plan as a draft, 86% after its round, at $9 and $13 respectively), and high fails as a drafter in both versions, once by skipping the citation check and once by truncating the plan. As a plan auditor, Grok 4.7 at high is quieter than Grok 4.6 at high (0.7 against 2.0 blocking findings per plan) and both are far below the cross-vendor auditor, so neither should gate alone. As an implementation auditor, Grok 4.7 at low gives full recall for about 23 cents an audit, the cheapest full-recall setting we have measured; if cost matters more than a Claude-family read, that is the rung.
The blog will have more of this. There's RSS if you want it in a reader.