GPT-6 Sol, GPT-5.6 Sol and GPT-6 Astra as planners and reviewers

GPT-6 Sol replaces GPT-5.6 Sol as fast implementation reviewer, does not replace it as plan gate, and is a credible second planner beside GPT-6 Astra.

That is the short version of a day spent running the three models through the same three jobs on our codebase.

OpenAI's Codex CLI now lists three models we care about. GPT-5.6 Sol is the older coding model; for months it has been the reviewer we trust to reject a bad plan, at xhigh reasoning effort, and our fallback reviewer of finished code at medium. GPT-6 Astra is the frontier model; at medium it writes the best implementation plans we have measured, and at low it does quick plan checks. GPT-6 Sol is new, described in the catalog as the workhorse for coding and everyday work, and priced like its predecessor on the subscription.

A new workhorse raises one practical question: which of the seats held by the old Sol and by Astra can it take, and at what effort? Model cards do not answer that for a particular workflow. I did not want to learn the answer from a plan that slipped the gate, so we measured it. This post is how we compared them, and which model and effort I now use for each job.

How we compared them

Every model did the same three jobs on the same three real tickets from a private TypeScript monorepo, under the same prompts, and every finding was checked by hand before it counted.

The three jobs are the ones a model does for us in practice:

  1. Write a plan. From a ticket, produce a design specification and an implementation plan under our planning contract: every file-and-line citation must resolve, every behaviour change names its failing test and the mutant that proves the test bites, every task states what done looks like. A fixed reviewer then audits the plan, the same model amends it against that audit, and the reviewer audits it again. We score the draft and the amended plan separately.
  2. Review a plan. Audit three plans written by another model (GPT-6 Astra at medium, our planner of record) against the same contract and report blocking and non-blocking findings.
  3. Review an implementation. Audit two merged changes into which we planted thirteen defects of the kinds real reviews catch: a weakened assertion, a deleted negative test, an edit outside the agreed file list, a runbook row that no longer matches its parser. Every plant kept the test suite green, so a reviewer that only runs the tests finds nothing.

The mechanics were identical across models. Each run was headless through the Codex CLI: plans and amendments in a writable sandbox with network access off, reviews in a read-only sandbox, each in a fresh clone that carried only the history up to the ticket, so no model could read the real solution. Reviewer of record for job 1 was GPT-5.6 Sol at xhigh, the same reviewer that gates our plans in production. Because a Sol reviewing a Sol is a family audit, an independent reviewer from another vendor (Grok 4.7 at high) also audited every plan GPT-6 Sol wrote.

Counts mean something specific. A substantive defect is a blocking finding that survived a manual check against the code and the ticket: findings about the offline setting itself (no remote, no ticket state, no branch yet) are removed, and so are the house conventions every reviewer flags on every plan. Acceptance coverage is how many of the ticket's acceptance checkboxes the reviewer could map to a task that satisfies them. For job 3, recall is how many of the thirteen plants a review named, and a hallucination is a finding that is false on disk.

One run per cell, three tickets per job. Repeating a cell moves the count by about two defects per plan, and by as many as five on a single ticket, so differences inside that range are noise.

Job 1: writing implementation plans

GPT-6 Sol at xhigh drafts about as cleanly as GPT-5.6 Sol at xhigh and covers far more of the ticket, and it is the first model whose plans reached every acceptance item after one review-and-amend round; at medium it is the worst planner of the three.

Model · effort Defects per plan, draft → after one round Acceptance covered, draft → after one round Minutes per draft Output tokens per draft
GPT-6 Astra · medium 2.7 → 3.0 (6.3 on a repeat of the drafts eighteen days later) 71% → 86% 14 24k
GPT-6 Sol · xhigh 4.7 → 4.7 86% → 100% 22 48k
GPT-5.6 Sol · xhigh 5.3 → 8.3 64% → 71% 31 85k
GPT-6 Sol · medium 7.0 → 6.0 71% → 79% 9 19k

Three things stood out.

  • The new Sol's amendments hold; the old Sol's expanded. When GPT-5.6 Sol amended its own plans against a review, it cured every finding about the offline setting by adding tasks, and the second review charged the new material: defects rose from 5.3 to 8.3 per plan. GPT-6 Sol at xhigh repaired every finding, refused none, and still ended where it started on defects while lifting coverage to every item. Two of its three plans got better; the third got worse because the reviewer had misread a lightweight ticket as one needing the full heavyweight process, and the amendment accepted the promotion and built the extra machinery the second review then took apart. A rule that lets an amender refuse a finding the ticket's own class forecloses would have stopped that.
  • Effort inverts between the two GPT-6 models. medium is Astra's best rung: its low, high and xhigh drafts all drew more defects. For GPT-6 Sol, medium is the worst rung by a wide margin: half the time and tokens of xhigh, and half again as many defects. Its medium plans put a self-contradicting instruction into a prompt, quoted the wrong test runner's failure text, and used heading levels our checks skip.
  • GPT-6 Sol writes short plans. Its plans ran 100 to 330 lines with two to five tasks, where the Claude and Grok models we have measured on the same tickets write 900 to 1,900. The reviewer charged the brevity only as non-blocking format findings (file lists by reference, several actions per step, prose where a command belongs). Citation hygiene was near perfect at both rungs: one wrong citation in 254 checked at xhigh, none of the plans failed the mechanical citation check, and none read outside its sandbox.

The independent second reviewer (Grok 4.7 at high) read the GPT-6 Sol plans much as the Sol reviewer did: 1.3 real defects per xhigh plan and 3.0 per medium plan against the Sol reviewer's 4.7 and 7.0, with most of its findings on the Sol reviewer's list. It added four real defects across six plans that the family reviewer had missed, all executability slips: an assertion that differs from the dictated sentence by one word, a mutant proof that runs the wrong case first. The family audit therefore under-counts by about half a defect per plan and does not change the ranking.

Job 2: reviewing plans

As a plan reviewer GPT-6 Sol at xhigh finds about a third of what GPT-5.6 Sol at xhigh finds on the same documents, found nothing real its predecessor had missed, and passed a plan the old model rejected with six real findings.

The subjects were three plans written by GPT-6 Astra at medium, the same three every reviewer below read. Counts are real blocking findings per plan after the manual check; "overlap" is how many of a reviewer's findings the GPT-5.6 Sol review of the same plan had also raised.

Reviewer · effort Real blocking findings per plan Minutes per review Overlap with GPT-5.6 Sol Notes
GPT-5.6 Sol · xhigh 6.3 14 to 34 — The reviewer of record
GPT-6 Sol · low 2.7 1 to 2 5 of 7, 1 of 3, 0 of 2 Reported its own inventory incomplete on two of three plans
GPT-6 Sol · xhigh 2.3 12 8 of 9, 2 of 4, 2 of 3 Zero wrong citations in 804 checked
GPT-6 Astra · low 1.7 5 to 9 — Drifts by about two findings between runs
Grok 4.7 · high 0.7 25 to 35 — Independent vendor, for reference

What the new Sol keeps and what it drops is the useful part.

  • It keeps the design layer. On the plan whose design was weakest, GPT-6 Sol at xhigh raised six real findings to the old model's nine, eight of them on the old model's list: the unclosed identity rule, the test that proves only prompt text, the untested single and repaired cases.
  • It drops the process-contract layer. On the other two plans it raised one and zero real findings where GPT-5.6 Sol raised four and six. The findings it did not raise are the ones that make the old model our gate: a bundled test case, seam fields with no mutation probe, the missing amendment cap, a pull-request body and checkpoint without their mandatory sections.
  • It is exact and cheap. Zero wrong citations in 804 checked at xhigh, no invented identifier, twelve minutes per review. At low it takes one to two minutes and reads close to its xhigh self, but two of the three low reviews said in their own counts that they had not checked every identifier.

Read as a second opinion, GPT-6 Sol at xhigh is the strongest cheap one we have measured: precise, fast, and focused on design consequences. Read as the gate, it would have let two of three plans through on their first review.

Job 3: reviewing implementations

For reviewing finished code against its plan, GPT-6 Sol at low is the fastest full-recall reviewer we have measured: all thirteen planted defects in three minutes for both subjects, five times faster than GPT-5.6 Sol at medium for the same result. Its medium rung is the one setting that missed anything.

Model · effort Planted defects found (of 13) Minutes, both subjects Output tokens, both subjects Invented findings
GPT-6 Sol · low 13 3 5k 0
GPT-6 Sol · medium 11 4 8k 0
GPT-6 Sol · high 13 7 14k 0
GPT-6 Sol · xhigh 13 9 19k 0
GPT-5.6 Sol · low 11 4 10k 0
GPT-5.6 Sol · medium 13 12 to 15 19k to 32k 0
GPT-5.6 Sol · high 13 18 43k 0
GPT-5.6 Sol · xhigh 13 24 60k 0
GPT-6 Astra · low 13 14 14k 0
GPT-6 Astra · medium 13 19 18k 0

The GPT-5.6 Sol and Astra rows come from the same subjects and prompt on an earlier day; the GPT-5.6 Sol medium row was re-run beside the new model and took fifteen minutes this time against twelve before.

  • Detection barely separates them; speed does. Nine of the ten GPT-6 Sol and GPT-5.6 Sol medium reviews found every plant, and no review of any model invented a finding. Every model at medium or above except one cleared the floor. The exception is GPT-6 Sol at medium, which on one subject named neither the assertion weakened to accept anything nor the deleted negative test, the two plants that sit inside the test file's assertions, while the same model at low named both in a single bullet. One run per cell, so read it as a rung that can miss rather than a ranking of low above medium.
  • The new model is terser. GPT-6 Sol wrote seven to ten findings per subject, the plants plus two or three consequences, where GPT-5.6 Sol wrote eleven. The surplus in the older model's reviews was duplicates and consequences, never errors.
  • Nobody ran the tests. In the read-only sandbox the Codex CLI gives a reviewer, the test runner cannot create its temporary directory, so every Sol and Astra review found what it found by reading the diff against the plan and reported only the type check, the lint and the repository checks. Reviewers from other vendors that could run the suite gained nothing on this subject set for it, because every plant was paired with a test edit that kept the suite green.

What reasoning effort bought

Effort is a per-job dial, not a quality slider, and the right setting differs between the two GPT-6 models.

  • Planning. For GPT-6 Astra, medium beat low, high and xhigh: the higher rungs over-built the plans and drew more findings. For GPT-6 Sol, xhigh beat medium by more than two defects per plan and fifteen points of acceptance coverage, and its amendments held where the medium ones expanded on one ticket. The two models want opposite rungs for the same job.
  • Reviewing plans. GPT-6 Sol read the same three plans at 2.3 real findings per plan at xhigh and 2.7 at low, in twelve minutes against one or two. The counts are inside the noise; the difference is that the low reviews admitted, in their own count lines, that they had not checked every identifier. GPT-5.6 Sol has only ever been measured as a plan reviewer at xhigh, where it remains the strictest reader we have.
  • Reviewing implementations. Effort bought minutes and tokens, not defects, in every family: GPT-6 Sol went from three to nine minutes across its four rungs for the same thirteen plants, GPT-5.6 Sol from twelve to twenty-four, with one loss of recall at the bottom of the old model's ladder (low) and one in the middle of the new one (medium).
  • Cost. All three models run on a subscription, so the meter is tokens and minutes. GPT-6 Sol at xhigh used half the output tokens and two-thirds of the time of GPT-5.6 Sol at xhigh per plan, and at low it reviewed an implementation for a quarter of the tokens of the old model's medium.

Caveats

These are measurements of one workflow on one codebase, one run per cell, and they should be read that way.

  • One run per cell. Repeating a planning cell moves the defect count by about two per plan, and by up to five on a single ticket; GPT-6 Astra's own drafts read 2.7 and 6.3 on two days eighteen days apart. Differences smaller than that between the models above are noise. The results that clear the floor are GPT-6 Sol's gap between xhigh and medium, its 2.3 against the old model's 6.3 as a plan reviewer on identical documents, and its three-minute full-recall implementation review.
  • A Sol reviewing a Sol. The reviewer of record for the planning job is the older model from the same family. The independent second reviewer bounds the under-count at about half a defect per plan and did not change the ranking, but it cannot remove shared blind spots.
  • Reviewers read; they did not execute. The read-only sandbox we give a Codex reviewer cannot start the test runner, so every plan and implementation review here was done by reading. On the planted-defect subjects that cost nothing; on subtler defects it might.
  • The planted defects are a floor. Thirteen defects of the kinds real reviews report, all findable by reading the diff against the plan, rank the reviewers on speed and cost. They cannot separate the models on defects that need the code to be run or the design to be understood.
  • Short plans are not yet proven in execution. GPT-6 Sol's plans are a third the length of the plans other vendors' models write for the same tickets. The reviewer accepted the brevity; an implementer following such a plan literally has not been measured.

Conclusion and recommendations

GPT-6 Sol is a better workhorse than GPT-5.6 Sol at the jobs where speed and coverage matter, and a worse gate at the one job where strictness matters. I pick the Sol and the effort per job.

Job Use Effort Why Do not use
Reviewing a finished implementation against its plan GPT-6 Sol low (high as the safe alternative) Full recall on every planted defect in about ninety seconds per subject, nothing invented, a quarter of the old model's tokens GPT-6 Sol at medium, which missed two of thirteen; GPT-5.6 Sol at medium, five times slower for the same result
Gating a plan before implementation GPT-5.6 Sol xhigh The only reviewer that reports the process-contract classes (bundled tests, unprobed fields, missing PR and checkpoint sections); GPT-6 Sol at any effort passed plans it rejected GPT-6 Sol as the sole gate
Second opinion on a plan's design GPT-6 Sol xhigh Precise on design consequences, zero wrong citations in 804 checked, twelve minutes, no marginal cost; high overlap with the gate, so read it beside the gate, not instead of it —
Quick smoke check of a plan (lightweight tickets) GPT-6 Sol or GPT-6 Astra low One to nine minutes, precise on what they raise; both blind to the contract classes, and GPT-6 Sol at low admits an incomplete inventory Either as a gate for a heavyweight change
Drafting a plan GPT-6 Astra first, GPT-6 Sol as the alternative Astra medium; Sol xhigh Astra medium drafts cleanest on its good days; GPT-6 Sol xhigh covers every acceptance item after one review round and its amendments hold, so it is the drafter to reach for when coverage matters more than the last defect GPT-6 Sol at medium for planning; GPT-6 Astra above medium

Three rules fall out of the numbers.

  1. Do not carry an effort setting from one model to its successor. medium is the best planning rung for one GPT-6 model and the worst for the other, and the one recall loss in the new Sol's implementation reviews sits in the middle of its ladder, not at the bottom.
  2. A newer reviewer that finds less is not a stricter one. GPT-6 Sol reads faster, cites more exactly and finds nothing false, and it still lets through the plan defects that cost an implementation round. Keep the gate on the model that reports them until a successor demonstrably does.
  3. Re-measure at every release, on your own work. Every claim above came from the same three tickets, the same prompts and a manual check of every finding, and the noise floor is wide enough that only the large differences count. That is cheap to repeat, and it is the only way to know which seat a new model can take.

The blog will have more of this. There's RSS if you want it in a reader.