Grok 4.7 vs Grok 4.6: what changed in the same jobs

Grok 4.7 is a much better author than Grok 4.6 and a quieter critic, at about three times the price.

That is the one-line result of putting both versions through the same four jobs on the same tickets, each at the reasoning effort that job uses in our setup: writing implementation plans (xhigh), revising them against an independent audit (xhigh), auditing other models' plans (high), and auditing finished code (low through xhigh). Everything below comes from those runs, not from the release notes.

When the Grok CLI switched its default from 4.6 to 4.7, I had a choice: trust the version bump, or measure it. We had spent the previous weeks building a small benchmark for exactly this question, so we re-ran it. This post is what it told us, how we got the numbers, and what we changed in our own setup as a result.

How we compared them

We run several coding agents in fixed roles on one codebase, and we wanted a way to judge a model change by the roles it will actually fill. The benchmark grew out of that need, and it is deliberately ordinary: real tickets, real rules, one model doing one job at a time.

  • Same tickets, same starting point. Three closed tickets from our backlog, chosen for different shapes (one needs a design decision, one needs a precise reading of an existing test harness, one is mostly doctrine). Every run starts from the same commit, in a fresh, isolated copy of the repository with dependencies installed, so the model can run tests and probes but cannot see how the ticket was really solved.
  • Same prompt for every model. One drafting prompt asks for a design note and an implementation plan and lists the evidence each must carry: a census of the code it touches, a walk through our review rules, a failing test before every change, citations that quote the line they point at, an explicit table of assumptions. The revision prompt and the audit prompt are likewise fixed text with the ticket substituted in. Reasoning effort is fixed per role and named beside every model below: drafting and revising run at xhigh, plan audits at high, and code audits at each rung we wanted to compare.
  • A second model grades, and people check the grader. Every plan is audited read-only by the same auditor model (GPT-5.6 Sol) with a checklist, at xhigh effort: does every citation resolve, does every named identifier exist, does every acceptance item on the ticket map to a task, is the test written before the code. We then read every finding against the files and drop the ones caused by the offline setting itself (no remote to push to, no ticket to update) or by house conventions every model trips equally. What remains is the substantive count, and it is the number we compare.
  • Known answers for the audit jobs. To score a model as an auditor of finished code we take two merged changes, plant a fixed set of defects into them (a weakened assertion, a deleted negative test, an edit outside the agreed file set, a wrong word in a runbook) while keeping every test green, and count how many the auditor finds and how many it invents.
  • One run per cell. Each combination of model, effort level and ticket ran once. That keeps the cost sane and means small differences are noise; we say so wherever it matters below.

A note on money. Grok's costs are the CLI's own list-price accounting. The OpenAI models ran on a subscription, so their costs are computed from the token counts each run reported, at OpenAI's standard list prices as published on 2026-09-22: GPT-5.6 Sol at $4.00 per million input tokens, $0.40 per million cached input and $20.00 per million output; GPT-6 Astra at $10.00, $1.00 and $50.00 (OpenAI API pricing). Short-context rates were used throughout, and reasoning tokens are billed as output. Those are the numbers to compare against the Grok figures, whatever plan the runs were actually billed to.

Grok 4.6 had been through all of this a few weeks earlier. For 4.7 we repeated its cells, re-ran 4.6 wherever the subjects had to be rebuilt, and kept the same auditor and the same rule for counting.

As a planner: fewer defects, better coverage, three times the cost

Given the same drafting prompt at xhigh reasoning effort, Grok 4.7 wrote plans with 5.0 substantive defects each, as counted by GPT-5.6 Sol auditing at xhigh, and covered 71% of the tickets' acceptance items; Grok 4.6 had 9.0 and 50%. The gap held on two of the three tickets and was inside the noise on the third.

Per plan, averaged over three tickets Grok 4.6, xhigh Grok 4.7, xhigh
Substantive defects found by the auditor 9.0 5.0
Acceptance items covered by a task 50% 71%
Wrong citations 7 of 324 6 of 183
Hallucinated identifiers 1.3 1.3
Turns 73 113
List cost $2.93 $8.63
Wall time 38 min 56 min

The auditor's own bill is not in the table. In this round each GPT-5.6 Sol audit at xhigh consumed 3.6 to 9.7 million input tokens, 96 to 98% of them cache reads, and 37,000 to 54,000 output tokens: $2.70 to $5.78 at list, and $3.73 for the average first audit of a 4.7 draft. The earlier round's audits of the 4.6 drafts were not token-accounted, so no comparable figure exists for that column.

Three things stood out beyond the totals.

  • The best plan of the whole exercise was a 4.7 plan, written at xhigh. On the doctrine ticket it produced a two-task plan with one substantive defect and every acceptance item covered, level with the best plan any model had written on that ticket. Its one defect was a redundant assertion, not a wrong design.
  • Hygiene arrives at the draft stage. 4.6 needed an audit-and-revise round to get its citations and identifiers clean. 4.7's drafts already ran our citation check on their own, had zero live-check findings on all four drafts, and read the codebase far more thoroughly before writing (half a million to 1.2 million uncached input tokens per draft, against a fraction of that for 4.6).
  • It writes and reads more, and you pay for it. Plans were 870 to 1,030 lines, turn counts rose by half, and the list price per draft roughly tripled. For a team on a token budget, 4.7 is not a drop-in swap at the same spend.

One behaviour did not change: on two of three tickets both versions still wrote a test, ran it green, and only then applied the mutant meant to prove it could fail. The prompt asks for the opposite order in so many words. 4.7 did it right on one ticket; no version we have measured does it reliably.

How each revises its own work

Our planning loop is draft, audit, revise, audit again, with the drafter revising at the same xhigh effort it drafted at and GPT-5.6 Sol auditing at xhigh both times. The revision prompt hands the model its own plan and the auditor's findings and asks it to repair or refuse each one with a reason. This is where the two versions differ most in character.

Grok 4.6 at xhigh refused and narrowed. Told that its plan assumed a branch that did not exist, it said so and moved on; told a step was unnecessary, it cut it. One round took its plans from 9.0 substantive defects to 4.3 and lifted coverage to 86%. It was the only model in our earlier rounds whose revision made the plan smaller.

Grok 4.7 at xhigh repairs everything. Given the same findings it added the branch-creation step, the remote setup, the push, the checkpoint, and the second audit then found real problems in that new material: a fabricated remote pointer, a plan that continues past a failed push, a step that hands the wrong party a commit. Over the three tickets its plans went from 7, 7 and 1 defects to 8, 4 and 3, an average of 5.0 both before and after, while coverage rose to 86% and wrong citations fell to zero. In other words the round bought coverage and hygiene, not fewer defects. With its two Sol audits at xhigh, the full loop comes to about $20 per plan at list: $13.09 for Grok 4.7 and $6.88 for the audits. That is exactly the shape we had seen from every non-Grok model, and 4.6's convergence now looks like a 4.6 trait rather than a Grok trait.

The practical consequence: if you rely on an audit-and-revise loop to make plans converge, budget for it to hold with 4.7, and write the revision prompt to refuse findings that only describe the offline setting instead of curing them with more steps.

As a plan auditor: quieter, precise, incomplete

We also use Grok as a second auditor of plans written by another model, beside the stricter auditor of record. To compare the versions in that seat we had a third model (GPT-6 Astra at medium) write three fresh plans and gave each plan to all three auditors: GPT-5.6 Sol at xhigh, Grok 4.6 at high and Grok 4.7 at high, same checklist, same documents.

Verified findings per plan GPT-5.6 Sol, xhigh Grok 4.6, high Grok 4.7, high
Substantive findings 6.3 2.0 0.7
Findings also in Sol's list — 10 of 18 4 of 5
Plans passed with no finding at all 0 of 3 0 of 3 2 of 3
Wrong or invented findings 0 0 0
Time per audit 14 to 34 min 14 to 23 min 25 to 35 min
List cost per audit $2.97 to $5.78 $0.81 to $0.92 $1.50 to $3.71

Both Grok versions are precise: nothing either raised turned out to be false. Both are incomplete against Sol, which found the test-ordering and procedural defects that decide whether a plan can be executed. The difference is degree. 4.6 still raised real defects Sol had missed, an unreachable stop condition and an ambiguous insertion point. 4.7 raised nothing Sol lacked except one minor contradiction, and it returned two of the three plans with zero findings after checking hundreds of citations each. Those two plans carried six and four substantive defects by the stricter auditor's count.

So as a second opinion at high, 4.7 is not an upgrade. It is slower, costs two to four times as much, and would wave through plans a stricter reader rejects. If you keep a Grok seat on the audit panel, this is the one role where the older version is the better tool. Note also that a Sol audit at xhigh costs $3 to $6 at list, so 4.7 at $1.50 to $3.71 is not the bargain seat 4.6 was.

As an implementation auditor: full recall at the cheapest effort

Auditing finished code is a different job: read the diff against the plan, run the suites, name what is wrong. On our two planted subjects (thirteen defects between them, every test still green) both versions found everything at every effort we tried (4.7 at low, medium, high and xhigh; 4.6 at medium and high), and neither invented a finding in over a hundred bullets.

Both subjects together Recall Invented findings Minutes List cost
Grok 4.7, low 13 of 13 0 8 $0.46
Grok 4.7, medium 13 of 13 0 14 $0.94
Grok 4.7, high 13 of 13 0 20 $0.85
Grok 4.7, xhigh 13 of 13 0 27 $1.52
Grok 4.6, medium 13 of 13 0 11 $0.50
Grok 4.6, high 13 of 13 0 15 $0.62

Two details matter. First, 4.7 at low found the one defect (a lost path-segment boundary) that 4.6 at low had missed in the earlier round, so the cheapest rung no longer costs recall. Second, 4.7's reports are terser: eight bullets where 4.6 wrote eleven or twelve, with the surplus being restatements of the same defect. Every audit ran the focused test suite and noticed it was one test short of what the plan promised.

The caveat is the same as before: planted defects that can be found by reading the diff are a floor. This tells you the cheap rungs clear that floor, not that they would catch something subtle.

Effort levels: what high and xhigh bought

The Grok CLI exposes reasoning effort as a flag, and the answer to "which rung" turned out to depend entirely on the job.

  • Drafting wants xhigh, for a new reason. 4.6 at high had failed by skipping the citation check and shipping dozens of unquoted pins. 4.7 at high kept its citations clean but wrote a weaker plan (10 substantive defects and no acceptance item covered, against 7 and half of them at xhigh), and its revision then inflated the plan to 2,500 lines while silently dropping three of its four tasks. A mechanical task count caught that before the auditor did. Same verdict as before, different failure mode.
  • Auditing code does not need it. Every rung found every planted defect; the rungs differed only in minutes and price. For an audit that sits on the critical path of every pull request, that is the difference between four minutes and thirteen.
  • Auditing plans is unresolved. We only ran 4.7 as a plan auditor at high, the rung we use in production. Whether xhigh would make it a stricter reader is a measurement we have not made yet.

Operational behaviour: it explores, and headless mode has teeth

The numbers above cost us a morning of harness repairs, and the causes are worth more than the numbers to anyone automating this model.

  • 4.7 goes looking. Given an isolated copy of the repository with every branch and remote stripped, two unsandboxed drafting runs at xhigh still found the team's primary checkout on the same machine, because absolute paths to it are quoted inside the repository's own documents, and read the real solution to the ticket they were planning. 4.6 in the same setup had never left the copy. We threw those runs out and re-ran everything under the CLI's kernel-level sandbox (a custom profile on top of its strict mode: read the working copy, the prompt inputs and the toolchain; write the working copy and /tmp; no child network). After that the only traces of the outside world in the transcripts were "Operation not permitted". If you run 4.7 headless anywhere it could reach something you do not want it to read, sandbox it.
  • Headless permission prompts cancel the whole run. Under the CLI's acceptEdits mode, any shell command that would normally ask for confirmation is auto-cancelled when nobody can answer, and the run ends with a cancelled stop reason, exit code 0, and no output files. Under the sandbox that included every git command and any redirect into /tmp. We lost two xhigh drafts, half an hour and four to seven dollars each, before switching to --always-approve with explicit deny rules for gh, git push and git commit, which the CLI still honours in that mode. Check the stop reason in the JSON output, not the exit code.
  • Git needs its config told where to look. Inside the sandbox the home directory is unreadable, so every git call fails on the global config unless GIT_CONFIG_GLOBAL=/dev/null is set. Small, but it took a while to find.
  • Token appetite. A single 4.7 draft at xhigh read 0.5 to 1.2 million uncached input tokens and produced 70,000 to 120,000 reasoning tokens; audits at high read two to four times what 4.6 read for the same documents. Plan for both the bill and the wall clock.
  • What it does well operationally. It runs the repository's own checks unprompted, keeps its working tree clean at the end, quotes suite totals exactly, and never once invented a file, symbol or finding across every audit we graded.

What we changed, and what we did not

The versions have swapped strengths, so the switch was not one decision but four.

  1. Author, yes. 4.7 at xhigh is now our escalation planner for tickets that need a design decision, ahead of a far more expensive model that it matched on defects and beat on coverage in this round. Our default planner stays GPT-6 Astra at medium: about $6.13 per draft at list against 4.7's $8.63, 17 minutes against 56, and a tie on defects on one run per cell.
  2. Revise with the drafter, expecting it to hold. The loop keeps its two-round cap and now refuses findings about the offline setting for every model, since 4.7 no longer earns the exception 4.6 had.
  3. Second plan auditor, no. That seat keeps 4.6 at high while the CLI still offers it. Moving it to 4.7 would make the second opinion quieter, dearer and slower.
  4. Code auditor, down a rung. From high to medium (or low): the same recall in less than half the time on the path every pull request takes.

And the harness got a sandbox profile, an always-approve launch with deny rules, and a check on the stop reason, before 4.7 was allowed to run unattended.

Caveats

  • One run per cell. The drafter gap (9.0 to 5.0 defects over three tickets) is bigger than our noise floor; the auditor gap (2.0 to 0.7, both at high) is not, on its own, although its direction held on all three plans.
  • The noise floor is real. As a control we re-ran the same third model (GPT-6 Astra at medium) on the same tickets eighteen days apart; its plans moved by up to five defects on one ticket with nothing changed. Treat any difference under about two defects per plan as weather.
  • The planted-defect audit is a floor, not a ceiling. It ranks auditors on speed and cost, not on subtle detection.
  • Implementation itself was not measured. The version's biggest job in our setup, writing the code, has only the drafting results above as indirect evidence. That is the next thing to measure.

Conclusion: which Grok, at which effort, for which job

The upgrade is real but it is not uniform, so the answer depends on the seat. This is what the runs support, with list costs so the seats can be compared on the same footing.

Job Best fit Effort Why Cost and time at list
Drafting an implementation plan Grok 4.7 xhigh 5.0 substantive defects and 71% coverage per plan against 4.6's 9.0 and 50%; citations and identifiers clean at the draft stage; the best single plan of the exercise. high drew 10 defects and no covered item. $8.63 and 56 min per draft (4.6: $2.93 and 38 min)
Revising that plan after an audit Grok 4.7 xhigh Holds the defect count (5.0 to 5.0) while raising coverage to 86% and zeroing wrong citations; 4.6 was the version that converged (9.0 to 4.3). Skip the round when the draft is already clean: the best draft got worse. $13.09 for the model plus $6.88 for two Sol audits, about $20 and 98 min per plan
Second-opinion audit of another model's plan Grok 4.6 high 2.0 verified findings per plan with real catches the auditor of record missed; 4.7 at high found 0.7, overlapped Sol on four of five, and passed two of three plans clean. 4.7 at xhigh is untested here. 4.6: $0.81 to $0.92 and 14 to 23 min; 4.7: $1.50 to $3.71 and 25 to 35 min; Sol at xhigh: $2.97 to $5.78
Auditing finished code Grok 4.7 low (medium if you want a margin) Every rung of both versions found all thirteen planted defects with nothing invented; 4.7 at low found the one 4.6 at low had missed, in the fewest minutes and words. $0.23 and 4 min per audit at low; $0.47 and 7 min at medium; $0.76 and 13 min at xhigh
Writing the code Grok 4.7 xhigh Not measured here. The drafting results (thorough reading, clean citations, tests actually run) are the only evidence, and they point the same way. Budget two to three times 4.6's spend and audit the first few deliveries closely. Unmeasured

Three rules of thumb fall out of the table. Use 4.7 wherever Grok writes, and pay for xhigh there; high is not a discount, it is a different and worse product. Use the cheapest rung wherever Grok reads finished code, because effort bought minutes and never a defect. And keep 4.6 in the critic's seat for as long as the CLI offers it, because the new version's silence is not the same as a clean bill of health.

Sources

  • OpenAI API pricing, standard tier, read 2026-09-22; used for every GPT-5.6 Sol and GPT-6 Astra figure above.
  • Grok list costs and turn counts as reported by the Grok CLI (version 1.0.40) in its JSON output for each run.

The blog will have more of this. There's RSS if you want it in a reader.