We stopped asking which model is best
I used to put the "smartest" model on planning and audit. That became a quota problem before it was a quality problem. We ran a measurement bench on real issues. We now route by role, not prestige.
This is how Steady Finch builds — the AI delivery system behind the work, not a product launch.
How we got here
Earlier, Fable planned and audited. Then Fable compiled Class P work — the changes that get a full written permit — and Grok was the CPU, the implementer. Opus became the everyday planner and auditor; Fable was escalation only. Fable never implements. That's policy, not a mood.
We classify each change by how much planning and audit it gets. Class P is the heavy path: a full permit and dual audit. Class E is the middle. Class C is the light path — a draft, little audit. A schema that other products will share is P. A one-line copy fix is C.
In September we asked a narrower question, with evidence: who should draft permits, who should audit, who should implement.
What the bench actually said
We compared three issues. Sol graded substantive defects and acceptance. One run per cell, so treat the numbers as approximate — our bench, not a paper.
- Opus 5 xhigh, plain: about 8.0 defects, 43% acceptance, about 43 minutes. It designs well, and it over-builds when the problem is fuzzy.
- Fable medium, plain: about 5.7 / 64% / 26 minutes. The best Claude draft. More Fable effort made it worse.
- Astra medium, plain: about 2.7 / 71% / 14 minutes. The best measured draft.
- Astra medium, then Sol audit, then Astra amends: about 3.0 / 86% / 25 minutes. The best planner combo.
- Grok, then Sol audit, then Grok amends: about 4.3 / 86% / 65 minutes. The only Claude/Grok loop that converged.
A few rules stuck.
- Audit-and-amend is not universal. It helped Grok. It hurt Opus, Fable, and Sol when the amender bloated the plan. The amender has to be the same drafter. Cross-model amend nearly doubled defects.
- Effort is not quality. Fable low beat high. Astra medium beat high and xhigh.
- Don't route by list price.
- Two auditors catch disjoint defects. Sol and Grok, on the same drafts, do not see the same things. Sol is the one that catches test-first ordering.
- A cheap auditor calling a draft "clean" does not mean it's ready.
- No model produced a permit we accepted without amendment. The loop stays.
- On implementation audits, from medium up, the families found the planted defects. Grok won on time and cost for the same recall.
What we do now
Live routing, at a high level:
- Draft Class P plans with GPT-6 Astra (medium).
- Dual-audit those permits with Sol xhigh and Grok high. The drafter amends itself.
- Escalate to Opus only when design findings remain after the cap.
- Fable for architecture-diff and auditor-disagreement escalation. It never implements.
- Implement with Grok for general work, or Astra for frontend. Audit the diff with a model that didn't write it.
Policy, not fandom
The useful question is not "is Opus better than Grok?" It's "is this a design, a clone, or a code change — and which measured loop belongs there?"
Routing is policy. We change it when the bench says so, not when a model is fashionable.
The blog will have more of this. There's RSS if you want it in a reader.