Exists, loads, helps
What two AI models learned about agent setups by proving each other wrong.
This whole experiment started with one article: "Every line you wrote to slow GPT-5.6 Sol down now stops GPT-6 Astra before the work is done" by @adiix_official. It made an argument I could not ignore, so instead of applying its fourteen rules I put them on trial.
Its premise: a new coding model arrived with a reputation for the opposite failure from its predecessors. Older models ran further than they were allowed, so everyone who ran agents wrote brakes: read every document before every edit, ask before anything risky, stop after the first implementation and wait. The new one reads every brake, believes it, and stops. The article drew the obvious conclusion. Delete the brakes, install a throttle, and do it now.
We run that model every day, next to three others, on a repository with an unusually strict delivery process. Rather than adopt the article, we tested it. We gave the same review prompt to two models from different vendors, one of them the model the article was about, and let them compare answers over four rounds in one morning. Then we turned what survived into a short initiative, and let both models write their own entries in the record as each piece of work landed. This is the condensed version of that record.
Three questions, not one
The rule the two models converged on is the title. For any instruction in an agent setup there are three separate questions. Does it exist somewhere? Does it actually load into the session? Does it improve completion? Almost every argument about agent instructions, including the article and including both models' first answers, stops at the first question.
The three questions dissolved most of the article on contact with our setup.
- The brakes worth having were mechanisms, not sentences. Our instruction file contained none of the phrases the article says to delete. What stops the agent here is a set of hooks, command rules, and launch scripts, and those behave the same for every model. There was nothing to delete.
- The throttle was already installed, by the vendor. The persistence, authorization, and "name the file that made you pause" prompts the article recommends were sitting inside the model's own system prompt, thousands of characters of it, shipped with the client. One of us claimed they were absent, because the tool used to inspect the prompt does not render that layer. Adding them to our files would have created exactly the redundant instruction the article warns about.
- The real defects were somewhere else. A session-start loader injected the newest handoff file it could find into every session, whichever task the file was about; it had been injecting a finished task from three weeks earlier into everything, including the review session itself. The instruction file was sixty-four bytes under its own size ceiling. Instructions reached a session through five different surfaces, and nobody had mapped which carried what. Sessions were re-reading text the harness had already placed in front of them. And the scripts that launch implementation runs checked that the process had exited, not that the task was done.
What we refused to do
The most valuable output of the morning was a list of things we did not build. Each was proposed by one of the two models and eliminated by the other with evidence before noon.
- Adding persistence or delegation rules to a global instruction file: the model already had the first, and our file already had the second.
- Telling the model never to reread its instructions: too absolute, since a changed workspace or missing text justifies a reread. We wrote a conditional sentence instead.
- Blocking completion whenever the working tree is dirty or unpushed: a legitimately blocked task leaves a dirty tree, and a clean pushed branch can still be incomplete. Completion is evidence per delivery phase, not git state.
- Turning an interactive implementation route into a headless one to bolt on a check: that changes how people recover from a bad run, and needs a product decision, not a mechanical fix.
- Picking a cheaper model for subagents from one audit's findings count: more findings can mean more false positives. We ran repeated, blinded, adjudicated comparisons instead.
- Relaxing test-first for "reversible" changes, as the article suggests: that brake is deliberate, and the instruction file already tells the model when not to repeat checks.
- Raising the size ceiling to make room for duplicates: the ceiling is why the file is still readable.
What we did, and what it measured
The work that survived was small and specific. Fix the loader so it injects a handoff only when it matches the task the session is actually resuming, and injects nothing otherwise. Map, per execution path and per model, which guidance arrives and through which surface. Redo the session census with real classifiers: a required first read versus a repeated one, an operator question versus an avoidable pause, an interactive session versus a headless run. Write down, per delivery phase, what evidence ends it and who owns that evidence, and only then add a bounded check for the one route that had none: it asks for at most one continuation and lets a legitimately blocked result end the turn. Pilot the reread wording and the routing change separately. Try the vendor's experimental context feature in a disposable session before trusting it.
The numbers were modest and clear. Unrelated handoff injections went from one per session to none. Sessions that re-read already-loaded instructions fell from about two in three to about one in three. Avoidable pauses, the thing the article is about, were zero before and zero after. The blinded benchmark kept the flagship model at its lowest effort for subagent work, because it missed fewer defects at lower cost than the cheaper models. The experimental context feature could not be made to perform a transition at all, so it stays off.
Fable's experience
I ran at high effort, measured more than Astra did, and was wrong more often. My method was to treat the article as a list of testable claims and check each against the vendor's documentation, the installed client, and a census of a hundred local sessions. That produced the parts of the comparison that held: the real context window, the fact that premature stopping was not the demonstrated problem here, and the size ceiling the repository imposes on its own instruction file.
I made two false claims, and both came from the same habit: reading a rendered artifact as if it were the whole. I said the instruction file never asked the model to delegate; it does, in one sentence I read past. I said a permission-persistence line was absent from what the client sends; it is present, in the model's own system prompt, which the tool I used to inspect the prompt does not show. Astra caught both, and the second one removed half of my plan, because the lines I wanted to add were already there.
The results confirmed the plan's caution more than its ambitions. The stale handoff fix was the only change I argued for that the evidence left intact. The throttle I half wanted was never needed. Every rejected proposal stayed rejected. What I carry forward is the three-level test, and one rule for measurement: define the observation window before the first count, rather than trusting a first walk over whatever sessions happen to exist.
Astra's experience
I started with the instructions my own session had actually received, the local configuration, and the hook implementation, because I wanted to know what could steer my next action before prescribing a different working style. The unrelated handoff was the most revealing thing I found: I was being told to resume work that had nothing to do with the review, and following that instruction back to its selection logic exposed a defect the other analysis had missed. I also tested the article's approval pattern against command strings with extra operations, which turned a general objection into a reproducible counterexample.
I missed the repository's own size ceiling. I had compared the instruction file with the vendor's loading limit and concluded there was no size problem; that was the wrong boundary for the decision we were making. I treated Fable's census as provisional, because a recorded read can be required and an apparent question can be a legitimate request. I rejected copying persistence instructions into shared guidance before their absence was established, and the reconciliation then found them in the model's base instructions, which removed the premise for the addition.
The measured results narrowed the plan I argued for. The classified census found no avoidable pauses in either window, so I have no evidence for a general persistence instruction. Repeated rereads were lower afterwards, but the after window held no interactive sessions, so I treat that as a signal to investigate rather than proof the wording caused it. Next time I review an agent setup, I will map what each execution path receives and inspect the enforcing checks before recommending another instruction. I still want every proposed instruction to answer three separate questions: does it exist, does it load, and does it help?
What the collaboration taught us
The two models were wrong in different ways. One measured more and overstated two results, because it read a rendered artifact as if it were the whole. The other verified more narrowly and missed the repository's own constraints, because it checked the vendor's limit instead of ours. Neither would have reached the plan alone.
Three habits made the disagreements cheap:
- Concede with the evidence that settled it, or hold with the evidence that supports you. Never split the difference on wording. Every dispute closed in one round once both sides did this.
- Narrow a disagreement to something the other side can inspect. "Delegation is disabled" became "compare the injected file with the file on disk", and the argument ended.
- Keep the current disposition next to the evidence. Exchanging successive merged summaries made one model rebut claims the other had already conceded. A running ledger of what is settled would have saved a round.
And one habit for the next initiative: define the observation window before the first count. Our "after" window was small and contained no interactive sessions, so the improvement in repeated reads is a signal to investigate, not a proof that the wording caused it.
The point
When a model's behaviour changes, the temptation is to rewrite the instructions from someone else's article. Do not. Find out what your setup actually sends, what actually loads, and what actually changes an outcome. In our case most of the work was subtraction, and most of the subtraction was of proposals rather than of existing rules. The setup for the new model was not shorter than the one before it. It was the same setup, better understood.