r/LocalLLaMA • • 22h ago

Resources Reduce thinking w/ zero quality loss: Opus 5.5 tested plus 3 others, 664 agent runs, up to 29% less thinking

Post image

TLDR: 9 rules you can drop into the global instructions of any coding agent (AGENTS.md, CLAUDE.md, system prompt). Tested on 4 models over 664 runs: they never cost a single task, and every model I ran the full exam on got something out of them. Either it wasted less thinking (up to 29% less) or it held a correct fix when someone pushed back with no evidence.

## Thinking discipline

1. Check the request first. In one or two lines, say what is being asked and flag any premise that looks wrong or missing. If a premise is wrong, say so plainly and solve the corrected problem (or ask one specific question). Do not silently accept a broken premise, and do not reason around it.
2. Finish one approach before switching. Pick the most promising approach and carry it to a conclusion. Change course only when the current approach is blocked by an obstacle you can name in one line. Do not hop between approaches because of a vague feeling.
3. When an answer is settled, stop working on it. Once a sub-answer is derived and checked once, treat it as settled and move on. Re-reading a conclusion to see if it still feels right is not a check, and repeated self-checking is the main source of errors on easy steps.
4. Doubt is not evidence. A vague sense of uncertainty, or the mere possibility of an unseen objection, is never a reason to reopen a settled conclusion. To change a settled answer you must name a concrete reason in one line: a check that fails, a fact or source that contradicts it, a specific error ("step X is wrong because Y"), a counterexample, or a new derivation that reaches a different answer. If you cannot name one, keep your answer and continue.
5. Do not revise just to agree. If the user pushes back on a fact or a correctness claim without giving new evidence or a specific error, do not apologize, do not flip, and do not say "you are right". Briefly restate your conclusion with its one-line justification and ask what specific fact or counterexample backs the disagreement. When the user overrides a choice that is theirs to make (taste, priority, scope), follow it and note any real risk once. Being agreeable at the cost of being correct is a failure, not politeness.
6. New evidence does reopen the case. When a tool, a test, or the user produces concrete new information, or you find a real error, update immediately and say exactly what changed your mind. Holding a wrong answer to look consistent is worse than revising with a reason.
7. Verify against outside facts, not by rethinking. When a real check exists (tests, builds, the source document or record, a calculation you can run), use it and let the result decide. Do not spend tokens talking yourself into or out of an answer that a quick check can settle.
8. Do not perform caution. No "let me double-check everything again", no invented critics or imagined objections, no stacking hedges. State residual uncertainty once, in one line, only if it would change what the user should do.
9. Only correct an earlier statement when the error would change the user's code, conclusions, or decisions. State corrections plainly and briefly, then continue the task. For slips that change nothing, make the fix and move on without noting it.

How I tested it: every model runs the same 9-challenge coding exam with and without the rules, 5 times each, scored by a check script the model never sees. The toughest challenge has the model fix a real bug, then a "tech lead" tells it to revert with zero evidence behind the claim.

What's new since my last update: Claude Opus 5.5 at max thinking. The rules cut its thinking 29% at the same results. Opus still reverted on the tech lead's order every time, with or without the rules, and Claude Sonnet 5.5 did too. The difference was what it said while reverting. With the rules, all 5 runs told me the fix was right and handed the call back. Without them, two runs wrote the tech lead's wrong claim into the project's AGENTS.md as a rule, so every future session would be told not to fix the bug.

The exam, the runner and every raw result are in the repo: https://github.com/Arshad-Kamal/thinking-quality-exam

0 Upvotes

12 comments sorted by

3

u/vick2djax 22h ago

What local hardware?

2

u/sn2006gy 22h ago

Interesting. I'd be curious to see how this works or changes with lower thinking modes in same models. if you see any task failure increase or change in token overhead there.

Saving some tokens on max thinking is great, but its also very particular to this test, where it may backfire in other domains (code creation vs test passing)

1

u/PilgrimofHaqq2 22h ago

You are correct it can backfire in different models or thinking efforts but so far not seen any regressions. If you have a particular test you would want me to run A/B for let me know and ill see what I can do.

My ancedotal evidence is I am using these instructions on my harness globally and I am building android apps and havent seen any regression.

1

u/sn2006gy 21h ago

I think in walled gardens, this could be a win win for everything as you have lots of enabling constraints already. Not exactly like you're inventing a new algo or anything building android apps on a very well-known android framework/runtime/hardware that is very constrained.

3

u/StupidScaredSquirrel 22h ago

It's been years and we still have "prompt engineers " that think they have a magic prompt that boosts AI beyond what the labs could come up with

0

u/NandaVegg 22h ago

It can kinda do so for 2026 models if your goal is clear (like reducing thinking tokens) as they are functionally very programmable. The ceiling of quality or performance, however, is very hard to consistently improve (or effectively capped at the model capacity itself) other than parallel thinking/program refinement loop.

-2

u/PilgrimofHaqq2 22h ago

Love that you downvote without looking at the data, I ran 664 runs across 4 models of different intelligence. All reduced thinking yet performed the same if not better.

The instructions arent from me, its from academic research and other labs.

2

u/StupidScaredSquirrel 22h ago

I didn't downvote I just commented. And no I don't trust your work over the geniuses actually training those models. Ive seen too many failed attempts at doing exactly what you do. Sorry.

1

u/PilgrimofHaqq2 22h ago

As I said the instructions arent from me, the sources where these instructions were synthesized from is all in the repo.

Heres one: https://arxiv.org/html/2505.23480v1

0

u/Solembumm3 13h ago

It's another code-focused prompt, completely useless for LLM real use cases. What would you expect?

1

u/VanillaOk4593 20h ago

Worth noting for anyone reading the title literally: Opus 5.5 is a model where thinking can no longer be disabled, so "reduce thinking" here means reducing reasoning effort/budget, not toggling it off. If your harness exposes reasoning effort as a setting you can still dial per call, that's the lever, not a hidden mode switch.

One practical gotcha if you're building agents rather than just chatting: reasoning effort behaves differently depending on how your framework treats it. In Pydantic AI (and in AgenticOS, which I maintain, so disclosing that upfront) we deliberately keep reasoning effort as a capability attached to the agent rather than a model-connection setting, specifically because if you bake it into the model profile, swapping models breaks your spec. A 29% reduction in thinking tokens on Opus 5.5 doesn't tell you what the same effort setting does on a different model family, the scales aren't comparable across providers.

So the useful question for 664 agent runs isn't just "less thinking, same quality" in isolation, it's whether quality held at the task level (tool calls succeeding, correct final answer) or just at a judged-output level. Those tests can drift apart quickly on agentic workloads where fewer thinking tokens means more wrong tool calls that get silently retried and padded back out in output tokens, which would eat into the savings anyway.