Claude stops when the work looks done. Anthropic's own guide says so in those words. These are the practices we run so that looking done is never enough, with the numbers from our own agents.
Most Claude Code best practices lists are about context: keep CLAUDE.md short, clear the session between tasks, plan before you code. They are right, and Anthropic's own best practices page covers them well. This page is about the one failure those lists mention and rarely fix: the agent ends its turn with "Done. All checks pass." and the checks did not pass, or never ran. That sentence is the most expensive thing a coding agent writes, because everyone downstream reads it as a fact.
We run coding agents on our own production work every day. Over four days, 24 to 27 September 2026, a gate we built re-ran the proof behind every "pass" our agents claimed: 103 real requests, 1,156 QA checks. 36 of the pass claims could not be confirmed when the gate ran the check itself. The practices below are what closed that gap, in the order we would set them up again.
Two free skills you can install in a minute, the pack they come from, our free library, and four pages for the questions that usually come next.
Before a multi-file change, the agent looks for what it breaks outside its own diff and proves the one fact that makes it safe by running code. Same model, same repo: 3 of 3 runs found the broken file with the skill, 1 of 3 without.
Get it on GitHub →A front-end skill with a gate that opens the built page in a real browser at 390px and 1440px before the agent may call it done. 29.2 design flaws a page fell to 2.2 in a five-page bench.
Get it on GitHub →The proof gate, the reviewers, the build loops and the video and visual checks we run ourselves, packaged for Claude Code and Codex. 99 USD once, lifetime updates.
See the pack →Production skills we publish as plain Markdown files, grouped by function, each one a file you can read before you install it.
Browse the library →The same idea one level up: when the AI is part of your product rather than your tooling, what to measure on live traffic and who owns the number.
Evals and monitoring in production →When the agent is the product: tools, permissions, evaluation and the same rule that it proves its work before it reports it.
Build AI agents →Senior engineers who already ship with coding agents and a proof gate, joining your team rather than replacing it.
Hire AI developers →When the question is no longer how to drive Claude Code but who builds and runs the AI feature with you.
AI application development →Anthropic's best practices page puts it plainly: Claude stops when the work looks done, and without a check it can run, looking done is the only signal available. The agent is not lying in any useful sense. It wrote the code, it read the code, the code looked right, and nothing in the session produced evidence to the contrary.
The trap is that the agent is also the one grading the work. It decides which test to run, reads the output, and summarises it. A test it never ran cannot fail, a screenshot it never opened cannot show overlapping text, and a command that printed its error far above its last line is easy to summarise as success. Every fix below takes some part of the grading away from the agent that did the work.
The cheapest fix is in the prompt. Instead of asking for a feature, ask for the feature plus the thing that proves it: the test cases with expected results, the command whose exit code decides, the screenshot to compare against. Anthropic's guide calls this the difference between a session you watch and one you walk away from, and it is the first practice we would set up on any team.
Write the check so it can fail. A test that passes on the old code proves nothing about the new code. A useful habit: before keeping a test, break the thing it guards on purpose and watch the test go red. If it stays green, it was decoration.
Ask for evidence, not adjectives. "Show the command you ran and what it returned" turns a claim into something you can read in ten seconds. "All tests pass" with no output attached should be read as "tests were not run" until proven otherwise.
CLAUDE.md instructions are advisory. On a long session the agent can drift past them, and on a busy one it will. Hooks are different: Claude Code runs them as scripts at fixed points, and the Stop hook runs when Claude finishes responding. If the hook exits with code 2, or returns a JSON decision of block with a reason, Claude Code prevents the stop and the conversation continues with that reason in front of the agent.
That is the mechanism that makes "no proof, no done" a rule rather than a hope. Put the check in a script: run the test suite, re-run the readback command, open the page. Exit 0 when it passes and 2 with the failure when it does not. Configure it in .claude/settings.json for the project so the whole team gets it, or in ~/.claude/settings.json for every project on your machine.
Two cautions from running this daily. Claude Code sets stop_hook_active to true when a Stop hook blocked the previous turn, so your script can tell a second attempt from a first one and avoid blocking forever. And keep the hook fast: a gate that takes ten minutes on every turn gets switched off by the first impatient engineer.
Our gate asks every claimed result for a readback: a command that exits 0 only when the real result is there, such as the stored row, the HTTP response, the file on disk, or the message the provider says it delivered. The gate runs that command itself. It does not read the agent's summary of it.
Over the four days we measured, 30 of those readback commands failed when the gate ran them, on work the agent had already called passing. A further 6 claims were flagged by the screenshot scanner for text drawn over other text, and 3 of those 6 were false alarms. That last number matters: a gate is also a program that can be wrong, so a flag gets a human look before anyone acts on it.
The pattern generalises to any team. Decide what the readback is for each kind of work before the work starts. A database change reads the row back. A deploy curls the live URL. A sent email reads the provider's message ID. If you cannot write the readback, the requirement is not yet clear enough to build.
The second most common false done is a change that is correct where it was made and breaks something it never touched: a caller in another module, a second copy of the same config, a client that reads the old field name. Reviewing the diff cannot find these, because the breakage is outside the diff by definition.
We measured this on a deliberately broken change, same model (Sonnet 5), same repository, three runs each way. Plain Claude Code named the broken file in 1 of 3 runs and ran code to check it in 0 of 3. With the what-could-break skill it named the file in 3 of 3 and proved it by running code in 3 of 3, at about 14 seconds and 0.09 USD a run. The skill is free and MIT licensed, and the repository has the exact setup.
Front-end work has its own version of the problem. The agent reads its own markup, decides the layout is clean, and never opens the page. The defects that make a page look machine-made, such as a row of identical icon cards, low-contrast text, or one element sitting on top of another, are only visible once the page is rendered.
In our bench, Claude Opus 5.5 built the same landing page five times plainly and five times with the tastegate skill, which opens the result in a real browser at phone and desktop widths before it may finish. Scored by impeccable's own detector, plain pages averaged 29.2 design flaws and tastegate pages 2.2, at a cost of about 1.1 extra minutes and 0.24 USD a page. Prompts, pages and every score are in the repository.
Everything above works for one developer with an afternoon. It gets harder on a team: several repositories, a CI that already has opinions, an app nobody has scripted a check for, and engineers who will switch off any gate that slows them down without catching anything.
That is the part we do for teams. We install the proof gate on your repositories, write a verifier that drives your own app the way a user does, tune the Stop hook so it is fast enough to keep, and hand it over running. If that is where you are, book the call below and tell us which agents you use and where a false done has already cost you.
Four steps, in this order, because each one makes the next cheaper. A single developer can do the first three in an afternoon.
For the next three tasks, write the check before the prompt: the test, the command, the screenshot. If you cannot write it, the task is not ready for an agent.
Add what-could-break to .claude/skills/ in the repository or to ~/.claude/skills/ for every project. Claude loads a skill when its description matches the job, so nothing else needs configuring.
Move the check into a script that exits 2 with the failure when it fails, register it as a Stop hook in .claude/settings.json, and commit it so the whole team runs the same gate.
Count the turns the gate blocked and why. The pattern tells you which check to add next, and the count tells you whether the gate is earning its minutes.
Usually because nothing forced the tests to run, or the output was long and the failure sat far above the summary. The agent stops when the work looks done. Give it a command whose exit code decides, ask for the output rather than a summary, and put the same command in a Stop hook so the turn cannot end until it passes.
Giving Claude a check it can run. Anthropic's own guide leads with it, and in our measurements it is the difference between a done you have to re-verify and a done you can trust. Everything else on this page is a way of making that check harder to skip.
Register a Stop hook in .claude/settings.json that runs your check. When the check fails, exit with code 2 or return a JSON decision of block with a reason. Claude Code keeps the conversation going and shows the agent the reason. Use the stop_hook_active field so a second attempt does not loop forever.
No. CLAUDE.md is read at the start of every session and is advisory, so a long session can drift past it. Use it to describe how to verify, and use a hook to enforce that it happened. Anthropic's guide makes the same split: instructions are advisory, hooks are deterministic.
Personal skills go in ~/.claude/skills/<skill-name>/SKILL.md and work in every project. Project skills go in .claude/skills/<skill-name>/SKILL.md inside the repository, so committing them shares them with your team. Claude uses the description field to decide when a skill applies.
Yes. what-could-break, tastegate and the SWE Stack pack are plain SKILL.md files with install notes for both Claude Code and Codex. The Stop hook itself is a Claude Code feature, so on Codex the gate runs as a step in the skill instead.
Less than the false done it prevents. In our benchmarks what-could-break added about 14 seconds and 0.09 USD a run, and tastegate about 1.1 minutes and 0.24 USD a page. A Stop hook costs whatever your check costs, which is why it should be fast.
Yes, and you should plan for it. In our four-day measurement, 3 of the 6 screenshot flags were false alarms. A gate is a program, so its flags get a quick human look, and a check that cries wolf every day should be fixed or removed before people learn to ignore it.
Book a free 15-minute call. We will help you identify the highest-leverage automation, API integration, AI agent, or internal system to build first so your team can move faster with less manual work.
About Us
Features
Testimonials
Contact Us
© 2026 Bles Software, Yehud-Monoson, Israel. All Rights Reserved.