CLAUDE CODE HOOKS

Claude Code hooks, and the one that refuses "done"

An instruction in CLAUDE.md is advice the agent can drift past. A hook is a script Claude Code runs itself, and on a blocking event its exit code 2 refuses the action.

Every event name, field and exit code on this page was read from Anthropic's own hooks reference and links to it. Every number is from our own gate's run log over the eight days to 1 October 2026, or from a benchmark whose prompts and scores are published in the skill's repository.

Most people arrive at Claude Code hooks for one of two reasons. Either they want something to happen automatically, such as a formatter after every edit or a notification when the agent is waiting on them, or they are tired of the agent ending its turn with "Done. All checks pass." when the checks did not run. The first is a convenience. The second is the reason hooks exist at all, because it is the one problem no instruction can fix.

We run coding agents on our own production work every day, with a gate built on exactly this mechanism. Over the eight days from 24 September to 1 October 2026 it handled 310 real requests and 3,093 QA checks, and 152 of the pass claims our agents wrote were not confirmed by the gate's own rerun of the check. This page is the contract that makes that possible, in the order you need it to wire one up.

The pieces, free first

The skill that decides what proof means for a change, the same skill as a zip, our free library, the practices page this one grew out of, and the pack they all come from.

what-could-break (free, MIT)

A hook can run your check. It cannot work out which check this particular change needs. That is what this skill does: it looks for what a change breaks outside its own diff and proves the one fact that makes it safe by running code.

Get it on GitHub →

The same skill as a zip

If you would rather not clone anything, the skill is a single SKILL.md file in a zip you can drop into .claude/skills/ and use in the next session. No signup, no email.

Download the zip →

Claude Code best practices

The wider version of this page: why the agent says done when the work is not, how to write a check that can fail, and the practices we would set up on a team in order.

Read the practices →

tastegate (free, MIT)

The front-end half of the same idea. Its gate opens the built page in a real browser at phone and desktop widths and scores it before the agent may call the work done.

Get it on GitHub →

Our free skills library

Production skills we publish as plain Markdown files, grouped by function, each one a file you can read in full before you install it.

Browse the library →

SWE Stack: all 23 skills

The proof gate, the reviewers, the build loops and the visual and video checks we run ourselves, packaged for Claude Code and Codex. 99 USD once, lifetime updates.

See the pack →

Evals for AI already in production

The same question one level up. When the AI is part of your product rather than your tooling, what gets measured on live traffic and who owns the number.

Evals and monitoring in production →

Developers who already work this way

Senior engineers who ship with coding agents behind a proof gate, joining your team rather than replacing it.

Hire AI developers →

What a hook is, and why an instruction is not one

CLAUDE.md is loaded into context and read like any other text. It is advice, and on a long session the agent drifts past it, not out of defiance but because the instruction is competing with everything else in the window. A hook is different in kind. Claude Code runs it as a program at a fixed point in the session, and nothing in the model's reasoning decides whether it runs.

That buys two things. It happens every single time, and on the events Anthropic marks as blocking it can refuse. A hook that formats a file is convenient. A hook that can say "no, this turn is not over" is the only way to make a verification rule hold on a session you are not watching.

The price is that a hook is code you own. It runs with your own credentials and whatever the script can reach, so treat it like any other script in the repository: read it before you install someone else's, keep it in version control, and keep the risky parts narrow.

The events, grouped by what they can actually stop

Anthropic's reference now lists more than thirty hook events, which is why most introductions to them feel like a menu with no recommendation. For verification, the useful way to read the list is by whether the event can block.

Before something happens: PreToolUse fires before a tool call and can block it or rewrite it. PermissionRequest fires when a call needs a permission decision. UserPromptSubmit fires before your prompt reaches Claude and can refuse it outright.

After something happens: PostToolUse after a tool call succeeds, PostToolUseFailure after one fails, and PostToolBatch once a whole batch of parallel calls resolves, which can stop the agentic loop before the next model call.

At the end: Stop when Claude finishes responding, SubagentStop when a subagent finishes, and TaskCompleted when a task is about to be marked done. All three can refuse, and those are the three a proof gate lives on.

The rest are reporting surfaces. Notification, SessionStart, SubagentStart, PostCompact and MessageDisplay cannot refuse anything, so a gate placed on one of them is a log line with ambition. Check an event against the blocking column before you build on it.

The exit code contract, which is the whole mechanism

A command hook gets JSON on stdin and answers with its exit code. Exit 0 means success, and on most events the stdout goes only to the debug log. Four events are the exception and feed plain stdout back as context Claude can see: UserPromptSubmit, UserPromptExpansion, SessionStart and PostModelSwitch. That is how a hook injects a fact into a session instead of blocking it.

Exit code 2 is the blocking error. On a blocking event the action is prevented and stderr goes to the agent as the reason. Any other exit code is non-blocking and the action proceeds, which is the trap worth knowing about before you write the script: a gate that dies on an unset variable exits 1 and gates nothing. Make every failure path exit 2 on purpose.

For finer control the hook can print JSON on stdout. A top-level decision of block with a reason works on Stop, SubagentStop, PostToolUse, PostToolBatch, UserPromptSubmit and several more. PreToolUse instead uses hookSpecificOutput with a permissionDecision of allow, deny, ask or defer, and can hand back updatedInput to rewrite the tool call rather than refuse it, which is how you sanitise a command instead of fighting the agent over it. A continue of false with a stopReason ends the turn outright.

One ordering rule saves an afternoon of confusion: exit code 2 overrides a JSON permissionDecision of allow. If your script both allows and exits non-zero, the block wins.

The Stop hook, the one worth setting up first

Stop fires when Claude finishes responding. Exit 2, or return a decision of block with a reason, and Claude Code prevents the stop: the conversation keeps going with your reason in front of the agent. That single behaviour is what turns "no proof, no done" from a request in a Markdown file into a rule the session cannot walk past.

The script is less clever than people expect. Run the thing that decides: the test command, the readback that reads the row or the HTTP response, the script that opens the page. Exit 0 when it passes. Exit 2 with the actual failure text on stderr when it does not. The agent then has a specific failure to fix rather than a scolding.

Guard the loop or you will never get a turn to end. The JSON on stdin carries stop_hook_active, which Claude Code sets to true when a Stop hook blocked the previous turn. Read it, and either let the second attempt through or escalate differently. Stop also receives last_assistant_message and turn_number, so a gate can look at what was claimed before it decides whether to ask for proof at all.

SubagentStop is the same contract one level down, and teams that fan work out to subagents need it more than they expect. Without it the parent agent inherits a confident summary of work nobody proved, and then reports that summary to you with its own confidence added.

What eight days behind a Stop gate actually look like

Our gate holds a requirement list for each request, and for every requirement the lane has to write what it did and the command that proves it. The gate then runs that command itself rather than reading the agent's summary of it. From 24 September to 1 October 2026 that came to 310 real requests and 3,093 QA checks.

152 of the pass claims were not confirmed by the gate's own rerun. 79 of them were a readback command that exited non-zero the moment the gate ran it. 48 were flagged by the screenshot scanner for text drawn over other text. The remaining 25 were screenshots the scanner could not read at all, too little legible text to judge, and an unjudged screenshot is not evidence, so the claim was not accepted on trust. 289 of the 310 requests finished with every requirement passing.

A gate is also a program, so it can be wrong, and the honest number matters more than the impressive one. 71 flagged rows came back with a written judgment of what the flagged crop actually showed, and in every one of those the flag turned out not to be text over text: a label clipped to an ellipsis on purpose, a heading halfway under a sticky header, particles sitting beside a word. That is why a flag on our gate asks for a look rather than failing the work by itself. A check that cries wolf every day and never explains itself is a check your team will switch off by Thursday.

Matchers, handler types and timeouts, which decide whether anyone keeps it

Hooks live under a hooks key in settings, grouped by event name, each group carrying a matcher and a list of handlers. On tool events the matcher is tested against the tool name. A matcher of "*", an empty string, or no matcher at all fires every time. Plain names, with pipes or commas between them, are read literally, so "Write|Edit" covers both. Add any other character and it becomes an unanchored JavaScript regular expression, which is how "mcp__memory__.*" catches a whole MCP server. Other events match other things: SessionStart on its source, which is startup, resume, clear, compact or fork, PreCompact on manual or auto, and ConfigChange on which settings file moved.

A handler does not have to be a shell command. The type can be command, http for a POST to a URL, mcp_tool to call a tool on an MCP server, or prompt and agent to hand the judgment to a model. Inside a matcher group, an if condition narrows it further using permission rule syntax, so "Bash(git *)" or "Edit(*.ts)" keeps an expensive gate off every unrelated call.

Timeouts are where good gates die. Command, http and mcp_tool handlers default to 600 seconds, a prompt handler to 30 and an agent handler to 60. Nobody tolerates a ten minute wait on every turn, so put a real timeout on the handler and keep the check fast. For work that genuinely takes minutes, async of true runs the handler in the background without blocking, and asyncRewake wakes Claude when that background script exits 2.

Where the file lives decides who gets the rule. .claude/settings.json in the repository is the one that matters, because it is committed and the whole team runs the same gate. ~/.claude/settings.json covers every project on one machine, .claude/settings.local.json stays out of git for your own experiments, and ${CLAUDE_PROJECT_DIR} in a command keeps the script path correct no matter which directory the session is working in.

Where the hook ends, the skill begins, and when to have it set up for you

A hook is deterministic and knows nothing about your work. It will run the command you gave it, every time, which is exactly what you want from a gate and useless for deciding what the right check is for one specific change. That judgment is the other half, and it belongs in a skill.

what-could-break is the one we give away because it is the half people skip. Before a multi-file change it looks for what the change breaks outside its own diff and proves the one fact that makes it safe by running code. Same model, same repository, three runs each way: plain Claude Code named the broken file in 1 of 3 runs and ran code to check it in 0 of 3. With the skill it was 3 of 3 and proved it by running code in 3 of 3, at about 14 seconds and 0.09 USD a run. It is free, MIT licensed, and the repository carries the exact setup so you can rerun the comparison on your own code.

The pair is the point. The skill works out what proof means for this change, and the hook makes sure that proof happened before the turn was allowed to end. If you want the pair on a team rather than a laptop, that is the part we do: the gate on your repositories, a verifier that drives your own app the way a user does, a Stop hook tuned fast enough that nobody disables it, handed over running. Everything we run ourselves ships in SWE Stack for 99 USD once, and the two free skills above are the same code as the ones inside it.

Wiring one up this afternoon

Four steps in this order, because each makes the next cheaper. A single developer gets through the first three in an afternoon on one repository.

1

1. Make the check a script, not a sentence

Put the thing that decides into one executable: the test run, the readback, the page open. It must exit 0 on pass and exit 2 with the failure on stderr, never exit 1 by accident.

2

2. Register it on Stop

Add a Stop entry under hooks in .claude/settings.json pointing at the script with ${CLAUDE_PROJECT_DIR}, commit it, and the whole team gets the same gate on the next session.

3

3. Read stop_hook_active before blocking

The field is true when your hook already blocked the previous turn. Branch on it so a second attempt can finish, or the session will never end and someone will delete your hook.

4

4. Count the blocks every week

Keep a line per block with its reason. The pattern tells you the next check to add, and the count tells you whether the gate is earning the seconds it costs.

The numbers behind this page

152
Pass claims not confirmed by the gate's own rerun, over 310 real requests from 24 September to 1 October 2026
79
Of those were a readback command that exited non-zero the moment the gate ran it instead of reading the summary
71
Screenshot flags that came back with a written judgment, and not one of them turned out to be text drawn over text
3 of 3
Runs that found the broken file with what-could-break, against 1 of 3 for plain Claude Code on the same model and repository

Questions developers ask about Claude Code hooks

What are Claude Code hooks?

Scripts Claude Code runs itself at fixed points in a session, configured under a hooks key in settings.json. Because the runtime invokes them rather than the model choosing to, they happen every time, and on the events Anthropic marks as blocking they can prevent the action that was about to happen.

Which hook stops Claude Code from ending a turn before the work is verified?

The Stop hook, which fires when Claude finishes responding. Exit with code 2, or return a JSON decision of block with a reason, and Claude Code prevents the stop and continues the conversation with that reason in front of the agent. SubagentStop does the same for a subagent.

What does exit code 2 mean in a Claude Code hook?

It is the blocking error. On a blocking event the action is prevented and stderr is shown to the agent as the reason. Exit 0 is success, and any other code is treated as non-blocking, so the action proceeds. That last part is why a script must fail with an explicit 2 rather than whatever code it happens to die with.

Where do I put a Claude Code hook?

Project hooks go in .claude/settings.json inside the repository, which is the one to use for a team gate because it is committed. Personal hooks go in ~/.claude/settings.json and apply to every project on that machine, and .claude/settings.local.json stays out of git. Plugins and skill frontmatter can also register hooks.

How do I keep a Stop hook from looping forever?

Read stop_hook_active from the JSON on stdin. Claude Code sets it to true when a Stop hook blocked the previous turn, so branch on it and let the retry through, or escalate to a different message instead of blocking again.

Can a hook change what Claude is about to do instead of blocking it?

Yes, on PreToolUse. Return hookSpecificOutput with a permissionDecision and an updatedInput object, and Claude Code runs the modified call. That is the clean way to sanitise a command or redirect a path, and it beats blocking because the agent is not left guessing what you wanted.

Are hooks better than CLAUDE.md or a skill?

They answer different questions. CLAUDE.md and a skill shape what the agent decides to do, and a skill is where the judgment lives about which check a change needs. A hook is deterministic enforcement and knows nothing about the work. A verification setup needs both: the skill decides what proof means, the hook makes sure it happened.

How slow can a hook be before the team turns it off?

Faster than you think. Command, http and mcp_tool handlers default to a 600 second timeout, which is far longer than anyone will accept on every turn, so set your own and keep the check to seconds. For a genuinely slow check use async of true to run it in the background, with asyncRewake to wake Claude if it exits 2.

No spam. Just a practical audit.

Ready to remove your biggest software bottleneck?

Book a free 15-minute call. We will help you identify the highest-leverage automation, API integration, AI agent, or internal system to build first so your team can move faster with less manual work.

© 2026 Bles Software, Yehud-Monoson, Israel. All Rights Reserved.