Every list ranks skills on how the author felt. A skill either changes an outcome you can count, or it does not.
Search for the best Claude Code skills and you get lists of names. The first page of results on 8 October 2026 was one developer's ten, a newsletter's six out of thirty-three tried, a Reddit thread, a curated GitHub list, a directory of community skills and a few videos, with Google's own summary stacked on top naming another six. Good for discovery. None of them tells you whether a skill changed anything, because none of them ran the same task twice.
That is the gap this page fills. We run coding agents on our own production work every day behind a gate that reruns the check instead of reading the agent's summary of it, and the two skills we give away publish their benchmarks. Below is the test a skill has to pass before it earns a place in our stack, what it said about ours, and the order we would install them in.
The skill that decides what proof a change needs, the same skill as a zip, the front-end half, the two pages this one grew out of, our free library, and the pack they all come from.
One of the two benchmarked below. Before a multi-file change it looks for what the change breaks outside its own diff, then proves the one fact that makes it safe by running code rather than describing it.
Get it on GitHub →A single skill file in a zip, if you would rather not clone anything. Drop it into .claude/skills/ and it loads in the next session. No signup, no email.
Download the zip →The front-end half, and the other benchmark on this page. Its gate opens the built page in a real browser at phone and desktop widths and scores it before the agent may call the work done.
Get it on GitHub →Why an agent reports done when the work is not, how to write a check that can fail, and the practices we would set up on a team in order.
Read the practices →The enforcement half. The blocking events, the exit code contract, and the Stop hook that refuses to let a turn end before the check has run.
Read the hooks reference →Production skills we publish as plain Markdown, grouped by function, each one a file you can read in full before you install it.
Browse the library →The proof gate, the reviewers, the build loops and the visual and video checks we run ourselves, packaged for Claude Code and Codex. 99 USD once, lifetime updates.
See the pack →Senior engineers who ship with coding agents behind a proof gate, joining your team rather than replacing it.
Hire AI developers →The lists are real work and worth skimming. One developer publishes the ten he uses at work, a newsletter publishes its six after trying thirty-three, Reddit carries a long thread of recommendations, and a curated GitHub list plus a directory of a couple of hundred community skills exist to be browsed. Google's answer box stacks them into another six names with a line each on what they are best for.
What none of them carries is a measurement. Every entry is a description of what the skill claims to do, graded by how the author felt while using it, which is the only grade available when nobody ran the same task twice. Take the top six from any two of those lists and you get a different top six, because a star count and a mention count are measuring attention rather than output.
There is a second problem that matters more once a few are installed. A skill is loaded into context by the model's own judgment that it is relevant, so it is advice, and on a long session the agent can drift past it. A skill that reads beautifully and a skill that changes an outcome look identical in a list, and they look identical in your terminal too, right up to the moment you measure.
Pick one outcome you can count. Not better code. Something a script can read: whether the broken file was named, how many findings a detector raises on the page, whether the readback exited zero. If you cannot write the scorer, you cannot judge the skill, and that is worth finding out before you install twenty of them.
Then run the same task twice, with the skill and without it, on the same model and the same repository, several runs each way, because a single run of a coding agent is mostly noise. Disable every other skill in both arms. Score it with something the skill does not run, because a skill judged by its own checker will pass its own checker.
Keep the script. What you own at the end of that afternoon is not an opinion about one skill, it is a scorer you can point at the next one, and at a new model version six weeks from now when the plain arm quietly gets better and the skill stops earning its seconds.
That is also the honest reason we publish two benchmarks and not twenty. Running one properly costs a few hours and real API money, and a number we did not measure is worth exactly what the lists are worth.
The backend one first. Claude Sonnet 5 in Claude Code, the same repository both ways, one renamed dictionary key in a pricing module that breaks the file importing it, three runs an arm. Plain Claude Code named the broken file in 1 of 3 runs and ran code to check it in 0 of 3. With the skill it was 3 of 3 on both, producing the actual KeyError rather than a paragraph predicting one. That cost about 14 seconds and 0.09 USD a run, against about 5 seconds and 0.06 USD without it.
The front-end one is scored by a skill tastegate never runs, which is the part worth reading. Same landing-page brief, Claude Opus 5.5, five pages an arm, every other skill disabled on both sides, judged by another design skill's own detector. Plain pages averaged 29.2 findings a page and tastegate pages 2.2, which is 13 times fewer. Its own browser gate, which opens the built page at phone and desktop widths, failed plain pages 45.4 times a page and tastegate pages 0 times. The extra cost was about 1.1 minutes and 0.24 USD a page: 2.9 against 4.0 minutes, 0.65 against 0.89 USD.
Both repositories ship the prompts, the per-run scores and the scorer, which is the only part of this you should actually trust. Rerun them. The arms are the comparison we think is fair on our brief, and if you disagree with the brief you can change it and get your own number instead of ours.
A benchmark tells you whether a skill helped on a task someone else chose. The question that survives into daily work is narrower: when the agent reported the work done, was it? Our gate holds a requirement list for every request, and for each requirement the lane has to write what it did and the command that proves it. The gate then runs that command itself.
Over the eight days from 1 to 8 October 2026 that came to 301 real requests and 2,821 QA checks, and 258 of the pass claims were not confirmed by the gate's own rerun. 62 were a readback command that exited non-zero the moment the gate ran it rather than reading the summary of it. The other 196 were a rendered screen the scanner refused. 282 of the 301 requests finished with every requirement passing.
The rendered half is where most skill stacks stop, so it is worth breaking open. 996 of those checks looked at a screen rather than a stored value: 608 in a web page, 241 in the iOS app, 94 on the Mac app, 49 on Android, and 4 whose platform the row did not record. A skill pack that only knows how to run your test suite cannot reach any of those, which is the practical reason the visual and device checks are half of ours.
A gate is also a program, so it can be wrong, and the honest number matters more than the impressive one. The most common reason ours refused a claim was not that it found a defect. 122 of the 196 rendered refusals were the scanner reporting that it could not judge the screenshot at all, too little legible text or a page it could not probe, and an unjudged screenshot is not evidence, so the claim was not accepted on trust. 59 were text drawn over other text, and 15 were text hidden under another element or clipped by a box edge. 226 flagged rows came back with a written judgment of what the crop actually showed, which is why a flag on our gate asks for a look rather than failing the work by itself.
Ordered by what each one changes, not by stars. First, the skill that decides what proof a change needs, because a gate without it runs the wrong check confidently. what-could-break is the one we give away, and its numbers are above.
Second, a verifier that drives your own application the way a user does, in your stack, against your fixtures. This is the one nobody can hand you finished, because it is specific to your app, so what ships instead is the skill that writes it for a repository and leaves a script behind.
Third, the front-end gate, if you ship an interface at all. A design skill with no browser check is a style guide, and the agent grades its own homework against it. tastegate is free and its numbers are above.
Fourth, a review pass that is allowed to say no, and then a skill creator, so that the second time your team makes the same correction it becomes a file instead of another paragraph in CLAUDE.md.
The well-known open skills are worth having alongside these. The methodology packs that enforce a spec, a plan and a test-driven loop, the official skill creator, and a front-end design skill all do real work, and the lists above will point you at them. The difference is not quality. It is that nobody has published a number for them, so you will have to run the afternoon yourself.
A skill shapes what the agent decides to do, and the same model decides whether to follow it. That is a ceiling no skill can lift, however well written, and it is the reason a pack of skills alone never fixes a confident done.
The refusal has to come from outside the model. In Claude Code that is a hook: a script the runtime invokes at a fixed point, so it happens every time. The Stop hook fires when Claude finishes responding, and exit code 2 prevents the stop, which puts your failure text in front of the agent instead of ending the turn. Our hooks page is the full contract, events, exit codes and the loop guard included.
The pair is the whole point. The skill works out what proof means for this particular change, and the hook makes sure that proof happened before the turn was allowed to end. Install only the first half and you have better advice. Install only the second and you have a gate running the wrong check.
Picking skills well is an afternoon. Having a stack that holds on a team is a different job: the gate on your repositories, a verifier that drives your own app, a front-end check at the widths your users actually hold, and a Stop hook fast enough that nobody quietly disables it.
Everything we run ourselves ships in SWE Stack for 99 USD once, and the two free skills on this page are the same code as the ones inside it, so you can judge the pack by its free half before paying for anything. If you would rather have the whole thing standing on your repositories and handed over running, that is the part we do.
Four steps in this order. One developer gets through all four on one repository in an afternoon, and what you keep at the end is the scorer, not the verdict.
Name the number before installing anything: the broken file named or not, findings per page, the readback's exit code. If no script can read it, the skill cannot be judged and you are back to how it felt.
Several runs each way, with the skill and without, every other skill disabled on both sides. One run of a coding agent is noise rather than a result.
Another tool's detector, a test suite the skill never sees, or a readback it cannot write. A skill marking its own work will pass its own marking.
Rerun it when the model version changes. The plain arm improves on its own, and a skill that has stopped earning its seconds should leave the stack.
The ones that change an outcome you can count on your own repository, which is not the same set for every team. Instead of another top six, this page gives the test: pick a countable outcome, run the same task with the skill and without it several times on the same model, and score it with something the skill does not run. Two with that work already published are what-could-break and tastegate, both free and MIT.
Measure it. Same model, same repository, several runs an arm, every other skill disabled on both sides, and a scorer the skill never runs. A star count measures attention and a description measures writing, so neither tells you what changed.
Anthropic's own documentation covers where a skill file lives and how it loads, and the plugin marketplace installs them from inside the session. Beyond that, curated lists on GitHub and a few directories collect community skills, and our own library publishes ours as plain Markdown you can read before installing.
Fewer than the lists suggest. Every installed skill competes for the model's judgment about which one is relevant right now, so a stack of twenty loosely scoped skills behaves less predictably than four that each do one thing and can be measured. Add one, measure it, then keep it or remove it.
A skill is instructions the model chooses to load when it judges them relevant, so it shapes what the agent decides to do. A hook is a script Claude Code runs itself at a fixed point, so it happens every time and on a blocking event it can refuse. A plugin is the package that distributes either. Verification needs the first two together.
The instruction half ports, because a skill is a Markdown file with frontmatter and any agent that reads a project file can read it, and both free skills on this page do. What does not port by itself is the enforcement, since each tool has its own hook or gate mechanism, so the refusal has to be wired up per tool.
Not on its own, and that is the most common disappointment with skill packs. A skill can say what proof to gather, but the same model decides whether to follow it. The refusal has to come from outside the model, which is the Stop hook's job: it fires when Claude finishes responding, and exit code 2 prevents the stop.
Only on the same test. Ask what the pack measures, whether its numbers can be reproduced, and whether its free half is good enough to judge the rest by, which is why both of ours are public and MIT. Ours keeps the same code in both halves: the two free skills here are the versions inside the pack, so its price and contents are one click away on the SWE Stack page and you can judge the rest by the free half first.
Book a free 15-minute call. We will help you identify the highest-leverage automation, API integration, AI agent, or internal system to build first so your team can move faster with less manual work.
About Us
Features
Testimonials
Contact Us
© 2026 Bles Software, Yehud-Monoson, Israel. All Rights Reserved.