Agent verification skills
Agent verification skills: proving a change actually works
Section titled “Agent verification skills: proving a change actually works”Upstream: 2026-10-08 Agent skills for splitting work into issues. Its per-unit “Verification / Test scenarios” field and Pocock’s “observation that would show it false” are where the oracle gets fixed before any work starts.
Gate: 2026-10-08 PM worker reviewer agent orchestration. Its §4 review-gate checklist (worker report contract → independent reviewer check → --match-head-commit) is the slot this note fills with concrete recipes and a record format.
Design gate: 2026-10-08 Design-first feature gate — pre-implementation design sign-off and post-implementation conformance review for new features.
PR body: 2026-10-09 Design PR description for humans — where the verification record sits in a skimmable PR body (below the fold) and how artifacts get inline.
Related: 2026-10-08 yetone magpie agent workflow (the “Not run:” record and gh pr merge --match-head-commit) · 2026-10-08 Agent lessons ledger lineage · 2026-09-28 Opinionated skill packs pstack shape (create/maintain-verification-skill) · 2026-09-28 Agent e2e verification Wayland browser computer-use · 2026-10-02 Agent screenshot and UI expression
Pinned 2026-10-08 ~23:40 (UTC+8). Skill files were read from shallow clones (default-branch HEAD as of today) or through cursor-github get_file_contents. Star counts come from the GitHub API at pin time, and only where I checked them. X dates are converted to UTC+8. Codex was not used.
1. Ranked picks
Section titled “1. Ranked picks”#1 Codewhale cw-gates → cw-dogfood → cw-land + AGENTS.md “Merging under a gate”
Section titled “#1 Codewhale cw-gates → cw-dogfood → cw-land + AGENTS.md “Merging under a gate””codewhale-hq/Codewhale docs/skills (cw-gates, cw-dogfood, cw-land, AGENTS.md). 41,082★, MIT, Rust, last pushed 2026-10-08 21:49 (UTC+8). The chain is cw-orient → cw-slice → cw-gates → cw-dogfood → cw-land → cw-handoff.
What it does.
cw-gatesis a 5-rung check ladder: climb only as far as the risk needs, and quote the real output.cw-dogfoodbuilds a binary stamped with the HEAD SHA, installs it, and drives the real product from a fresh login shell.cw-landchecks mergeability against the real head withgit merge-treeand needs an artifact that literally says PASS.
“Assertions without command output are not evidence.” / “Climb only as far as the risk requires. Say where you stopped and what you skipped.” / “Quote the real
test result: N passed; M failedline, and confirmN > 0.” / “Prefer proving a regression test fails without the fix. A test that passes either way pins the implementation, not the defect.” (cw-gates)“Green gates prove the code compiles and asserts. They do not prove the product works.” / “The version string must contain the short HEAD SHA you just built.” (cw-dogfood)
“A gate is its artifact… the record must literally say PASS at merge time.” / “Read the review thread, not the check rollup.” / “When the artifact is ambiguous, resolve the ambiguity — never the merge.” (AGENTS.md)
Evidence produced.
- A checklist with each command, pass/fail and its salient line, plus “Name explicitly what you did not run and why.”
- A stamped build (
CODEWHALE_BUILD_SHA=$(git rev-parse HEAD)). The installer refuses an unstamped or dirty tree. - Dogfood output that lists the scenarios not exercised.
- Evidence goes in the commit message and PR body, because “Agents do not comment on issues or PRs.”
Pros:
- Covers all four of Noa’s asks in one pack: causation, the run/not-run record, the real path, and a gate tied to the head.
- Catches the zero-tests-ran trap (“
cargo test <filter>exits 0 having run zero tests”). - Says outright “Don’t report a green gate as permission.”
- Has a line on visual proof: “A screenshot proves layout and color; only live observation or a recording proves motion.”
- Asks before spending provider tokens.
- Plain markdown, so cheap to port.
Cons:
- The commands are Rust/CLI-specific (
cargo, a stamped installer), and the web/E2E path is thin. - AGENTS.md says “never practice TDD here — this overrides… superpowers test-driven-development”. It still demands that a test written after the fix be shown failing without it, so it is test-after plus fail-without-fix, not red-first.
- No mutation step.
- The SHA stamp only makes sense for things that build into an artifact.
Gate fit.
- The worker fills the cw-gates checklist plus the dogfood “not exercised” list in the PR body.
- The reviewer or Freezy applies “a gate is its artifact”: read the record on the PR at the head SHA, not the check rollup.
- Ambiguous or missing evidence means INCONCLUSIVE (this matches the PM digest’s §4). Then merge with magpie’s
--match-head-commit.
Vs pstack / magpie.
- pstack
create-verification-skillcovers the Drive/Evidence half: how to launch and drive this app, and real-user-path standards. Codewhale adds what pstack lacks: fail-without-fix, a not-run list, and a SHA tie-in. - Magpie’s record (“Not run: …”, “Reviewed at head…”,
--match-head-commit) is the same idea, written from the reviewer’s side. Codewhale is the worker-side and pre-merge version, and is more explicit about N>0 and artifact literalness.
#2 superpowers verification-before-completion (+ alpine review-pr stash recipe, + dsh negative controls)
Section titled “#2 superpowers verification-before-completion (+ alpine review-pr stash recipe, + dsh negative controls)”obra/superpowers verification-before-completion (MIT, 8ca22db 2026-09-26 02:06 UTC+8), paired with test-driven-development. Reviewer-side counterpart: alpinejs/alpine .claude/skills/review-pr. Negative-control rule: deepseek-harness dsh-ci-test-reliability.
What it does. A completion gate: IDENTIFY the proving command → RUN it fresh → READ the output → VERIFY it supports the claim → ONLY THEN claim. A table maps each claim to the evidence it needs.
“NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE” / “Skip any step = lying, not verifying” / “Regression test works | Red-green cycle verified | Test passes once” / ”✅ Write → Run (pass) → Revert fix → Run (MUST FAIL) → Restore → Run (pass)” / “Agent reports success → Check VCS diff → Verify changes → Report actual state”
alpine: “Actually verify regression. Don’t just reason about whether the test fails without the fix — prove it. Stash the fix (
git stash -- <fix files>), rebuild (npm run build), run the test. If it passes without the fix, the test is not testing the fix… This is non-negotiable for bug fix PRs.”dsh: “For a new static or corpus guard, temporarily introduce the rejected case and observe the intended failure… Report exact commands and observed results; do not describe retries, skipped tests, or pending CI as passing.”
Evidence produced. Fresh command output quoted in the claim, and a revert→fail→restore→pass cycle. Alpine posts a <!-- claude-review --> verdict comment with Test Results, and a human (Caleb) merges.
Pros:
- The clearest causation recipe, and the most widely installed (superpowers was ~296k★ per the PM digest).
- Treats sub-agent self-reports as claims to check against the VCS diff.
- The alpine recipe is the same check done by the reviewer, and is the easiest independent check for a verifier agent.
- dsh adds negative controls for new guards and lints, plus “report pending checks as pending”.
Cons:
- A self-check by the same agent: no commit SHA, no “Not run:” section, no artifact format.
- No E2E or user-path guidance.
- Alpine says “NEVER run the full test suite” (a repo-specific speed choice) and is a maintainer bot, not a reusable pack.
Gate fit.
- The worker runs the revert-fail-restore cycle and pastes it.
- An independent verifier (cheap second cloud agent) repeats alpine’s
git stash -- <fix files>→ rerun at the reviewed head. - Freezy can’t run code, so it only checks that the verifier’s record names the head SHA and shows a real FAIL. Per magpie, a compile error or a skipped test doesn’t count.
Vs pstack / magpie.
- This is magpie’s “a test that fails without the change… real FAIL” and “Break it on purpose”, packaged as a portable skill.
- pstack doesn’t do causation at all.
- deslop already has “deliberate break… suite must go red, then green with the fix” for test changes. This generalises it to every behaviour change.
#3 Every ce-dogfood / ce-test-browser (+ Showboat verify for re-checkable proof)
Section titled “#3 Every ce-dogfood / ce-test-browser (+ Showboat verify for re-checkable proof)”EveryInc/compound-engineering-plugin ce-dogfood and ce-test-browser (MIT, 67035e9 2026-10-08 07:40 UTC+8). Re-checkable evidence: simonw/showboat (1,231★, Apache-2.0, last push 2026-03-15 UTC+8).
What it does.
ce-dogfoodis diff-scoped. It maps the affected flows (Mermaid), builds a persona × scenario matrix, and drives each scenario through theagent-browserCLI.- It fixes what breaks, with a regression test per fix.
- It writes a committed report at
<root>/dogfood-reports/<YYYY-MM-DD>-<branch-slug>-dogfood.md, which doubles as a resume checkpoint. It never pushes. ce-test-browseris the lighter per-route Pass/Fail/Skip pass.
“Done: every matrix scenario is Pass, Fixed, Skipped, or in a terminal Blocked state; the project’s automated suite has been run once and its result recorded… A green matrix over a red suite finalizes as a not-ready verdict” / “A fix is not done until a regression test fails before it and passes after, or the report says why no automated test was meaningful.” (ce-dogfood)
“every affected route marked Pass, Fail, or Skip and each Skip carrying its reason… dropping a route from the summary because nobody could reach it, is the failure” (ce-test-browser)
Showboat: “A verifier can re-execute all code blocks and confirm the outputs still match.”
verify“Re-runs every code block… Prints diffs and exits with code 1 if any output has changed”. Simon: “Theexeccommand… is designed to discourage the agent from cheating and writing what it hoped had happened into the document.”
Evidence produced.
- The report sections are: Diff Summary, Personas, Flows, Test Matrix (Status / Issue / Fix / Commit), What Was Fixed (regression test fails before / passes after), Console Errors, Human Verifications, Decisions for a Human, and Final Status.
- Anything needing OAuth, email, payments or SMS is marked
Blocked (needs human verify). - Showboat adds a markdown proof doc whose
execblocks a reviewer can re-run withshowboat verify.
Pros:
- The best real-user-path discipline found. A Skip has to carry a reason, so nothing quietly drops out, which is the not-run record applied to E2E.
- Ties the matrix to the suite result and to per-fix commits.
- Has a “do not hard-code main” rule for the diff base.
ce-workalready assumes “the host orchestrator inspects actual changes and owns authoritative verification”, which is Noa’s model.- Showboat is the only artefact in this survey that a reviewer can mechanically re-run.
Cons:
- Heavy: a multi-phase reference set, and it needs agent-browser.
ce-test-browsersays “Do not introduce a third browser stack.” - Screenshots go to OS temp, so they’re lost on a cloud VM unless they get attached.
- The report records commits per fix but not the final reviewed head.
- Showboat
verifyonly re-runs deterministic shell output. It can’t re-run browser clicks, and it hasn’t been pushed since March.
Gate fit.
- A cloud worker runs a trimmed ce-test-browser pass (affected routes only) using Cursor’s computer use. Screenshots and video are attached to the PR as Cursor artifacts (capabilities).
- The PR body carries the matrix with reasons for every Skip.
- Freezy checks the matrix for unexplained gaps and for any “green matrix over red suite”.
- A verifier re-runs
showboat verifywhere CLI proof exists. Browser paths still need CI Playwright or a second agent run.
Vs pstack / magpie.
- pstack’s
verify-<app>skill tells the agent how to launch and drive the app. ce-dogfood tells it what to cover for this diff and how to report it. They stack:verify-<app>as the driver, a ce-dogfood-style matrix as the record. - Magpie’s
ui-preview.yml(DeepSeek plans scenes, Playwright records into the PR body, split-trust jobs) is the CI version of the same evidence, and it is stronger as a gate because the worker doesn’t produce it.
2. Other candidates
Section titled “2. Other candidates”| Candidate | What it adds | Causation | Run/not-run | Real path | Independent gate | Fit for cloud worker / PM bot | Port effort | License · activity |
|---|---|---|---|---|---|---|---|---|
deepseek-harness dsh-pre-push-checks | ”narrowest… check that would fail for its regression”; check the selected test count; git rev-parse HEAD origin/<branch> then gh pr checks; “Commit hashes… from before the rewrite are not current evidence” | ✅ | ✅ pending ≠ pass | ✗ | ✅ remote-ref = HEAD | Good worker pre-push step | Low | 5badb15 2026-10-03; already covered in the lineage digest |
mattpocock tdd | ”Red before green.”; “Expected values must come from an independent source of truth”; “Test only at pre-agreed seams” | ✅ red-first | ✗ | ✗ | ✗ | Pairs with to-tickets seams | Low | MIT · b0618bc today |
anthropics/skills webapp-testing | with_server.py + Python Playwright; “Reconnaissance-Then-Action” | ✗ | ✗ | ✅ | ✗ | Driver only | Low | LICENSE.txt in skill · 683bc88 |
claude-code pr-review-toolkit pr-test-analyzer | Rates test gaps 1–10; “good tests are those that fail when behavior changes unexpectedly” | Reasoned, not run | ✗ | ✗ | Read-only review | Fine for Freezy (no exec), weak as proof | Low | 71cddde today |
vercel-labs/agent-browser dogfood | Exploratory QA with record start …webm, per-step screenshots, console/errors; “Verify reproducibility before collecting evidence”; “Never read the target app’s source code” | ✗ (finds bugs) | ✗ | ✅ | Blind-tester stance | Good second-agent QA | Low | 0207911 today |
ChromeDevTools chrome-devtools skill | navigate → wait → take_snapshot (uid) → interact; screenshots, CSS, a11y/LCP/memory siblings | ✗ | ✗ | ✅ | ✗ | Driver | Low | 2744afa today |
| Playwright Test Agents | planner → specs/*.md, generator “verifies selectors and assertions live”, healer re-runs until green | ✗ | ✗ | ✅ CI-runnable | ✅ if specs run in CI | Output is real .spec.ts that CI can gate | Medium | Healer may skip a test “if the healer believes that functionality is broken”: an oracle-edit risk |
| Cursor cloud agent computer use (changelog, blog) | “start dev servers, open the app in a browser, click through UI flows, and verify their changes work before pushing a PR”; screenshots/videos/logs attached to PR | ✗ | ✗ | ✅ | ✗ produced by worker | Native to Noa’s workers | None | ”Allow posting artifacts to GitHub” uses public unguessable URLs; CI autofix is GitHub Actions only |
| Anthropic long-running harness | feature_list.json "passes": false; “It is unacceptable to remove or edit tests”; init.sh + e2e smoke each session | Partial | Feature list | ✅ Puppeteer | ✗ | Pattern, not a skill | Low | 2025-11-26 post |
| simonw guides (first-run-the-tests, agentic-manual-testing) | “confirm that the tests fail before implementing”; “Never assume that code generated by an LLM works until that code has been executed.” | ✅ | ✗ | ✅ (Rodney, Showboat) | Showboat | Prompt lines | Trivial | Blog |
| keyboardsamurai/mutagate | ”Mutation testing as a hook: blocks coding agents until their tests catch the bugs” | ✅ mutation | ✗ | ✗ | Hook, local | Worker-side only | Medium | MIT · 2★ · created 2026-10-06 |
| lkc-studio/claude-skill-mutants | Diff-scoped mutate.py --since main, non-zero exit on survivors, triage reference | ✅ | ✗ | ✗ | CI-runnable | Could run as a CI job | Medium | Stars not verified |
| mikemartincode/mutation-gate, xiaolai/tdd-guardian-for-claude | mutmut pre-commit gate; TaskCompleted hook + optional mutation gate | ✅ | ✗ | ✗ | Hook | Local only | Medium | Stars not verified |
| duoglas/simple-harness-kit | .harness/verify-evidence.json “binds Git commit/tree/dirty state… SHA-256 digest”; reports DEGRADED honestly | ✗ | ✅ | ✗ | ✅ SHA-bound | Idea worth stealing | Low | 0★ |
| dsifry/metaswarm | Orchestrator verifies itself, never trusts the worker’s self-report (see PM digest) | — | — | — | ✅ | Already in PM digest | — | 33d39f7 2026-06-20; not re-read |
Signal from X and HN (context, not packs):
- @poteto 2026-07-31 / 2026-08-29: the pstack verification skill is “the most important part of creating an agent loop you can trust”. @RenanBattu asked the right follow-up: is it a runnable harness or a rubric graded by the same agent?
- @johnroodepic 09-27: “only counts if it fails without the fix. keep expected outputs wired to a source the agent can’t edit”. @inferencegod 06-18: “the bite: revert the fix, rerun the test”. @simonlo 10-06: the
git stash push -- src/check. - @kanjun 03-06: “Claude Code lied about tests passing, but didn’t run them at all”. @DanylenkoM 09-25: 100% branch coverage, but only 5 of 12 planted bugs caught. @gokulmenons 09-24: agents edit assertions. @crazymonkeysai 10-06: “a check the agent cannot write or skip”.
- @ChrisSimpson 09-16: cloud agents attaching review videos to PRs. @keyboardsamurai 10-07: mutagate launch.
- Kent Beck: Genie Wants to Leap (“deleting assertions from tests, deleting whole tests, & faking large swathes of implementation”) and Augmented Coding: Beyond the Vibes (watch for “Any indication that the genie was cheating, for example by disabling or deleting tests”).
- HN was quiet: CodeLeash (TDD state machine, 12 pts), rootcause/autofix, KeelTest, Ask HN: How to Claude Like Anthropic.
3. Synthesized verification record (template)
Section titled “3. Synthesized verification record (template)”Sources: magpie’s record + Codewhale cw-gates/dogfood + ce-dogfood matrix + dsh pre-push + PM digest §4. The worker pastes it into the PR body. The reviewer appends its own block at the reviewed head.
## Verification recordHead: <full sha> Base: <base sha / branch> Pushed ref = HEAD: yes (`git rev-parse HEAD origin/<branch>`)Oracle: <issue AC / test file / spec> — not edited by this PR: yes | no (explain)
### Causation (per behaviour change)- <change> → test `<path::name>` - without fix: FAIL — `<salient line>` (cmd: `git stash -- <fix files> && <test cmd>`) - with fix: PASS — `<salient line>` - (compile error / skipped / 0 tests selected ≠ FAIL)- New guard/lint: negative control introduced → observed failure `<line>`- Mutation (optional): <tool> on diff — <k>/<n> killed; survivors: <list or "none">
### Checks run (exact commands, N>0)| Command | Result | Salient line ||---|---|---|| `<cmd>` | PASS/FAIL | `test result: N passed; M failed` (N=…) |
### Real user path| Route / scenario | Driver (Playwright / browser / CLI / API) | Status (Pass/Fail/Skip/Blocked) | Evidence (screenshot / video / log link) | Reason if Skip/Blocked ||---|---|---|---|---|Build under test reports SHA: <short sha> (stamped build / version endpoint), or "n/a: <why>"
### Not run- <check> — <why> (e.g. needs secrets, OAuth, payments, slow suite, provider tokens)
### CI at head<check name>: success/failure/pending @ <sha> (pending is pending, not pass)
---## Reviewer check (independent, at head <sha>)- [ ] Record head == PR head now (else stale → INCONCLUSIVE)- [ ] CI check runs for <sha> all success (list_check_runs_for_ref), none skipped silently- [ ] Causation re-run by verifier: stash fix → FAIL observed / not re-run (why)- [ ] Oracle untouched: diff of tests/fixtures/specs reviewed; no deleted/weakened assertions, no new skips- [ ] Every AC maps to a check or a user-path row; every gap is in "Not run"- [ ] Artifacts open and match claims (screenshot/video shows the asserted state)Verdict: PASS / FAIL / INCONCLUSIVE — merge only with `gh pr merge --match-head-commit <sha>`Rules that go with it:
- The artifact has to literally say PASS (Codewhale).
- Missing evidence means INCONCLUSIVE, never PASS (PM digest §4).
- A Cursor worker’s own screenshots count as claims. Only CI or a separate verifier counts as the gate.
4. Overlap with Noa’s skills and where a mirror would sit
Section titled “4. Overlap with Noa’s skills and where a mirror would sit”- Already covered:
deslop“Clean with proof” (deliberate break → red → green; hand-backproof: <test/command/result>)repo-sloppiness(evidence cites file:line or a command)project-map(“Do not treat green cards as truth without a test/CI/demo link”)lessons-ledger(“verified too thinly?”; back a promoted rule with a test or lint)critique(implementer AC to tick)project-pmonboarding step 3 (offer pstack/create-verification-skill, only on an explicit yes)
- Missing: a record format (as the magpie digest noted), a fail-without-fix step outside deslop’s test-touching case, an N>0 check, head-SHA binding, and a reviewer-side gate checklist that Freezy can run without executing code.
- Where a mirror would sit (proposal only; no skill was created or edited):
- Worker side: a
verification-recordskill that cloud-agent briefs reference. It holds the §3 template plus the cw-gates ladder and the superpowers revert-fail-restore cycle, and calls the repo’s pstackverify-<app>skill as its driver when one exists. - Reviewer side: a
pm-gatesection insideproject-pmor a sibling. It holds the reviewer checklist from §3, thelist_check_runs_for_refread, the dispatch of a cheap verifier agent for the stash re-run, and the--match-head-commitmerge (Noa still merges). - Feedback:
lessons-ledgeralready asks “verified too thinly?”. Have it read the “Not run” sections as input.
- Worker side: a
5. Gaps
Section titled “5. Gaps”- SHA binding is rare. Only magpie’s
--match-head-commit, Codewhale’s stamped build and the 0★ simple-harness-kit bind evidence to a commit. No popular skill pack makes the worker record the head SHA. - Freezy can’t re-run code, so independence depends on CI or a second agent. Nothing found packages “verifier agent re-runs the stash check at head” as a reusable skill. Alpine’s maintainer bot is the closest.
- Browser evidence isn’t re-checkable. Cursor artifacts, agent-browser videos and ce-dogfood screenshots are all produced by the worker. Showboat
verifyonly covers deterministic shell output. Magpie’s CIui-preview.ymlis the only split-trust UI evidence seen. - Oracle protection is only stated as prose (“It is unacceptable to remove or edit tests”; mattpocock’s “independent source of truth”). No pack checks it mechanically, and the Playwright healer can skip tests.
- Mutation gates are tiny and new (mutagate 2★, created 2026-10-06; the others unverified). Nothing is mature enough to adopt as-is.
- Not re-read this run: the OpenAI harness-engineering post (403 to WebFetch), metaswarm and pstack’s own files (relied on prior digests). Reddit wasn’t searched because Exa returned HTTP 417. Kanjun’s “Vet” repo wasn’t verified.
6. Sources
Section titled “6. Sources”- Codewhale: docs/skills · AGENTS.md
- obra/superpowers: verification-before-completion · test-driven-development · systematic-debugging
- alpinejs: review-pr skill
- deepseek-harness: dsh-pre-push-checks · dsh-ci-test-reliability
- EveryInc: ce-dogfood · ce-test-browser · ce-work
- Showboat and Simon Willison: simonw/showboat · red-green-tdd · first-run-the-tests · agentic-manual-testing
- mattpocock: tdd
- Anthropic: webapp-testing · pr-test-analyzer · effective harnesses for long-running agents
- Browser drivers: agent-browser dogfood · chrome-devtools skill · Playwright Test Agents
- Cursor: cloud agent capabilities · changelog 02-24-26 · agent computer use
- Mutation and evidence tools: mutagate · claude-skill-mutants · mutation-gate · tdd-guardian-for-claude · simple-harness-kit
- Kent Beck: Genie Wants to Leap · Augmented Coding: Beyond the Vibes
- X and HN: see the links in §2.
- Vault: 2026-10-08 yetone magpie agent workflow · 2026-10-08 Agent lessons ledger lineage · 2026-10-08 PM worker reviewer agent orchestration · 2026-10-08 Agent skills for splitting work into issues · 2026-09-28 Opinionated skill packs pstack shape