跳转到内容

Agent verification skills

Agent verification skills: proving a change actually works

Section titled “Agent verification skills: proving a change actually works”

Upstream: 2026-10-08 Agent skills for splitting work into issues. Its per-unit “Verification / Test scenarios” field and Pocock’s “observation that would show it false” are where the oracle gets fixed before any work starts. Gate: 2026-10-08 PM worker reviewer agent orchestration. Its §4 review-gate checklist (worker report contract → independent reviewer check → --match-head-commit) is the slot this note fills with concrete recipes and a record format. Design gate: 2026-10-08 Design-first feature gate — pre-implementation design sign-off and post-implementation conformance review for new features. PR body: 2026-10-09 Design PR description for humans — where the verification record sits in a skimmable PR body (below the fold) and how artifacts get inline.

Related: 2026-10-08 yetone magpie agent workflow (the “Not run:” record and gh pr merge --match-head-commit) · 2026-10-08 Agent lessons ledger lineage · 2026-09-28 Opinionated skill packs pstack shape (create/maintain-verification-skill) · 2026-09-28 Agent e2e verification Wayland browser computer-use · 2026-10-02 Agent screenshot and UI expression

Pinned 2026-10-08 ~23:40 (UTC+8). Skill files were read from shallow clones (default-branch HEAD as of today) or through cursor-github get_file_contents. Star counts come from the GitHub API at pin time, and only where I checked them. X dates are converted to UTC+8. Codex was not used.


#1 Codewhale cw-gates → cw-dogfood → cw-land + AGENTS.md “Merging under a gate”

Section titled “#1 Codewhale cw-gates → cw-dogfood → cw-land + AGENTS.md “Merging under a gate””

codewhale-hq/Codewhale docs/skills (cw-gates, cw-dogfood, cw-land, AGENTS.md). 41,082★, MIT, Rust, last pushed 2026-10-08 21:49 (UTC+8). The chain is cw-orient → cw-slice → cw-gates → cw-dogfood → cw-land → cw-handoff.

What it does.

  • cw-gates is a 5-rung check ladder: climb only as far as the risk needs, and quote the real output.
  • cw-dogfood builds a binary stamped with the HEAD SHA, installs it, and drives the real product from a fresh login shell.
  • cw-land checks mergeability against the real head with git merge-tree and needs an artifact that literally says PASS.

“Assertions without command output are not evidence.” / “Climb only as far as the risk requires. Say where you stopped and what you skipped.” / “Quote the real test result: N passed; M failed line, and confirm N > 0.” / “Prefer proving a regression test fails without the fix. A test that passes either way pins the implementation, not the defect.” (cw-gates)

“Green gates prove the code compiles and asserts. They do not prove the product works.” / “The version string must contain the short HEAD SHA you just built.” (cw-dogfood)

“A gate is its artifact… the record must literally say PASS at merge time.” / “Read the review thread, not the check rollup.” / “When the artifact is ambiguous, resolve the ambiguity — never the merge.” (AGENTS.md)

Evidence produced.

  • A checklist with each command, pass/fail and its salient line, plus “Name explicitly what you did not run and why.”
  • A stamped build (CODEWHALE_BUILD_SHA=$(git rev-parse HEAD)). The installer refuses an unstamped or dirty tree.
  • Dogfood output that lists the scenarios not exercised.
  • Evidence goes in the commit message and PR body, because “Agents do not comment on issues or PRs.”

Pros:

  • Covers all four of Noa’s asks in one pack: causation, the run/not-run record, the real path, and a gate tied to the head.
  • Catches the zero-tests-ran trap (“cargo test <filter> exits 0 having run zero tests”).
  • Says outright “Don’t report a green gate as permission.”
  • Has a line on visual proof: “A screenshot proves layout and color; only live observation or a recording proves motion.”
  • Asks before spending provider tokens.
  • Plain markdown, so cheap to port.

Cons:

  • The commands are Rust/CLI-specific (cargo, a stamped installer), and the web/E2E path is thin.
  • AGENTS.md says “never practice TDD here — this overrides… superpowers test-driven-development”. It still demands that a test written after the fix be shown failing without it, so it is test-after plus fail-without-fix, not red-first.
  • No mutation step.
  • The SHA stamp only makes sense for things that build into an artifact.

Gate fit.

  • The worker fills the cw-gates checklist plus the dogfood “not exercised” list in the PR body.
  • The reviewer or Freezy applies “a gate is its artifact”: read the record on the PR at the head SHA, not the check rollup.
  • Ambiguous or missing evidence means INCONCLUSIVE (this matches the PM digest’s §4). Then merge with magpie’s --match-head-commit.

Vs pstack / magpie.

  • pstack create-verification-skill covers the Drive/Evidence half: how to launch and drive this app, and real-user-path standards. Codewhale adds what pstack lacks: fail-without-fix, a not-run list, and a SHA tie-in.
  • Magpie’s record (“Not run: …”, “Reviewed at head…”, --match-head-commit) is the same idea, written from the reviewer’s side. Codewhale is the worker-side and pre-merge version, and is more explicit about N>0 and artifact literalness.

#2 superpowers verification-before-completion (+ alpine review-pr stash recipe, + dsh negative controls)

Section titled “#2 superpowers verification-before-completion (+ alpine review-pr stash recipe, + dsh negative controls)”

obra/superpowers verification-before-completion (MIT, 8ca22db 2026-09-26 02:06 UTC+8), paired with test-driven-development. Reviewer-side counterpart: alpinejs/alpine .claude/skills/review-pr. Negative-control rule: deepseek-harness dsh-ci-test-reliability.

What it does. A completion gate: IDENTIFY the proving command → RUN it fresh → READ the output → VERIFY it supports the claim → ONLY THEN claim. A table maps each claim to the evidence it needs.

“NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE” / “Skip any step = lying, not verifying” / “Regression test works | Red-green cycle verified | Test passes once” / ”✅ Write → Run (pass) → Revert fix → Run (MUST FAIL) → Restore → Run (pass)” / “Agent reports success → Check VCS diff → Verify changes → Report actual state”

alpine: “Actually verify regression. Don’t just reason about whether the test fails without the fix — prove it. Stash the fix (git stash -- <fix files>), rebuild (npm run build), run the test. If it passes without the fix, the test is not testing the fix… This is non-negotiable for bug fix PRs.”

dsh: “For a new static or corpus guard, temporarily introduce the rejected case and observe the intended failure… Report exact commands and observed results; do not describe retries, skipped tests, or pending CI as passing.”

Evidence produced. Fresh command output quoted in the claim, and a revert→fail→restore→pass cycle. Alpine posts a <!-- claude-review --> verdict comment with Test Results, and a human (Caleb) merges.

Pros:

  • The clearest causation recipe, and the most widely installed (superpowers was ~296k★ per the PM digest).
  • Treats sub-agent self-reports as claims to check against the VCS diff.
  • The alpine recipe is the same check done by the reviewer, and is the easiest independent check for a verifier agent.
  • dsh adds negative controls for new guards and lints, plus “report pending checks as pending”.

Cons:

  • A self-check by the same agent: no commit SHA, no “Not run:” section, no artifact format.
  • No E2E or user-path guidance.
  • Alpine says “NEVER run the full test suite” (a repo-specific speed choice) and is a maintainer bot, not a reusable pack.

Gate fit.

  • The worker runs the revert-fail-restore cycle and pastes it.
  • An independent verifier (cheap second cloud agent) repeats alpine’s git stash -- <fix files> → rerun at the reviewed head.
  • Freezy can’t run code, so it only checks that the verifier’s record names the head SHA and shows a real FAIL. Per magpie, a compile error or a skipped test doesn’t count.

Vs pstack / magpie.

  • This is magpie’s “a test that fails without the change… real FAIL” and “Break it on purpose”, packaged as a portable skill.
  • pstack doesn’t do causation at all.
  • deslop already has “deliberate break… suite must go red, then green with the fix” for test changes. This generalises it to every behaviour change.

#3 Every ce-dogfood / ce-test-browser (+ Showboat verify for re-checkable proof)

Section titled “#3 Every ce-dogfood / ce-test-browser (+ Showboat verify for re-checkable proof)”

EveryInc/compound-engineering-plugin ce-dogfood and ce-test-browser (MIT, 67035e9 2026-10-08 07:40 UTC+8). Re-checkable evidence: simonw/showboat (1,231★, Apache-2.0, last push 2026-03-15 UTC+8).

What it does.

  • ce-dogfood is diff-scoped. It maps the affected flows (Mermaid), builds a persona × scenario matrix, and drives each scenario through the agent-browser CLI.
  • It fixes what breaks, with a regression test per fix.
  • It writes a committed report at <root>/dogfood-reports/<YYYY-MM-DD>-<branch-slug>-dogfood.md, which doubles as a resume checkpoint. It never pushes.
  • ce-test-browser is the lighter per-route Pass/Fail/Skip pass.

“Done: every matrix scenario is Pass, Fixed, Skipped, or in a terminal Blocked state; the project’s automated suite has been run once and its result recorded… A green matrix over a red suite finalizes as a not-ready verdict” / “A fix is not done until a regression test fails before it and passes after, or the report says why no automated test was meaningful.” (ce-dogfood)

“every affected route marked Pass, Fail, or Skip and each Skip carrying its reason… dropping a route from the summary because nobody could reach it, is the failure” (ce-test-browser)

Showboat: “A verifier can re-execute all code blocks and confirm the outputs still match.” verify “Re-runs every code block… Prints diffs and exits with code 1 if any output has changed”. Simon: “The exec command… is designed to discourage the agent from cheating and writing what it hoped had happened into the document.”

Evidence produced.

  • The report sections are: Diff Summary, Personas, Flows, Test Matrix (Status / Issue / Fix / Commit), What Was Fixed (regression test fails before / passes after), Console Errors, Human Verifications, Decisions for a Human, and Final Status.
  • Anything needing OAuth, email, payments or SMS is marked Blocked (needs human verify).
  • Showboat adds a markdown proof doc whose exec blocks a reviewer can re-run with showboat verify.

Pros:

  • The best real-user-path discipline found. A Skip has to carry a reason, so nothing quietly drops out, which is the not-run record applied to E2E.
  • Ties the matrix to the suite result and to per-fix commits.
  • Has a “do not hard-code main” rule for the diff base.
  • ce-work already assumes “the host orchestrator inspects actual changes and owns authoritative verification”, which is Noa’s model.
  • Showboat is the only artefact in this survey that a reviewer can mechanically re-run.

Cons:

  • Heavy: a multi-phase reference set, and it needs agent-browser. ce-test-browser says “Do not introduce a third browser stack.”
  • Screenshots go to OS temp, so they’re lost on a cloud VM unless they get attached.
  • The report records commits per fix but not the final reviewed head.
  • Showboat verify only re-runs deterministic shell output. It can’t re-run browser clicks, and it hasn’t been pushed since March.

Gate fit.

  • A cloud worker runs a trimmed ce-test-browser pass (affected routes only) using Cursor’s computer use. Screenshots and video are attached to the PR as Cursor artifacts (capabilities).
  • The PR body carries the matrix with reasons for every Skip.
  • Freezy checks the matrix for unexplained gaps and for any “green matrix over red suite”.
  • A verifier re-runs showboat verify where CLI proof exists. Browser paths still need CI Playwright or a second agent run.

Vs pstack / magpie.

  • pstack’s verify-<app> skill tells the agent how to launch and drive the app. ce-dogfood tells it what to cover for this diff and how to report it. They stack: verify-<app> as the driver, a ce-dogfood-style matrix as the record.
  • Magpie’s ui-preview.yml (DeepSeek plans scenes, Playwright records into the PR body, split-trust jobs) is the CI version of the same evidence, and it is stronger as a gate because the worker doesn’t produce it.

CandidateWhat it addsCausationRun/not-runReal pathIndependent gateFit for cloud worker / PM botPort effortLicense · activity
deepseek-harness dsh-pre-push-checks”narrowest… check that would fail for its regression”; check the selected test count; git rev-parse HEAD origin/<branch> then gh pr checks; “Commit hashes… from before the rewrite are not current evidence”✅✅ pending ≠ pass✗✅ remote-ref = HEADGood worker pre-push stepLow5badb15 2026-10-03; already covered in the lineage digest
mattpocock tdd”Red before green.”; “Expected values must come from an independent source of truth”; “Test only at pre-agreed seams”✅ red-first✗✗✗Pairs with to-tickets seamsLowMIT · b0618bc today
anthropics/skills webapp-testingwith_server.py + Python Playwright; “Reconnaissance-Then-Action”✗✗✅✗Driver onlyLowLICENSE.txt in skill · 683bc88
claude-code pr-review-toolkit pr-test-analyzerRates test gaps 1–10; “good tests are those that fail when behavior changes unexpectedly”Reasoned, not run✗✗Read-only reviewFine for Freezy (no exec), weak as proofLow71cddde today
vercel-labs/agent-browser dogfoodExploratory QA with record start …webm, per-step screenshots, console/errors; “Verify reproducibility before collecting evidence”; “Never read the target app’s source code”✗ (finds bugs)✗✅Blind-tester stanceGood second-agent QALow0207911 today
ChromeDevTools chrome-devtools skillnavigate → wait → take_snapshot (uid) → interact; screenshots, CSS, a11y/LCP/memory siblings✗✗✅✗DriverLow2744afa today
Playwright Test Agentsplanner → specs/*.md, generator “verifies selectors and assertions live”, healer re-runs until green✗✗✅ CI-runnable✅ if specs run in CIOutput is real .spec.ts that CI can gateMediumHealer may skip a test “if the healer believes that functionality is broken”: an oracle-edit risk
Cursor cloud agent computer use (changelog, blog)“start dev servers, open the app in a browser, click through UI flows, and verify their changes work before pushing a PR”; screenshots/videos/logs attached to PR✗✗✅✗ produced by workerNative to Noa’s workersNone”Allow posting artifacts to GitHub” uses public unguessable URLs; CI autofix is GitHub Actions only
Anthropic long-running harnessfeature_list.json "passes": false; “It is unacceptable to remove or edit tests”; init.sh + e2e smoke each sessionPartialFeature list✅ Puppeteer✗Pattern, not a skillLow2025-11-26 post
simonw guides (first-run-the-tests, agentic-manual-testing)“confirm that the tests fail before implementing”; “Never assume that code generated by an LLM works until that code has been executed.”✅✗✅ (Rodney, Showboat)ShowboatPrompt linesTrivialBlog
keyboardsamurai/mutagate”Mutation testing as a hook: blocks coding agents until their tests catch the bugs”✅ mutation✗✗Hook, localWorker-side onlyMediumMIT · 2★ · created 2026-10-06
lkc-studio/claude-skill-mutantsDiff-scoped mutate.py --since main, non-zero exit on survivors, triage reference✅✗✗CI-runnableCould run as a CI jobMediumStars not verified
mikemartincode/mutation-gate, xiaolai/tdd-guardian-for-claudemutmut pre-commit gate; TaskCompleted hook + optional mutation gate✅✗✗HookLocal onlyMediumStars not verified
duoglas/simple-harness-kit.harness/verify-evidence.json “binds Git commit/tree/dirty state… SHA-256 digest”; reports DEGRADED honestly✗✅✗✅ SHA-boundIdea worth stealingLow0★
dsifry/metaswarmOrchestrator verifies itself, never trusts the worker’s self-report (see PM digest)———✅Already in PM digest—33d39f7 2026-06-20; not re-read

Signal from X and HN (context, not packs):


3. Synthesized verification record (template)

Section titled “3. Synthesized verification record (template)”

Sources: magpie’s record + Codewhale cw-gates/dogfood + ce-dogfood matrix + dsh pre-push + PM digest §4. The worker pastes it into the PR body. The reviewer appends its own block at the reviewed head.

## Verification record
Head: <full sha> Base: <base sha / branch> Pushed ref = HEAD: yes (`git rev-parse HEAD origin/<branch>`)
Oracle: <issue AC / test file / spec> — not edited by this PR: yes | no (explain)
### Causation (per behaviour change)
- <change> → test `<path::name>`
- without fix: FAIL — `<salient line>` (cmd: `git stash -- <fix files> && <test cmd>`)
- with fix: PASS — `<salient line>`
- (compile error / skipped / 0 tests selected ≠ FAIL)
- New guard/lint: negative control introduced → observed failure `<line>`
- Mutation (optional): <tool> on diff — <k>/<n> killed; survivors: <list or "none">
### Checks run (exact commands, N>0)
| Command | Result | Salient line |
|---|---|---|
| `<cmd>` | PASS/FAIL | `test result: N passed; M failed` (N=…) |
### Real user path
| Route / scenario | Driver (Playwright / browser / CLI / API) | Status (Pass/Fail/Skip/Blocked) | Evidence (screenshot / video / log link) | Reason if Skip/Blocked |
|---|---|---|---|---|
Build under test reports SHA: <short sha> (stamped build / version endpoint), or "n/a: <why>"
### Not run
- <check> — <why> (e.g. needs secrets, OAuth, payments, slow suite, provider tokens)
### CI at head
<check name>: success/failure/pending @ <sha> (pending is pending, not pass)
---
## Reviewer check (independent, at head <sha>)
- [ ] Record head == PR head now (else stale → INCONCLUSIVE)
- [ ] CI check runs for <sha> all success (list_check_runs_for_ref), none skipped silently
- [ ] Causation re-run by verifier: stash fix → FAIL observed / not re-run (why)
- [ ] Oracle untouched: diff of tests/fixtures/specs reviewed; no deleted/weakened assertions, no new skips
- [ ] Every AC maps to a check or a user-path row; every gap is in "Not run"
- [ ] Artifacts open and match claims (screenshot/video shows the asserted state)
Verdict: PASS / FAIL / INCONCLUSIVE — merge only with `gh pr merge --match-head-commit <sha>`

Rules that go with it:

  • The artifact has to literally say PASS (Codewhale).
  • Missing evidence means INCONCLUSIVE, never PASS (PM digest §4).
  • A Cursor worker’s own screenshots count as claims. Only CI or a separate verifier counts as the gate.

4. Overlap with Noa’s skills and where a mirror would sit

Section titled “4. Overlap with Noa’s skills and where a mirror would sit”
  • Already covered:
    • deslop “Clean with proof” (deliberate break → red → green; hand-back proof: <test/command/result>)
    • repo-sloppiness (evidence cites file:line or a command)
    • project-map (“Do not treat green cards as truth without a test/CI/demo link”)
    • lessons-ledger (“verified too thinly?”; back a promoted rule with a test or lint)
    • critique (implementer AC to tick)
    • project-pm onboarding step 3 (offer pstack /create-verification-skill, only on an explicit yes)
  • Missing: a record format (as the magpie digest noted), a fail-without-fix step outside deslop’s test-touching case, an N>0 check, head-SHA binding, and a reviewer-side gate checklist that Freezy can run without executing code.
  • Where a mirror would sit (proposal only; no skill was created or edited):
    • Worker side: a verification-record skill that cloud-agent briefs reference. It holds the §3 template plus the cw-gates ladder and the superpowers revert-fail-restore cycle, and calls the repo’s pstack verify-<app> skill as its driver when one exists.
    • Reviewer side: a pm-gate section inside project-pm or a sibling. It holds the reviewer checklist from §3, the list_check_runs_for_ref read, the dispatch of a cheap verifier agent for the stash re-run, and the --match-head-commit merge (Noa still merges).
    • Feedback: lessons-ledger already asks “verified too thinly?”. Have it read the “Not run” sections as input.

  • SHA binding is rare. Only magpie’s --match-head-commit, Codewhale’s stamped build and the 0★ simple-harness-kit bind evidence to a commit. No popular skill pack makes the worker record the head SHA.
  • Freezy can’t re-run code, so independence depends on CI or a second agent. Nothing found packages “verifier agent re-runs the stash check at head” as a reusable skill. Alpine’s maintainer bot is the closest.
  • Browser evidence isn’t re-checkable. Cursor artifacts, agent-browser videos and ce-dogfood screenshots are all produced by the worker. Showboat verify only covers deterministic shell output. Magpie’s CI ui-preview.yml is the only split-trust UI evidence seen.
  • Oracle protection is only stated as prose (“It is unacceptable to remove or edit tests”; mattpocock’s “independent source of truth”). No pack checks it mechanically, and the Playwright healer can skip tests.
  • Mutation gates are tiny and new (mutagate 2★, created 2026-10-06; the others unverified). Nothing is mature enough to adopt as-is.
  • Not re-read this run: the OpenAI harness-engineering post (403 to WebFetch), metaswarm and pstack’s own files (relied on prior digests). Reddit wasn’t searched because Exa returned HTTP 417. Kanjun’s “Vet” repo wasn’t verified.