跳转到内容

PM worker reviewer agent orchestration

PM → worker → reviewer agent orchestration

Section titled “PM → worker → reviewer agent orchestration”

Before this: 2026-10-08 Agent skills for splitting work into issues. Its top picks are to-tickets (mattpocock), then slice (AgentiveStack: AFK/HITL plus “more than a day → split”), then Every’s ce-plan fields for per-unit Verification and Test scenarios. Those issue bodies (What to build · Acceptance criteria · Blocked by · AFK/HITL) are the input to everything below.

Related: Cloud agent orchestrator (archived Coordy: “one run ↔ worktree ↔ branch… PR remains the integrate gate”) · 2026-09-28 Persistent project coordinator model · 2026-09-26 Cloud agent orchestrators in the wild · 2026-09-28 Worktree and PR by design orchestrators · 2026-09-28 Pi SDK sessions worktrees and TS coordinator · 2026-10-02 Overnight and long-running agents · 2026-10-02 Dual-model Sol Codex Claude UI advisor prompts · 2026-10-08 yetone magpie agent workflow · 2026-10-08 Agent lessons ledger lineage

Verification: 2026-10-08 Agent verification skills — concrete recipes and a verification-record template for the §4 review gate (Codewhale cw-gates, superpowers revert-fail, ce-dogfood matrix). Design gate: 2026-10-08 Design-first feature gate — pre-implementation design sign-off and post-implementation conformance review for new features.

Pinned 2026-10-08 ~23:30 (UTC+8). Skill and agent files were read through cursor-github get_file_contents or raw fetches at default-branch HEAD. Star counts come from the GitHub API at pin time. Model names in the quotes are as the sources printed them.


#1 subagent-driven-development + task-reviewer-prompt: obra/superpowers

Section titled “#1 subagent-driven-development + task-reviewer-prompt: obra/superpowers”
Repoobra/superpowers: 296,660★, MIT, HEAD 8ca22db
Filesskills/subagent-driven-development/SKILL.md · task-reviewer-prompt.md · implementer-prompt.md · re-review-prompt.md · scripts/{task-brief,review-package,sdd-workspace} · siblings executing-plans, requesting-code-review, writing-plans, using-git-worktrees, verification-before-completion

What it does. A controller session runs a written plan. For each task it sends a fresh implementer subagent, then a task reviewer, then a fix loop. One whole-branch review runs at the end, on the most capable model.

Key instructions (quoted):

  • “Fresh subagent per task + task review (spec + quality) + broad final review = high quality, fast iteration”
  • Model selection: “Use the least powerful model that can handle each role… Mechanical implementation tasks (isolated functions, clear specs, 1-2 files): use a fast, cheap model… Architecture and design tasks: use the most capable available model. The final whole-branch review is one of these.” and “Always specify the model explicitly when dispatching a subagent. An omitted model inherits your session’s model.”
  • “Turn count beats token price. … the cheapest models routinely take 2-3× the turns on multi-step work — costing more overall. Use a mid-tier model as the floor for reviewers and for implementers working from prose descriptions.”
  • Reviewer: “Treat the implementer’s report as unverified claims about the code… a stated rationale never downgrades a finding’s severity.” The spec check looks for Missing / Extra / Misunderstood, and also lists “⚠️ Cannot verify from diff” items.
  • Fix loop: “Five rounds maximum per task… Rounds 1-3 — resume the original implementer… Rounds 4-5 — dispatch a fresh implementer on a more capable model.” “Never fix findings yourself in the controller session.”
  • Stop conditions: “Four things stop you, and only these: an irreversible or destructive operation; a security-sensitive action; a side effect outside this worktree… (a merge, a push to a shared branch, a publish); and a plan so broken that every path forward is a guess.”
  • Ledger: “Track progress in a ledger file, not only in todos.” Every ruling is written as Ruling: <what> — <why> — <what it costs if wrong> and reported back at the end under “Rulings I made”.

Loop shape

SeatWhoModelOutput
Planwriting-plans (upstream)most capableplan + Global Constraints
Implementfresh subagent per taskcheap for complete-code transcription; mid for prose specs; top for designstatus DONE / DONE_WITH_CONCERNS / NEEDS_CONTEXT / BLOCKED, commits, report file
Task reviewread-only subagent, different seatscaled to diff risk; mid-tier floorSpec ✅/❌/⚠️ + Critical/Important/Minor + “Task quality: Approved / Needs fixes”
Bouncesame implementer for R1–3, then a stronger fresh one for R4–5—scoped re-review: each finding ADDRESSED / NOT ADDRESSED
Breakercontroller adjudicates after R5—park with a ruling, or rule on load-bearing findings and carry on
Finalwhole-branch reviewermost capableONE fix wave, one re-review, then finishing-a-development-branch (merge/PR options go to the human)

Pros. It is the most concrete gate prompt found (file:line evidence; severity calibration; “do not pre-judge findings”). It has explicit model tiers and an escalation rule. Its context hygiene suits a long-lived PM: briefs and reports are passed as files, never pasted history (“a real session’s dispatch hit 42k chars of which 99% was pasted history”). It survives compaction because of the ledger. The rulings list is a ready-made audit trail.

Cons. It is built for one session and one worktree: “Never dispatch multiple implementation subagents in parallel (conflicts).” The reviewer is told not to re-run tests (“the implementer’s report carries the test evidence”). That is risky when the implementer is a remote cloud agent, so metaswarm’s independent-validate rule (#2) should be added. Five rounds is generous for paid cloud runs. It also assumes the plan holds the exact values, which the issue bodies must now carry.

Pairs with issue splitting. A to-tickets issue (What to build · AC · Blocked by) maps directly onto the SDD task brief, and the AC become the reviewer’s [GLOBAL_CONSTRAINTS]. slice’s AFK/HITL type decides whether the PM dispatches at all. A HITL issue waits for Noa. ce-plan’s per-unit Verification / Test scenarios fill the “report contract” the implementer must return.

How it maps onto Noa’s setup

SDD piecePM bot (Freezy etc.) + Cursor cloud agents
controllerthe project PM bot. It still writes no code (project-pm anti-job)
task brief filethe GitHub issue body + one dispatch line (“where this fits”) + the report contract
implementer subagentone Cursor cloud agent per issue (POST /v1/agents with model.id, autoCreatePR: true), on its own cursor/... branch
review packagethe PR diff (cursor-github get_pull_request_diff) + the agent’s report/PR body
task reviewera read-only reviewer agent on a different family from the implementer (pstack interrogate for high-risk diffs; one SDD-style reviewer for routine ones)
fix roundPOST /v1/agents/{id}/runs follow-up on the same agent with the findings verbatim (one active run per agent; 409 agent_busy otherwise). A stronger model means a new agent started with prUrl set to the same PR
ledgerper-issue comment or a sidecar file, not the project note (project-pm keeps one note with “Unfolded Ideas” as the only catch-all)
finishingNoa merges. The PM posts the verdict and the reviewed head SHA

It extends project-pm (a new outbound lane after “next”), pstack figure-it-out Phase C (“Pair delegated work with a judge. If a worker games the gate, reset and harden the contract”; verdicts VERIFIED / NOT VERIFIED / INCONCLUSIVE), and interrogate (multi-model reviewer for the gate). lessons-ledger would read the bounce-back history each night.


#2 orchestrated-execution (IMPLEMENT → VALIDATE → ADVERSARIAL REVIEW → COMMIT): dsifry/metaswarm

Section titled “#2 orchestrated-execution (IMPLEMENT → VALIDATE → ADVERSARIAL REVIEW → COMMIT): dsifry/metaswarm”
Repodsifry/metaswarm: 419★, MIT, HEAD 33d39f7 (last push 2026-06-20 UTC+8) · Show HN (“127 PRs to Prod this wknd with 18 AI agents”)
Filesskills/orchestrated-execution/SKILL.md · skills/plan-review-gate/ · skills/design-review-gate/ · rubrics/adversarial-review-rubric.md · skills/external-tools/ (Codex/Gemini delegation) · skills/pr-shepherd/

What it does. An issue (or a spec with DoD items) becomes a BEADS epic. The epic is split into work units, each with its own DoD, file scope, dependencies and human-checkpoint flag. Each unit then runs a gate state machine.

Key instructions (quoted):

  • “Core principle: Trust nothing. Verify everything. Review adversarially.”
  • VALIDATE: “The orchestrator independently runs quality gates. Never trust subagent self-reports… The orchestrator does NOT ask the coding subagent ‘did the tests pass?’ and accept the answer.” The checks are tsc, eslint, the full test suite, coverage thresholds and git diff --name-only against the declared file scope.
  • Reviewer: “Adversarial — your job is to FIND FAILURES, not to approve… Check EACH DoD item. Cite file:line evidence… Any single BLOCKING issue means overall FAIL.”
  • “Fresh reviewer rule: On re-review after FAIL, the orchestrator MUST spawn a new review subagent. Never pass previous findings to the new reviewer.”
  • “Max retries: 3 attempts per gate, then ESCALATE to human with full failure history.” Escalation lists the attempts, phase, error, fix tried, root-cause assessment, options and a recommendation.
  • Human checkpoints are “planned pauses… Do NOT continue past a checkpoint without human response.” Typical triggers: schema changes, security code, the first unit of a new pattern, and anything that needs external credentials.
  • Coder template: “Do NOT modify files outside your file scope… Do NOT self-certify.”

Loop shape. The orchestrator plans. A 3-reviewer plan review gate and a 5-agent design review gate (PM, Architect, Designer, Security, CTO; capped at 3 iterations) run before any code. Coder subagents implement, optionally delegated to Codex/Gemini CLI “for cost savings”. The orchestrator validates by itself. A fresh adversarial reviewer returns PASS/FAIL, and per the README a “writer always reviewed by different model”. The bounce limit is 3, then the human. Commits are made only after PASS and carry Reviewed-by: adversarial-review (PASS). A final cross-unit review runs at the end.

Pros. It has the clearest gate as a state machine (“Quality gates are BLOCKING STATE TRANSITIONS, not advisory”). It names anti-patterns (“Coverage is close enough at 92%, proceeding”). The file-scope check is mechanical. The escalation report is ready to post to Noa. And it starts from an issue, which matches the PM-bot shape.

Cons. The project is small and has been quiet since June 2026, so mirror the rules, not the framework. It is heavy (18 agents, BEADS bd CLI, mandatory TDD/coverage). “Never pass previous findings” conflicts with superpowers’ scoped re-review, which passes the findings list so each can be ticked ADDRESSED. The trade-off is anchoring versus cost. The validation commands assume a TS/vitest stack, so a PM bot needs the repo’s own verify-* skill instead.

Pairs with issue splitting. A metaswarm work unit is the same thing as a to-tickets issue plus three extra fields: file scope, DoD items, and a human-checkpoint flag. Add the scope and checkpoint fields to the split skill’s template. AFK/HITL is close to the checkpoint flag.

Maps onto Noa’s setup. The VALIDATE step is the hard part for a PM bot that does not run code. Two options: (a) a second cheap cloud agent (or the repo’s pstack verify-<repo> skill) runs the checks on the PR head and reports “what was run, at which head”. (b) The PM reads the required CI checks through cursor-github list_check_runs_for_ref and treats red or missing checks as a FAIL. This extends project-pm onboarding step 3, which already asks about a verify-* skill. Without one the gate cannot be independent.


#3 lfg → ce-work mode:return-to-caller → ce-code-review mode:agent: EveryInc/compound-engineering-plugin

Section titled “#3 lfg → ce-work mode:return-to-caller → ce-code-review mode:agent: EveryInc/compound-engineering-plugin”
RepoEveryInc/compound-engineering-plugin: 25,431★, MIT, HEAD 67035e9 (2026-10-08 07:40 UTC+8). Ships plugin manifests for Claude, Codex, Cursor, Grok, Pi, OMP, OpenCode, Kimi, Devin, Cline
Filesskills/lfg/SKILL.md · lfg/references/stage-routing.md · skills/ce-work/SKILL.md · ce-work/references/{execution-engines,execution-strategy,cross-model-execution}.md · skills/ce-code-review/SKILL.md · ce-code-review/references/persona-catalog.md · cross-model-review.md · ce-babysit-pr, ce-compound

What it does. lfg is a hands-off pipeline. Its steps: work source (an implementation-ready ce-plan) → ce-work in return-to-caller mode (implementation and local verification only) → ce-simplify-code → ce-code-review mode:agent plan:<path> → apply eligible fixes → residual handoff → ce-compound (learning) → commit/push/PR → ce-babysit-pr to CI-decided.

Key instructions (quoted):

  • lfg: “A code change… ends in an open pull request… reviewed with the eligible findings applied and the rest recorded… Merging stays with the user unless they granted it for this run.”
  • ce-work intent: “Workers receive bounded units; the host orchestrator inspects actual changes and owns authoritative verification and canonical commits.” Execution strategy: “Cap concurrency at a bounded batch (~3-5 workers)”. Workers must report “the unit’s verification evidence… the evidence fields are not reconstructable from the tree afterward.”
  • Stage routing: “Planning routes to ce-plan as a plan_model:<alias> carrier… Implementation routes to ce-work as an implementation_engine object”, where target is one of codex, claude, grok, cursor, composer, opencode. An unscoped “use codex” binds to “the implementation stage only”.
  • ce-code-review: “Report-only by default; never push… Never push, open PRs, or file tickets in any mode.” Personas are chosen per diff: correctness always, project-standards when a standards file exists, and testing, maintainability, security, performance, api-contract, data-migration, reliability, adversarial (≥50 changed lines or risky surfaces) or previous-comments when relevant. The adversarial lens can run a second time “through a different model (the ‘peer’)… only when its identity record… says independence_verified: true.”
  • Residuals: “A residual at this point is undecided, not accepted debt… The record therefore goes where that reviewer already looks, the PR body,” as a ## Unapplied review findings checkbox list. It files no ticket per nit (“how a run of small nits floods a tracker”).

Loop shape. Planning uses plan_model. Implementation goes to whichever harness or model implementation_engine names (Cursor/Composer included), with the host verifying. Review uses diff-selected personas plus an optional cross-model adversarial peer. The gate is that an actual ce-code-review receipt must exist (“Never substitute a mental self-review”). Bounce-back is a single apply pass, after which residuals are recorded rather than looped. The human approves the merge.

Pros. It is the most production-hardened, it is multi-harness, and it has explicit per-stage model routing grammar. The reviewer roster scales with the diff. “Unapplied review findings” in the PR body is a clean final handoff to Noa. Return-to-caller mode is exactly the “worker returns, caller owns the gates” contract.

Cons. It is very large: many reference files that “must be read, never approximated”. The design assumes one host session that owns the canonical commits, not a remote PM. There is no numeric bounce cap like superpowers’ 5 or metaswarm’s 3; it applies once and records the rest. It would also need porting into Grok Bot skills rather than installing as is.

Pairs with issue splitting. ce-plan is #3 in the splitting digest. The same plan feeds ce-work, and its per-unit Verification and Test scenarios become the reviewer’s check list (ce-code-review … plan:<path>).

Maps onto Noa’s setup. Borrow three things. (1) Write routing carriers into the dispatch, e.g. implementation: cursor model=<cheap>, review: different family. (2) Use the persona-selection table as the PM’s reviewer roster. (3) Put residual findings in the PR body as checkboxes for Noa, not in the project note. This pairs with critique (the UI reviewer persona) and lessons-ledger (Every’s ce-compound plays the same role).


CandidateWhat’s concreteWhy not top 3 / use
Cursor Projects + Cloud Agents API (docs, API)Coordinator “doesn’t write code itself. It plans the work, delegates… and brings the finished work back to you to check.” API: model.id per agent, autoCreatePR, prUrl, follow-up POST /v1/agents/{id}/runs (one active run per agent).It’s the substrate Noa’s PM bots already drive, not a gate recipe. Use it for the dispatch and bounce mechanics.
Cursor Bugbot (docs).cursor/BUGBOT.md rules; “reads .cursor/config/bugbot.yaml… from the PR’s base branch… a PR cannot change how Bugbot reviews itself”; findings default to neutral, so a required check alone “does not block merges on findings”.A good always-on second reviewer. Enable fail-on-unresolved if available.
Claude Code subagents / agent teams (sub-agents, agent teams)model: per subagent (haiku/sonnet/opus/inherit). Team hooks: TaskCompleted “Exit with code 2 to prevent completion and send feedback”, and TeammateIdle. Caution: teammate plan approval is auto-granted: “Claude Code approves the plan in the lead’s session… without the lead reviewing it.”Primitives, not a playbook. The superpowers and CE skills run on top of them.
Claude Code opusplan (model-config)“uses opus during plan mode, then switches to sonnet for execution”.One-session plan/execute split. No reviewer seat.
Claude advisor tool (blog, CC docs)Executor consults a stronger advisor “before committing to an approach, when stuck… or before declaring a task complete”; /advisor opus; max_uses cap.An in-loop advisor, not a post-hoc gate. Evidence in §3.
Anthropic multi-agent research system (post)Orchestrator–worker; lead saves the plan to memory; “Teach the orchestrator how to delegate… objective, an output format, guidance… clear task boundaries”.Backing only. It warns that “most coding tasks involve fewer truly parallelizable tasks than research”.
Anthropic long-running harness (post)Initializer + coding agent; a JSON feature list with "passes": false that agents may only flip; “It is unacceptable to remove or edit tests”; one feature per session.Backing for “AC as a machine-checkable list” and “don’t let workers edit the oracle”.
OpenAI harness engineering + Codex cloud (post)“Request additional specific agent reviews both locally and in the cloud… iterate in a loop until all agent reviewers are satisfied (effectively this is a Ralph Wiggum Loop)”; “Humans may review pull requests, but aren’t required to.”The most radical agent-to-agent gate. Its preconditions are custom lints and structural tests, which Noa’s repos mostly lack.
Ralph loop (ghuntley, snarktank/ralph 21,932★)`while :; do cat PROMPT.mdclaude-code ; done`; “one thing per loop”; tune with “signs”.Single-agent outer loop; no reviewer seat. It fits inside one cloud-agent run.
Factory Missions / droids (Missions, droid exec, droids)droid exec --mission --worker-model <id> --validator-model <id>; custom droids with model: + tools: read-only reviewer example; Missions “runs user-facing QA testing… to validate each feature”.A productized worker/validator split with separate models. Vendor-bound. Factory openly lists “Cost vs. quality tradeoffs” as still being tested.
Amp oracle / subagents / Puck (tools docs, GPT-5 oracle)oracle = second-opinion reasoning model; “Use the oracle to review the last commit’s changes”; “We intentionally do not force the main agent to always use the oracle, due to higher costs”. Puck “launches and coordinates agents” (manual page behind sign-in).An optional reviewer, not a mandatory gate.
oh-my-pi (OMP) (can1357/oh-my-pi 34,638★)Nine model roles (default, smol for cheap fan-out, slow, plan, advisor…); the advisor “reads every turn the main agent takes, injecting notes inline — a quiet aside, a concern, or a hard blocker”; /review spawns parallel reviewer subagents with P0–P3 ranking.Pi-family harness. Its role config is the cleanest local “plan/cheap/review” model map.
Aider architect/editor (post)--architect; Architect describes the change, Editor writes the edits.Two-model edit step, no gate. Evidence in §3.
Cline Plan/Act (docs)“Use different models for Plan and Act”. Example rows: Claude Opus → Claude Sonnet; GLM 4.6 → Grok Code Fast.IDE mode split; no reviewer.
Devin “manage Devins” (blog)Main session coordinates child Devins in VMs (from 2026-09-26 Cloud agent orchestrators in the wild).Not re-read this run. Vendor-bound.
GitHub Copilot cloud agent (docs)Assign an issue → PR → “request a review from you”; “Copilot will not be aware of… further comments that are added to the issue”; bounce via @copilot in PR comments (“batch them by clicking Start a review”).An issue-queue worker. Useful lesson: post corrections on the PR, not the issue.
OpenHands resolver (docs)fix-me label or @openhands-agent → draft PR (or branch on failure) → remove label; MAX_ITERATIONS default 50.Label-triggered worker with no reviewer of its own. The cleanest “issues labeled for agents” pattern.
claude-squad (repo 8,580★) · Conductor (site) · Vibe Kanban (repo 28,292★) · Sculptor (repo 236★)Parallel agents in worktrees/workspaces, with a human reviewing and merging the diffs.Human-as-reviewer workbenches with no automated gate. Local, not cloud.
Terragon (terragon-oss)“Snapshot notice (January 16, 2026)… at the time of shutdown.”Dead.
Sweep (repo)Repo now describes itself as “AI coding assistant for JetBrains”; hosted issue→PR bot shut down (prior digest).Dead as an issue worker.
maddog (dev.to, repo 1★)Routes “by decision type, not difficulty… Verdict before a merge: an Opus reviewer with no edit tools”.Tiny, but its numbers are useful (§3).
doordash-oss/agentic-orchestrator (repo 112★, HN)“Turn a goal into a supervised run.” TUI for long-running agents.Not read beyond the description.

3. Plan big / implement cheap: the evidence

Section titled “3. Plan big / implement cheap: the evidence”
SourceSetupNumberCaveat
Anthropic multi-agent research (2025-06)Opus 4 lead + Sonnet 4 subagents vs single Opus 4+90.2% on internal research eval; multi-agent uses ~15× chat tokensResearch, not coding: “most coding tasks involve fewer truly parallelizable tasks”
Anthropic advisor strategy (2026-04)Sonnet 4.6 executor + Opus 4.6 advisor+2.7 pp SWE-bench Multilingual, −11.9% cost/task vs Sonnet aloneAdvisor is consulted mid-run (~400–700 tokens), not as a post-hoc reviewer. Vendor benchmark
sameHaiku 4.5 + Opus advisor (BrowseComp)19.7% → 41.2%; trails Sonnet solo by 29%, costs 85% lessBrowsing, not coding. Platform docs say the benefit shrinks as the executor nears the advisor
Aider architect/editor (2024-09)o1-preview architect + DeepSeek or o1-mini editor85% (SOTA then) on aider’s editing benchmark; Sonnet/GPT-4o/4o-mini also improved when paired with themselvesOld models; edit-format task, not end-to-end
superpowers SDDCheapest tier only for “complete code to write” transcription”cheapest models routinely take 2-3× the turns on multi-step work — costing more overall”Practitioner rule from observed sessions; no dataset
maddog (2026-10-07)49 advisor-mode vs 30 plain Claude Code sessions57% of output on Haiku/Sonnet vs 31%; output per session +~60%; top-model output flat (~124K) → “more work routed to cheap models, not a smaller bill”; “Haiku fails on tasks that only look mechanical”n=1 developer; output tokens, not quality
@pescatios on X (2026-09-29 UTC+8)Opus 5.5 orchestrator; Haiku fast-worker; Sonnet 5.5 nuanced-worker; Codex as “independent senior engineer… adversarial review""I’m trialing Haiku… If its retries outweigh the savings, Sonnet 5.5 stays there.” “Don’t use an expensive model merely to validate another model’s output. Prefer actual tests”Setup guide, not measured
@mooroobee on XOrchestrator “reason, plan, decide, and review”; Sonnet explore/implement; Opus “deliberately, never by default""better than Opus… with about half the cost”Self-report, non-coder
@drewdil on XPM agent ↔ architect → per-ticket workers → Codex review subagents → PR agent (Sonnet) → always-on Opus reviewer calling Codex”I mostly keep within my weekly limits”Workflow, no numbers
@Road_Kill11 on XSol-high orchestrator + Luna-max subagents per stacked-PR group; parent issue with sub-issues ↔ PRs; human skims and clicks merge”small PRs were pretty easy to review”Anecdote
HN: Opus 5.5 threadOpus plans, Fable subagent reviews the plan; Fable 5.1 advisor, Sonnet 5.5 for mechanical changes12 PRs, CI ~10 → ~4 min, ~$500 token-equivalent; another commenter: “Having Opus do the work, and have fable do reviews… is a good combo”Anecdotes
Factory MissionsSeparate --worker-model / --validator-model”Cost vs. quality tradeoffs: How aggressive should the orchestrator be?… We are testing this.”Admits it is unmeasured

Reading. The public evidence supports a strong planner or advisor plus a mid-tier executor: small quality gains at equal or lower cost. Large gains appear only when the executor is weak and the advisor is consulted during the run. Cheapest-tier implementers pay off only for fully specified, transcription-like work. Several independent sources say they lose on turns and retries otherwise (superpowers, maddog, pescatios). No public head-to-head was found for “cheap cloud implementer + strong post-hoc reviewer gate” on coding, with the rework cost counted. Noa would have to measure it: log the model tier, rounds to PASS, and cost per issue in the dispatch ledger.

Model-per-seat knobs found: Claude Code model: per subagent, opusplan, /advisor. Cursor Cloud Agents model.id per agent. Cline separate Plan/Act models. OMP modelRoles (smol/slow/plan/advisor). Factory --worker-model / --validator-model and droid complexity routing. Every CE plan_model: / implementation_engine. Amp oracle (fixed by mode).


Built from superpowers (S), metaswarm (M), compound-engineering (CE), the Anthropic harness (A), OpenAI harness (O), magpie (Mg, via 2026-10-08 yetone magpie agent workflow), Cursor/Bugbot (C) and pstack (P).

Inputs the gate needs (from the issue and the dispatch)

  • Issue AC as enumerated, independently checkable items. “not ‘code looks good’” (M). Each says what observation would show it false (to-tickets FAQ, via the split digest).
  • File scope / blast radius declared (M); the reviewer flags anything outside it.
  • A worker report contract: status (DONE / DONE_WITH_CONCERNS / NEEDS_CONTEXT / BLOCKED, S), files changed, and a verification record: commands run, at which head SHA, results, red-before-green where behavior changed, and an explicit Not run: line (Mg, CE). “the evidence fields are not reconstructable from the tree afterward” (CE).
  • The worker may not edit the oracle: tests, AC lists, CI config and reviewer rules are out of scope unless the issue says otherwise (A: “It is unacceptable to remove or edit tests”; C: Bugbot reads rules from the base branch).

What the reviewer checks

  • Spec compliance: Missing / Extra / Misunderstood against each AC item, with file:line evidence (S, M).
  • Diff scope: git diff --name-only against the declared scope; no unrelated churn, no lockfile or format sweeps (M, CE “hidden write surfaces”).
  • Independent verification: required CI checks green on the reviewed head SHA, or a separate verifier ran the repo’s verify-* skill. Never accept “tests pass” from the worker alone (M, P figure-it-out “Verify by inspecting the artifact, never a self-report”). Missing evidence means INCONCLUSIVE, not PASS (P).
  • Quality: error handling, tests assert real behavior, no swallowed errors or duplicated logic. Severity: Critical/Important block, Minor is deferred (S).
  • Risk-selected lenses: security, data-migration, API contract, frontend races or UI (critique) only when the diff touches them (CE persona catalog).
  • Reviewer independence: a different model family from the implementer (CE independence_verified, M “writer always reviewed by different model”, pstack interrogate “adversarial signal comes from model diversity”), read-only tools, and no subagents of its own (S).
  • The reviewer does not trust the report’s rationales (“a stated rationale never downgrades a finding’s severity”, S), and the PM never tells it “do not flag X” (S).

Bounce-back

  • A fix round = one follow-up run with the findings verbatim, then one scoped re-review (S). Cursor: POST /v1/agents/{id}/runs.
  • Cap: 3 rounds (M) for paid cloud runs. Round 3 may switch to a fresh agent on a stronger model (S R4–5 rule, compressed). After the cap → escalate to Noa with the attempt table and options (M).
  • On BLOCKED or NEEDS_CONTEXT: change something (context, model, or split the issue). “Never… force the same model to retry without changes” (S).
  • Minor and contestable findings → a PR-body ## Unapplied review findings checklist (CE), not the loop and not the project note.
  • Re-review style is a fork: a scoped re-review with the findings list (S, cheaper) versus a fresh reviewer with no prior findings (M, against anchoring). Suggested default: scoped for routine issues, fresh for risky ones (auth, data, money, migrations).

Final approval

  • Automatic PASS lets the PM mark the issue “ready”; Noa merges (project-pm anti-job; CE “Merging stays with the user”; S’s four stop conditions include “a merge”).
  • Merge only the reviewed head: gh pr merge --match-head-commit <reviewed sha> (Mg). Post the verdict and what ran before the merge (Mg).
  • HITL issues and planned checkpoints (schema, security, first unit of a new pattern, external credentials) always pause for Noa (M).
  • Agent-only merge (O) only after the repo has mechanical guardrails: custom lints, structural tests, required checks. Not a default.

After

  • Ledger line per issue: model tier, rounds, verdict, cost if available, and rulings with “what it costs if wrong” (S).
  • A nightly lessons-ledger pass over bounce-backs, so a recurring finding becomes a rule after 3 days (magpie / 2026-10-08 Agent lessons ledger lineage); CE’s ce-compound is the same idea.

5. Suggested shape for Noa (draft, not created)

Section titled “5. Suggested shape for Noa (draft, not created)”

A PM-side skill (working name pm-dispatch) that sits after pm-fold and the issue-split skill:

  1. Frontier. Read open AFK issues with no open blockers. HITL issues → ask Noa.
  2. Route. Pick a tier per issue: complete spec and 1–2 files → cheap; prose spec or multi-file → mid; design judgment → strong, or run pstack architect/arena first. Reviewer = mid or strong, from a different family. Cap 3–5 concurrent agents (CE) and parallelize only across disjoint scopes.
  3. Dispatch. One cloud agent per issue (autoCreatePR: true). The prompt is the issue body + global constraints + report contract + “do not edit tests/AC/CI; do not spawn reviewers”.
  4. Gate. On finish: PR diff + report → independent checks (CI or verify-*) → read-only reviewer → verdict comment on the PR.
  5. Bounce. Up to 3 follow-up rounds as above, then escalate to Noa.
  6. Hand back. Post “ready for merge @ ” with the residual checklist. Move the project note’s “next/status” only (current-truth), and add no logs to it.
  7. Nightly. Run lessons-ledger over the day’s verdicts.

Noa’s skills this extends: project-pm (new lane; the anti-job “open/merge PRs” stays for merge), pm-fold (fold before split), figure-it-out (judge-paired delegation, VERIFIED/NOT VERIFIED/INCONCLUSIVE), interrogate (high-risk reviewer), arena/architect (one-way-door issues before dispatch), critique (UI lens), lessons-ledger. crowd-discuss is for brainstorming forks with cheap OpenCode models (“not code bakeoffs”), not for the gate.


  • No measured data on cheap cloud implementers plus a strong post-hoc reviewer for coding, rework included. All coding-side cost claims are anecdotes or vendor in-loop advisor benchmarks. Noa’s own ledger would be the first real data.
  • The PM can’t run code. Independent validation depends on CI or a verify-* skill per repo (project-pm onboarding step 3). Repos without one have no real gate.
  • Cursor specifics not tested live: follow-up runs on an agent whose PR is already open, prUrl hand-off to a stronger-model agent, and Bugbot fail-on-unresolved availability. Launching agents is an external action and was not done here.
  • Re-review anchoring vs cost (superpowers scoped vs metaswarm fresh) has no data either way.
  • Agent teams caveat: Claude Code teammate plan approval is auto-granted by the lead. Don’t treat it as a gate.
  • Not re-read this run: Devin managed Devins, Codex cloud docs, Amp Puck (manual page behind sign-in), doordash agentic-orchestrator, superpowers implementer-prompt.md / re-review-prompt.md (only referenced from the SKILL).
  • Reddit: Exa site:reddit.com returned nothing; HN had only comment-level reports. X was reached via the x namespace (search_posts_all) with no rate-limit issues.
  • The skill sketch in §5 is a draft only. No skill or project note was created or edited.

Skill / agent files (cursor-github or raw, HEAD at pin)

Vendor docs

Backing posts

X (UTC+8 dates)

HN