Agent FV toolbox and DIY playbook
Agent FV toolbox and DIY playbook
Section titled “Agent FV toolbox and DIY playbook”Related: 2026-09-27 Formal verification in agent-driven development · 2026-09-27 Lean 4 agents for verified software · 2026-09-27 Bend2 for verified parallel agents · Cloud agent orchestrator · 2026-09-26 Herdr for v0 cloud orchestrator
Confidence
Section titled “Confidence”| Claim | Confidence | Note |
|---|---|---|
| Agents propose / solvers verify is the durable pattern | High | Converges across BMC-Agent, Schwarz, AxDafny, AutoVerus, Specula, Cedar Analysis |
| Single stack covers all fleet code | Low | Different trust surfaces (BMC bounds, SMT, ITP, TLC models, authz policy) |
| High bench % ⇒ ready for Herdr production merge | Low–Medium | Spec quality, bounds, model≠runtime, verify≠perf still dominate |
| Cherny-style Lean/TLA+ = full SDK FV | Low as FV | Clarified as model→CEx→fix; RV three-trust chain required |
1. Propose → verify → repair map
Section titled “1. Propose → verify → repair map”NL / issue / AGENTS.md / LAWS │ ▼┌───────────────────┐│ Specifier agent │ contracts, TLA invariants, LAWS, Cedar policies└─────────┬─────────┘ ▼┌───────────────────┐│ Coder / prover │ code + annotations / Lean terms / Bend proofs└─────────┬─────────┘ ▼┌───────────────────┐│ Hard checker (CI) │ outside the LLM└─────────┬─────────┘ pass │ fail ▼┌───────────────────┐│ Repair agent │ diagnostics / CEx / obligation snapshots → patch└───────────────────┘| Tool | Spec role | Checker | Repair signal | Best fit |
|---|---|---|---|---|
| BMC-Agent / AProver | Top-down per-fn DSL → CBMC/Kani assumes/asserts | CBMC (C) / Kani (Rust) | Multi-stage CEx validation + CEGAR-at-spec | C/Rust fleets; bugfinding with realism tiers |
| Schwarz | Front-end contracts (C, Verus) | cvc5 / Bitwuzla / Z3 | Obligation-local SMT snapshots + lemmas + theory policies | When “timeout/unknown” thrash agents |
| AutoVerus | Specs assumed given | Verus | Error-type → specialized agents; Lynette cheat-filter; merge partials | Rust/Verus annotation restore |
| AxDafny | Joint program+proof; requires/ensures superset gate | dafny verify (+ CEx) | Reflection memory ≤20 iters; bypass denylist | Dafny codegen + proof repair |
| Specula | Agents write TLA+ invariants + models | TLC (BFS + sim) | Bidirectional trace-val + MC anti-reward-hack | Concurrent/distributed system bugfinding |
| Cedar / AgentCore | NL→Cedar + schema from MCP tools | Cedar Analysis (symbolic) + Gateway enforce | Analysis rejects vacuous/conflicting policies | Tool-gateway authz envelope |
| WybeCoder | Loom/Velvet specs | cvc5 + Lean | Subgoal decomp / sequential pass@k | Imperative prove-as-you-generate (CC-BY-NC) |
Lean lake / Bend PROOF | Human specs / LAWS | Kernel / Bend checker (+ --safe) | Elaborator / check errors | ITP / parallel-laws gates — see sibling digests |
2. Comparison table (headline numbers)
Section titled “2. Comparison table (headline numbers)”| Project | Stack | Headline result | Caveat |
|---|---|---|---|
| BMC-Agent | LLM + CBMC/Kani | 62 confirmed real bugs; VibeOS 34/145 CEx realistic | Boundedness k=4; weak specs miss silently; evidence tiers required |
| Schwarz | Agent + SMT | 95.2% agentic-verif (452/475); 91.5% SV-COMP ReachSafety subset | Ablation: lemmas dominate (−15 pp without); ~80% failures = missing frontend semantics |
| AutoVerus | Verus multi-agent | 137/150 (~91%) Verus-Bench | Specs given; small single-fn tasks; Schwarz later 69.3% on KVerus-File vs Schwarz 95.5% |
| AxDafny | Dafny agent | 725/782 (92.7%) DafnyBench; LCB-Pro-Dafny 56.4% vs 11.6% pass@1 | Verify ≠ runtime (TLE dominant on Python transpile) |
| Specula | TLA+/TLC agents | 249 bugs / 48 projects; 68 confirmed / 24 fixed; 0 FP (reproduced); median ~3.7h / $57 | Model-checking / bugfinding — not product theorem proving |
| Cedar AgentCore | Cedar + Analysis | Architectural: default-deny tool traffic; NL→policy + symbolic checks | Authz envelope ≠ functional correctness of tools |
| Vero / CLEVER / VeriBench | Lean 4 | See 2026-09-27 Lean 4 agents for verified software | Repo / E2E / autoformalization hardness |
3. Tool deep notes (for wiring)
Section titled “3. Tool deep notes (for wiring)”BMC-Agent / AProver (arXiv:2605.21434 · agentic-prover/aprover)
Section titled “BMC-Agent / AProver (arXiv:2605.21434 · agentic-prover/aprover)”- Commitments: restricted DSL → deterministic
__CPROVER_*/kani::*; compositional assume-guarantee; CEx are not bug reports until 4-stage validate (reachability → callee feasibility → GCC replay → realism audit). - Evidence tiers for reporting:
confirmed_dynamic/confirmed_system_entry/confirmed_bmc— don’t collapse to one “verified” bit. - Defaults: unwind k=4, per-fn timeout 120s; inconclusive rather than false “clean.”
- Adopt: CBMC/Kani gate for C/Rust; compositional harness; realism pipeline; persistent spec store. Skip: treating bounded clean as unbounded FV.
- Turns coarse timeout/unknown into obligation-local repair: program-point snapshots; SMT-local lemmas (must prove before use); theory-aware policies (LIA/NIA vs BitVec vs FP).
- Trust: model owns search; checker owns authority; final clean-gate regenerates all obligations (snapshots ignored).
- Ablation (aggregate 92.7%): no theory policies 88.2%; no SMT-local lemmas 77.7%; both off 73.3% — lemmas dominate.
- Claim: no incorrect Schwarz-accepted verdicts (authors). DIY: expose
O_l = ⟨l, Γ_l, φ_l, τ_l⟩not just “Verus failed.”
AutoVerus (arXiv:2409.13082 · microsoft/verus-proof-synthesis)
Section titled “AutoVerus (arXiv:2409.13082 · microsoft/verus-proof-synthesis)”- Three phases: preliminary invariants → generic refinement → ~10 error-type debug agents; Lynette rejects
assume/ code-spec edits; Houdini prune; merge complementary partials. -
half tasks <30s or ≤3 LLM calls. Adopt error-type routing + cheat filter; skip one-shot prompting.
AxDafny (arXiv HTML 2606.32007 · Axiomatic-AI/ax-dafny)
Section titled “AxDafny (arXiv HTML 2606.32007 · Axiomatic-AI/ax-dafny)”- Proposer → deterministic+LLM review → reflection memory; budget 20 iters.
- Gates: requires/ensures superset (no weakening); regex ban
{:axiom},{:verify false},assume,{:extern}. - Gains concentrated in first ~5 iterations on DafnyBench. Secondary eval: verified→Python often TLE — specs enforce functional correctness, not asymptotics.
Specula (arXiv:2607.25333 · specula-org/Specula)
Section titled “Specula (arXiv:2607.25333 · specula-org/Specula)”- Push-button agentic TLA+: invariants + models + TLC; scenario projection fights state explosion; bidirectional grounding (trace validation ↔ MC) against reward-hack.
- 80.3% of bugs via MC (median CEx length 9); 99.1% safety invariants; agent sensitivity extreme (Haiku 0 / Opus finds).
- Framing: system-level bugfinding on code-grounded specs — aligns with Cherny/RV, not unbounded FV.
Cedar / AgentCore Policy (AWS Security Blog)
Section titled “Cedar / AgentCore Policy (AWS Security Blog)”- Treat LLM as untrusted; controls outside the model at Gateway: default-deny, forbid-wins, analyzable Cedar.
- Neuro-symbolic: NL→Cedar only with schema validation + Cedar Analysis (always-permit/deny, contradictions, conflicts). Partial eval hides always-denied tools from
list tools. - Complements code-level FV: runtime authorization envelope, not business-logic proofs.
4. Herdr-relevant DIY patterns
Section titled “4. Herdr-relevant DIY patterns”Fits Cloud agent orchestrator / 2026-09-26 Herdr for v0 cloud orchestrator: verification is a skill + CI gate, not a vibe.
Minimal fleet recipe
Section titled “Minimal fleet recipe”Herdr / orchestrator job → coding agent edits worktree → verify skill in sandbox: (dafny | verus | lake build | cbmc|kani | tlc | bend PROOF | cedar analyze) → on fail: repair agent ≤ N rounds (prefer obligation-local / error-type routing) → on pass: PR + optional human on spec / law / policy diffOptional: Code Mode sandbox (2026-09-27 Cloudflare Code Mode) to run verifiers as typed tools without stuffing MCP schemas into context.
When to verify vs leave informal
Section titled “When to verify vs leave informal”| Verify / model-check | Leave informal / PBT / review |
|---|---|
| Authz, money, crypto, parsers on untrusted input | UI copy, glue, one-off scripts |
| Concurrency / state machines / protocols | Soft real-time heuristics |
| Public API contracts of agent-generated libraries | Internal helpers with few callers |
| Policy at tool gateway (Cedar-style) | Prompt text itself |
| Invariants in LAWS / AGENTS “must never” | Aesthetics, latency tuning |
Cost / latency knobs
Section titled “Cost / latency knobs”- Bounded checks first (BMC k, TLC state bound, short
laketargets). - Compositional per-fn contracts before whole-program.
- Cache specs/proofs (AProver spec store; Schwarz snapshots with clean-gate).
- Tiered: unit tests → PBT → FV on hot modules → human for spec diffs.
- Prefer local obligations (Schwarz) over dumping whole SMT logs into context.
- Expect unfinished Vero-style runs can cost more than finishes — budget caps on untouchable classes.
Human gates (recommended)
Section titled “Human gates (recommended)”- Spec / law / policy review before trusting green CI (Cherny/RV: model fidelity).
- Threat-model label on findings (BMC-Agent: active vs latent public-API).
- No auto-merge on first proof of a new invariant class.
- Dual-source specs or triangulation when LLM writes contracts (BMC-Agent dual-source; AxDafny LLM reviewer for vacuity).
- Evidence tiers in reports — never one boolean “verified.”
Stack pick for Noa’s fleet (practical)
Section titled “Stack pick for Noa’s fleet (practical)”| Need | First pick | Second / experimental |
|---|---|---|
| Rust contracts | Verus + AutoVerus-style repair (± Schwarz-like local SMT) | Kani BMC-Agent |
| C systems | CBMC via BMC-Agent | Schwarz C frontend |
| Dafny / teaching FV loop | AxDafny / AutoProv CI | — |
| Concurrent/distributed | Specula TLA+/TLC | Cherny-style Lean models for slices |
| Deep proofs / mathlib | Lean lake — 2026-09-27 Lean 4 agents for verified software | WybeCoder (NC license) |
| Parallel + laws greenfield | Bend2 — 2026-09-27 Bend2 for verified parallel agents | — |
| Tool gateway authz | Cedar Analysis + default-deny | Prompt-only (do not) |
| Agent workflow meta | Lean4Agent FormalAgentLib contracts | Lightweight linters only (weak) |
5. Overclaim traps
Section titled “5. Overclaim traps”- Wrong or weak specs — proof of the wrong property; silent miss if permissive (BMC-Agent admits pipeline cannot self-detect).
- Model ≠ code — Lean/TLA model of TS/Python finds real bugs and misses runtime gaps (rv_inc): need faithful translation + trusted semantics + trusted checkers.
- Bounds — BMC unwind / TLC finite instances ≠ unbounded correctness.
- Verify ≠ runtime / perf — AxDafny verified→Python TLE-dominated.
- IC1 / per-spec % flattery — VeriBench IC1=1.0 with skill ≪0.3; Vero 87% specs ≠ 27/43 repos.
- Clever-Loom ≠ CLEVER E2E — 62% vs ≤1/161.
- Buzz “formally verified the SDK” — Cherny clarified model→CEx→fix (clarification).
- Compiler/metatheory bugs — Bend2 checker↔Lean mismatches; source proof ≠ toolchain trust.
- Authorization ≠ functional correctness — Cedar envelopes tools; doesn’t prove business logic.
- LLM-as-judge as FV — Lean4Agent shows weak SWE alignment; FormalJudge / kernel/SMT required for authority.
- LLMExec / conditional soundness sold as absolute “verified agent.”
- Proof debt denial — Taelin: 1 LOC → tens of proof lines; AI changes economics, doesn’t erase it.
6. Bottom line for Researchy / Noa
Section titled “6. Bottom line for Researchy / Noa”Wire fail-closed checkers into Herdr jobs (Dafny / Verus / Lean lake / CBMC|Kani / TLC / Cedar / Bend LAWS); route repair by error type or obligation locality; spend humans on specs, laws, policies, and threat models; use Lean for deep proofs, Bend2 experimentally for parallel+laws greenfields, Specula/Cherny for concurrency bugfinding, Cedar for tool authz — and never confuse a green agent story with product-level formal verification.
Sources
Section titled “Sources”Papers / products
Section titled “Papers / products”- BMC-Agent — arXiv:2605.21434 ·
sources/arxiv-2605.21434-bmc-agent.md - Schwarz — arXiv:2608.30803 ·
sources/arxiv-2608.30803-schwarz.md - AutoVerus — arXiv:2409.13082 ·
sources/arxiv-2409.13082-autoverus.md - AxDafny — arXiv HTML 2606.32007 ·
sources/arxiv-2606.32007-axdafny.md - Specula — arXiv:2607.25333 ·
sources/arxiv-2607.25333-specula.md - Cedar AgentCore — AWS blog ·
sources/aws-agentcore-cedar-policy.md - WybeCoder / CLEVER / Lean4Agent / Vero / VeriBench / Cherny / RV — see sibling Lean digest +
sources/