跳转到内容

Agent FV toolbox and DIY playbook

Related: 2026-09-27 Formal verification in agent-driven development · 2026-09-27 Lean 4 agents for verified software · 2026-09-27 Bend2 for verified parallel agents · Cloud agent orchestrator · 2026-09-26 Herdr for v0 cloud orchestrator


ClaimConfidenceNote
Agents propose / solvers verify is the durable patternHighConverges across BMC-Agent, Schwarz, AxDafny, AutoVerus, Specula, Cedar Analysis
Single stack covers all fleet codeLowDifferent trust surfaces (BMC bounds, SMT, ITP, TLC models, authz policy)
High bench % ⇒ ready for Herdr production mergeLow–MediumSpec quality, bounds, model≠runtime, verify≠perf still dominate
Cherny-style Lean/TLA+ = full SDK FVLow as FVClarified as model→CEx→fix; RV three-trust chain required

NL / issue / AGENTS.md / LAWS
│
▼
┌───────────────────┐
│ Specifier agent │ contracts, TLA invariants, LAWS, Cedar policies
└─────────┬─────────┘
▼
┌───────────────────┐
│ Coder / prover │ code + annotations / Lean terms / Bend proofs
└─────────┬─────────┘
▼
┌───────────────────┐
│ Hard checker (CI) │ outside the LLM
└─────────┬─────────┘
pass │ fail
▼
┌───────────────────┐
│ Repair agent │ diagnostics / CEx / obligation snapshots → patch
└───────────────────┘
ToolSpec roleCheckerRepair signalBest fit
BMC-Agent / AProverTop-down per-fn DSL → CBMC/Kani assumes/assertsCBMC (C) / Kani (Rust)Multi-stage CEx validation + CEGAR-at-specC/Rust fleets; bugfinding with realism tiers
SchwarzFront-end contracts (C, Verus)cvc5 / Bitwuzla / Z3Obligation-local SMT snapshots + lemmas + theory policiesWhen “timeout/unknown” thrash agents
AutoVerusSpecs assumed givenVerusError-type → specialized agents; Lynette cheat-filter; merge partialsRust/Verus annotation restore
AxDafnyJoint program+proof; requires/ensures superset gatedafny verify (+ CEx)Reflection memory ≤20 iters; bypass denylistDafny codegen + proof repair
SpeculaAgents write TLA+ invariants + modelsTLC (BFS + sim)Bidirectional trace-val + MC anti-reward-hackConcurrent/distributed system bugfinding
Cedar / AgentCoreNL→Cedar + schema from MCP toolsCedar Analysis (symbolic) + Gateway enforceAnalysis rejects vacuous/conflicting policiesTool-gateway authz envelope
WybeCoderLoom/Velvet specscvc5 + LeanSubgoal decomp / sequential pass@kImperative prove-as-you-generate (CC-BY-NC)
Lean lake / Bend PROOFHuman specs / LAWSKernel / Bend checker (+ --safe)Elaborator / check errorsITP / parallel-laws gates — see sibling digests

ProjectStackHeadline resultCaveat
BMC-AgentLLM + CBMC/Kani62 confirmed real bugs; VibeOS 34/145 CEx realisticBoundedness k=4; weak specs miss silently; evidence tiers required
SchwarzAgent + SMT95.2% agentic-verif (452/475); 91.5% SV-COMP ReachSafety subsetAblation: lemmas dominate (−15 pp without); ~80% failures = missing frontend semantics
AutoVerusVerus multi-agent137/150 (~91%) Verus-BenchSpecs given; small single-fn tasks; Schwarz later 69.3% on KVerus-File vs Schwarz 95.5%
AxDafnyDafny agent725/782 (92.7%) DafnyBench; LCB-Pro-Dafny 56.4% vs 11.6% pass@1Verify ≠ runtime (TLE dominant on Python transpile)
SpeculaTLA+/TLC agents249 bugs / 48 projects; 68 confirmed / 24 fixed; 0 FP (reproduced); median ~3.7h / $57Model-checking / bugfinding — not product theorem proving
Cedar AgentCoreCedar + AnalysisArchitectural: default-deny tool traffic; NL→policy + symbolic checksAuthz envelope ≠ functional correctness of tools
Vero / CLEVER / VeriBenchLean 4See 2026-09-27 Lean 4 agents for verified softwareRepo / E2E / autoformalization hardness

  • Commitments: restricted DSL → deterministic __CPROVER_* / kani::*; compositional assume-guarantee; CEx are not bug reports until 4-stage validate (reachability → callee feasibility → GCC replay → realism audit).
  • Evidence tiers for reporting: confirmed_dynamic / confirmed_system_entry / confirmed_bmc — don’t collapse to one “verified” bit.
  • Defaults: unwind k=4, per-fn timeout 120s; inconclusive rather than false “clean.”
  • Adopt: CBMC/Kani gate for C/Rust; compositional harness; realism pipeline; persistent spec store. Skip: treating bounded clean as unbounded FV.
  • Turns coarse timeout/unknown into obligation-local repair: program-point snapshots; SMT-local lemmas (must prove before use); theory-aware policies (LIA/NIA vs BitVec vs FP).
  • Trust: model owns search; checker owns authority; final clean-gate regenerates all obligations (snapshots ignored).
  • Ablation (aggregate 92.7%): no theory policies 88.2%; no SMT-local lemmas 77.7%; both off 73.3% — lemmas dominate.
  • Claim: no incorrect Schwarz-accepted verdicts (authors). DIY: expose O_l = ⟨l, Γ_l, φ_l, τ_l⟩ not just “Verus failed.”
  • Three phases: preliminary invariants → generic refinement → ~10 error-type debug agents; Lynette rejects assume / code-spec edits; Houdini prune; merge complementary partials.
  • half tasks <30s or ≤3 LLM calls. Adopt error-type routing + cheat filter; skip one-shot prompting.

  • Proposer → deterministic+LLM review → reflection memory; budget 20 iters.
  • Gates: requires/ensures superset (no weakening); regex ban {:axiom}, {:verify false}, assume, {:extern}.
  • Gains concentrated in first ~5 iterations on DafnyBench. Secondary eval: verified→Python often TLE — specs enforce functional correctness, not asymptotics.
  • Push-button agentic TLA+: invariants + models + TLC; scenario projection fights state explosion; bidirectional grounding (trace validation ↔ MC) against reward-hack.
  • 80.3% of bugs via MC (median CEx length 9); 99.1% safety invariants; agent sensitivity extreme (Haiku 0 / Opus finds).
  • Framing: system-level bugfinding on code-grounded specs — aligns with Cherny/RV, not unbounded FV.
  • Treat LLM as untrusted; controls outside the model at Gateway: default-deny, forbid-wins, analyzable Cedar.
  • Neuro-symbolic: NL→Cedar only with schema validation + Cedar Analysis (always-permit/deny, contradictions, conflicts). Partial eval hides always-denied tools from list tools.
  • Complements code-level FV: runtime authorization envelope, not business-logic proofs.

Fits Cloud agent orchestrator / 2026-09-26 Herdr for v0 cloud orchestrator: verification is a skill + CI gate, not a vibe.

Herdr / orchestrator job
→ coding agent edits worktree
→ verify skill in sandbox:
(dafny | verus | lake build | cbmc|kani | tlc | bend PROOF | cedar analyze)
→ on fail: repair agent ≤ N rounds (prefer obligation-local / error-type routing)
→ on pass: PR + optional human on spec / law / policy diff

Optional: Code Mode sandbox (2026-09-27 Cloudflare Code Mode) to run verifiers as typed tools without stuffing MCP schemas into context.

Verify / model-checkLeave informal / PBT / review
Authz, money, crypto, parsers on untrusted inputUI copy, glue, one-off scripts
Concurrency / state machines / protocolsSoft real-time heuristics
Public API contracts of agent-generated librariesInternal helpers with few callers
Policy at tool gateway (Cedar-style)Prompt text itself
Invariants in LAWS / AGENTS “must never”Aesthetics, latency tuning
  • Bounded checks first (BMC k, TLC state bound, short lake targets).
  • Compositional per-fn contracts before whole-program.
  • Cache specs/proofs (AProver spec store; Schwarz snapshots with clean-gate).
  • Tiered: unit tests → PBT → FV on hot modules → human for spec diffs.
  • Prefer local obligations (Schwarz) over dumping whole SMT logs into context.
  • Expect unfinished Vero-style runs can cost more than finishes — budget caps on untouchable classes.
  1. Spec / law / policy review before trusting green CI (Cherny/RV: model fidelity).
  2. Threat-model label on findings (BMC-Agent: active vs latent public-API).
  3. No auto-merge on first proof of a new invariant class.
  4. Dual-source specs or triangulation when LLM writes contracts (BMC-Agent dual-source; AxDafny LLM reviewer for vacuity).
  5. Evidence tiers in reports — never one boolean “verified.”
NeedFirst pickSecond / experimental
Rust contractsVerus + AutoVerus-style repair (± Schwarz-like local SMT)Kani BMC-Agent
C systemsCBMC via BMC-AgentSchwarz C frontend
Dafny / teaching FV loopAxDafny / AutoProv CI—
Concurrent/distributedSpecula TLA+/TLCCherny-style Lean models for slices
Deep proofs / mathlibLean lake — 2026-09-27 Lean 4 agents for verified softwareWybeCoder (NC license)
Parallel + laws greenfieldBend2 — 2026-09-27 Bend2 for verified parallel agents—
Tool gateway authzCedar Analysis + default-denyPrompt-only (do not)
Agent workflow metaLean4Agent FormalAgentLib contractsLightweight linters only (weak)

  1. Wrong or weak specs — proof of the wrong property; silent miss if permissive (BMC-Agent admits pipeline cannot self-detect).
  2. Model ≠ code — Lean/TLA model of TS/Python finds real bugs and misses runtime gaps (rv_inc): need faithful translation + trusted semantics + trusted checkers.
  3. Bounds — BMC unwind / TLC finite instances ≠ unbounded correctness.
  4. Verify ≠ runtime / perf — AxDafny verified→Python TLE-dominated.
  5. IC1 / per-spec % flattery — VeriBench IC1=1.0 with skill ≪0.3; Vero 87% specs ≠ 27/43 repos.
  6. Clever-Loom ≠ CLEVER E2E — 62% vs ≤1/161.
  7. Buzz “formally verified the SDK” — Cherny clarified model→CEx→fix (clarification).
  8. Compiler/metatheory bugs — Bend2 checker↔Lean mismatches; source proof ≠ toolchain trust.
  9. Authorization ≠ functional correctness — Cedar envelopes tools; doesn’t prove business logic.
  10. LLM-as-judge as FV — Lean4Agent shows weak SWE alignment; FormalJudge / kernel/SMT required for authority.
  11. LLMExec / conditional soundness sold as absolute “verified agent.”
  12. Proof debt denial — Taelin: 1 LOC → tens of proof lines; AI changes economics, doesn’t erase it.

Wire fail-closed checkers into Herdr jobs (Dafny / Verus / Lean lake / CBMC|Kani / TLC / Cedar / Bend LAWS); route repair by error type or obligation locality; spend humans on specs, laws, policies, and threat models; use Lean for deep proofs, Bend2 experimentally for parallel+laws greenfields, Specula/Cherny for concurrency bugfinding, Cedar for tool authz — and never confuse a green agent story with product-level formal verification.


  • BMC-Agent — arXiv:2605.21434 · sources/arxiv-2605.21434-bmc-agent.md
  • Schwarz — arXiv:2608.30803 · sources/arxiv-2608.30803-schwarz.md
  • AutoVerus — arXiv:2409.13082 · sources/arxiv-2409.13082-autoverus.md
  • AxDafny — arXiv HTML 2606.32007 · sources/arxiv-2606.32007-axdafny.md
  • Specula — arXiv:2607.25333 · sources/arxiv-2607.25333-specula.md
  • Cedar AgentCore — AWS blog · sources/aws-agentcore-cedar-policy.md
  • WybeCoder / CLEVER / Lean4Agent / Vero / VeriBench / Cherny / RV — see sibling Lean digest + sources/