veribench-neurips-2026
Extract — VeriBench: End-to-End Formal Verification Benchmark for AI Coding Agents in Lean 4
Section titled “Extract — VeriBench: End-to-End Formal Verification Benchmark for AI Coding Agents in Lean 4”- URL: https://cs.stanford.edu/people/brando9/professional_documents/papers/NeurIPS_2026_VeriBench.pdf
- Authors: Brando Miranda† et al. (Stanford / Harvard / UC Berkeley / HKBU / Galois / LLNL; † brando9@stanford.edu). Nie‡ work before Google DeepMind.
- Venue label on PDF: NeurIPS 2026 preprint (filename); Preprint banner in body
- Read via: curl PDF + pdftotext (defuddle N/A for PDF)
Claims
Section titled “Claims”- End-to-end Python→Lean 4 autoformalization under agentic verifier feedback — not proof-completion against fixed specs. Agents must emit Lean implementation + tests + theorem statements + proofs from executable Python source.
- 884 tasks total: 602-task canonical core + VeriBench-IndustrySet (282) across 14 industry-elicited high-assurance domains. Headline SCSC table uses 176 Lean evaluation units (multi-formalization of canonical five splits); IndustrySet scored separately at compile-rate (Appendix I).
- Canonical core families: HumanEval 56; EasySet 41; CSSet 13; SecuritySet 460 (MIT 6.858–derived; 227 safe/vulnerable pairs → 454 + extras); RealCodeSet 32 (stdlib).
- Novel score SCSC = geometric mean of five factors: IC1 (typecheck), IC2 (sorry-free proof of stated thms), TE1 (thm↔gold semantic equivalence), D1/D2 (gold validity gates). Zero any factor → score 0 (conjunctive).
- Headline agent-skill is three-factor only: (S_{e,\mathrm{skill}}=(IC_1\cdot IC_2\cdot TE_1)^{1/3}). Five-factor (S_{e5}) mixes agent + gold quality.
- Theorem-equivalence gap: agents typecheck easily but state theorems that don’t match gold obligations. Formulation ≥ as hard as proof search.
- TE1 = Claude Sonnet 4.6 LLM judge (k=3 median 0–10); separate Codex CLI rubric vs humans: Pearson r≈0.61, rank ρ≈0.38 — proxy, not kernel bi-implication.
Key numbers (Table 2, canonical core)
Section titled “Key numbers (Table 2, canonical core)”| System | IC1 | IC2 | TE1 | n | (S_{e,\mathrm{skill}}) | (S_{e5}) |
|---|---|---|---|---|---|---|
| Codex (GPT-5.4) | 1.000 | 0.237 | 0.102 | 165 | 0.289 | 0.417 |
| Claude Code (Sonnet 4.6) | 1.000 | 0.114 | 0.098 | 165 | 0.224 | 0.358 |
| Single-call Sonnet 4.6 | 0.310 | 0.282 | 0.105 | 168 | 0.209 | 0.344 |
| Gemini 3-flash-preview† | 0.206 | 0.043 | 0.156 | 141 | 0.111 | 0.237 |
| Leanstral v2 | 0.292 | 0.065 | 0.058 | 168 | 0.103 | 0.225 |
Gold quality (Q_e^{\mathrm{gold}}\approx 0.72)–0.74 across systems (D1≈0.92, D2≈0.57). Abstract rounds skill to 0.29 / 0.22 / 0.10. Overall TE1 ≤ 0.156; headline systems ≤ 0.105.
Harness: Harbor (Laude/Terminal-Bench) — fresh Docker/task, domain blocking, harbor run -d veribench@1.1; budgets 1h agent / 1h verify / 1h env build.
Quotes
Section titled “Quotes”“Codex and Claude Code typecheck all generated files, yet estimated theorem–gold equivalence remains at ≤ 0.156 across evaluated systems.”
“…surfacing theorem formulation as a bottleneck at least as severe as proof search.”
“SCSC’s theorem-equivalence factor TE1 is computed by a Claude Sonnet 4.6 LLM judge… treating TE1 as a semantic equivalence estimate, not a formal bi-implication proof.”
DIY / fleet relevance
Section titled “DIY / fleet relevance”- Don’t score FV agents on typecheck alone (IC1=1.0 can still mean (S_{e,\mathrm{skill}})≪0.3).
- Gate on stated obligations matching intended property (TE1-like review or human spec review) + sorry-free proofs.
- Prefer conjunctive / all-or-nothing gates over compensatory averages.
- Harbor-style sandbox + gold-mirror blocking is a reusable anti-cheat pattern.
- Contrast vs Vero: VeriBench = source-conditioned autoformalization (Python→Lean); Vero = repo-level code+proof fill against Lean scaffolds. Both show high local skill ≠ full closure.
Caveats
Section titled “Caveats”- TE1 is LLM-judged; Sonnet judging Claude Code is a named confound.
- D2≈0.57 means gold proofs themselves often still have sorry — gold quality is imperfect by design (exposed via D gates).
- IndustrySet not in headline SCSC denominator.
- No stable arXiv abs linked from this PDF path (Map gap still open).