跳转到内容

veribench-neurips-2026

Extract — VeriBench: End-to-End Formal Verification Benchmark for AI Coding Agents in Lean 4

Section titled “Extract — VeriBench: End-to-End Formal Verification Benchmark for AI Coding Agents in Lean 4”
  • End-to-end Python→Lean 4 autoformalization under agentic verifier feedback — not proof-completion against fixed specs. Agents must emit Lean implementation + tests + theorem statements + proofs from executable Python source.
  • 884 tasks total: 602-task canonical core + VeriBench-IndustrySet (282) across 14 industry-elicited high-assurance domains. Headline SCSC table uses 176 Lean evaluation units (multi-formalization of canonical five splits); IndustrySet scored separately at compile-rate (Appendix I).
  • Canonical core families: HumanEval 56; EasySet 41; CSSet 13; SecuritySet 460 (MIT 6.858–derived; 227 safe/vulnerable pairs → 454 + extras); RealCodeSet 32 (stdlib).
  • Novel score SCSC = geometric mean of five factors: IC1 (typecheck), IC2 (sorry-free proof of stated thms), TE1 (thm↔gold semantic equivalence), D1/D2 (gold validity gates). Zero any factor → score 0 (conjunctive).
  • Headline agent-skill is three-factor only: (S_{e,\mathrm{skill}}=(IC_1\cdot IC_2\cdot TE_1)^{1/3}). Five-factor (S_{e5}) mixes agent + gold quality.
  • Theorem-equivalence gap: agents typecheck easily but state theorems that don’t match gold obligations. Formulation ≥ as hard as proof search.
  • TE1 = Claude Sonnet 4.6 LLM judge (k=3 median 0–10); separate Codex CLI rubric vs humans: Pearson r≈0.61, rank ρ≈0.38 — proxy, not kernel bi-implication.
SystemIC1IC2TE1n(S_{e,\mathrm{skill}})(S_{e5})
Codex (GPT-5.4)1.0000.2370.1021650.2890.417
Claude Code (Sonnet 4.6)1.0000.1140.0981650.2240.358
Single-call Sonnet 4.60.3100.2820.1051680.2090.344
Gemini 3-flash-preview†0.2060.0430.1561410.1110.237
Leanstral v20.2920.0650.0581680.1030.225

Gold quality (Q_e^{\mathrm{gold}}\approx 0.72)–0.74 across systems (D1≈0.92, D2≈0.57). Abstract rounds skill to 0.29 / 0.22 / 0.10. Overall TE1 ≤ 0.156; headline systems ≤ 0.105.

Harness: Harbor (Laude/Terminal-Bench) — fresh Docker/task, domain blocking, harbor run -d veribench@1.1; budgets 1h agent / 1h verify / 1h env build.

“Codex and Claude Code typecheck all generated files, yet estimated theorem–gold equivalence remains at ≤ 0.156 across evaluated systems.”

“…surfacing theorem formulation as a bottleneck at least as severe as proof search.”

“SCSC’s theorem-equivalence factor TE1 is computed by a Claude Sonnet 4.6 LLM judge… treating TE1 as a semantic equivalence estimate, not a formal bi-implication proof.”

  • Don’t score FV agents on typecheck alone (IC1=1.0 can still mean (S_{e,\mathrm{skill}})≪0.3).
  • Gate on stated obligations matching intended property (TE1-like review or human spec review) + sorry-free proofs.
  • Prefer conjunctive / all-or-nothing gates over compensatory averages.
  • Harbor-style sandbox + gold-mirror blocking is a reusable anti-cheat pattern.
  • Contrast vs Vero: VeriBench = source-conditioned autoformalization (Python→Lean); Vero = repo-level code+proof fill against Lean scaffolds. Both show high local skill ≠ full closure.
  • TE1 is LLM-judged; Sonnet judging Claude Code is a named confound.
  • D2≈0.57 means gold proofs themselves often still have sorry — gold quality is imperfect by design (exposed via D gates).
  • IndustrySet not in headline SCSC denominator.
  • No stable arXiv abs linked from this PDF path (Map gap still open).