跳转到内容

arxiv-2505.13938-clever

Extract — CLEVER: A Curated Benchmark for Formally Verified Code Generation

Section titled “Extract — CLEVER: A Curated Benchmark for Formally Verified Code Generation”
  • 161 HumanEval-derived Lean 4 problems for end-to-end verified code gen. Success = both QED stages pass Lean kernel:
    1. Spec certification: NL → Lean spec ψ + proof ψ ≅ held-out human ψ* (isomorphism / semantic equivalence).
    2. Impl certification: Lean impl π matching signature + proof that π satisfies ψ* (ground truth, not the model’s ψ — so impl eval is independent of spec-gen quality).
  • Specs are deliberately non-computable Prop predicates (quantifiers/inductives) so models cannot copy spec syntax into impl and simp a vacuous proof (computable specs leak — shown with Fibonacci / palindrome+sum examples vs GPT-4o).
  • Avoids test-case supervision, LLM-generated annotations, and incomplete property sets (FVAPPS critique: proving weak bounds can be satisfied by return 0).
  • Staged pipeline enables fine-grained diagnosis (spec gen / equiv proof / impl / corr proof fail independently). Metric: pass@k-seconds (default k=600s) with retries per stage until compile or timeout.
  • Also releases 5-problem hand-curated few-shot prompt set (proofs up to 309 LOC; iso proofs 29–82 LOC).
  • End-to-end full solve ≈ 0–1/161 (~0–0.621%) across all evaluated setups — frontier benchmark.
  • Few-shot (FS all stages): GPT-4o 0% E2E; o4-mini / Claude-3.7 / DeepSeek-R1 each 0.621% (1/161). Spec compile high (71–87%) but spec proved ~0.6–1.2%; impl compile 61–83%, proved 0.6–5.6% (R1 best on impl prove).
  • COPRA on equiv+corr proofs: GPT-4o / Claude-3.7 still 0.621% E2E; Claude+COPRA reaches 8.696% impl-proved (best partial) but iso remains the bottleneck.
  • Hybrid GPT-5-mini + Kimina FS: 0% E2E (spec proved 0% despite 90% spec compile).
  • Headline intro claim: few-shot solves up to 1/161 end-to-end.
  • Anti-leak checklist for DIY Lean benches: non-computable specs; held-out ψ*; prove impl against ψ* not ψ; isomorphism gate before counting “verified”; no test-only supervision as FV substitute.
  • Adopt: CLEVER as stress test / regression for Lean agent fleets (expect near-zero E2E until harnesses improve); staged gates mirror propose→check→repair; COPRA-style proof agents help corr proofs more than iso.
  • Contrast with WybeCoder Clever-Loom 62%: different task (imperative Loom VCs with SMT+Lean hybrid on translated problems) — do not equate to CLEVER E2E rates when citing.
  • Hard cases called out: numerical search termination (poly root); open math (nth prime Fibonacci).
  • Html defuddle used v4 content (table numbers). Agentic “Hybrid” row sparse. Confidence high on design rationale and ~1/161 hardness; use as calibration against overclaimed Lean agent solve rates (cf. Vero 27/43 repo full-solves — different scope).