跳转到内容
- 161 HumanEval-derived Lean 4 problems for end-to-end verified code gen. Success = both QED stages pass Lean kernel:
- Spec certification: NL → Lean spec ψ + proof ψ ≅ held-out human ψ* (isomorphism / semantic equivalence).
- Impl certification: Lean impl π matching signature + proof that π satisfies ψ* (ground truth, not the model’s ψ — so impl eval is independent of spec-gen quality).
- Specs are deliberately non-computable
Prop predicates (quantifiers/inductives) so models cannot copy spec syntax into impl and simp a vacuous proof (computable specs leak — shown with Fibonacci / palindrome+sum examples vs GPT-4o).
- Avoids test-case supervision, LLM-generated annotations, and incomplete property sets (FVAPPS critique: proving weak bounds can be satisfied by
return 0).
- Staged pipeline enables fine-grained diagnosis (spec gen / equiv proof / impl / corr proof fail independently). Metric: pass@k-seconds (default k=600s) with retries per stage until compile or timeout.
- Also releases 5-problem hand-curated few-shot prompt set (proofs up to 309 LOC; iso proofs 29–82 LOC).
- End-to-end full solve ≈ 0–1/161 (~0–0.621%) across all evaluated setups — frontier benchmark.
- Few-shot (FS all stages): GPT-4o 0% E2E; o4-mini / Claude-3.7 / DeepSeek-R1 each 0.621% (1/161). Spec compile high (71–87%) but spec proved ~0.6–1.2%; impl compile 61–83%, proved 0.6–5.6% (R1 best on impl prove).
- COPRA on equiv+corr proofs: GPT-4o / Claude-3.7 still 0.621% E2E; Claude+COPRA reaches 8.696% impl-proved (best partial) but iso remains the bottleneck.
- Hybrid GPT-5-mini + Kimina FS: 0% E2E (spec proved 0% despite 90% spec compile).
- Headline intro claim: few-shot solves up to 1/161 end-to-end.
- Anti-leak checklist for DIY Lean benches: non-computable specs; held-out ψ*; prove impl against ψ* not ψ; isomorphism gate before counting “verified”; no test-only supervision as FV substitute.
- Adopt: CLEVER as stress test / regression for Lean agent fleets (expect near-zero E2E until harnesses improve); staged gates mirror propose→check→repair; COPRA-style proof agents help corr proofs more than iso.
- Contrast with WybeCoder Clever-Loom 62%: different task (imperative Loom VCs with SMT+Lean hybrid on translated problems) — do not equate to CLEVER E2E rates when citing.
- Hard cases called out: numerical search termination (poly root); open math (nth prime Fibonacci).
- Html defuddle used v4 content (table numbers). Agentic “Hybrid” row sparse. Confidence high on design rationale and ~1/161 hardness; use as calibration against overclaimed Lean agent solve rates (cf. Vero 27/43 repo full-solves — different scope).