跳转到内容

arxiv-2606.32007-axdafny

Extract — AxDafny: Agentic Verified Code Generation in Dafny

Section titled “Extract — AxDafny: Agentic Verified Code Generation in Dafny”
  • AxDafny = verifier-guided agentic loop for joint Dafny program + proof synthesis (impl, invariants, asserts, termination) — not proof-hint-only.
  • Architecture adapted from Axiomatic’s minimal Lean ATP agent: proposer → deterministic+LLM review → reflection memory; budget 20 iterations/instance.
  • Deterministic gates: (1) proposed requires/ensures must superset original (no spec weakening); (2) regex ban on bypass: {:axiom}, {:verify false}, assume, {:extern}. Then dafny verify (+ optional --counterexamples). LLM reviewer catches vacuous predicate rewrites.
  • Memory: previous attempt verbatim + rolling scratchpad regenerated each iter (preserve lessons).
  • New bench LCB-Pro-Dafny: 250 LiveCodeBench-Pro problems → Dafny (100 easy / 100 med / 50 hard) with NL + signature + formal spec; curated (LLM-as-judge flagged 14/250 = 5.6% semantic mismatches, all fixed). Dafny 4.11.0.
MethodVerify %
AxDafny (Gemini-3.1-Pro)92.7% (725/782)
AxDafny (GPT-5.5 medium)88.9% (695/782)
AxDafny (GPT-5.5 low)85.3%
DafnyPro (full, w/ hint library)86.2%
DafnyPro w/o retrieval76.2%
Gemini-3.1-Pro pass@168.5%
DafnyBench paper baseline68.0%
GPT-5.5 med/low pass@154.6% / 54.1%

+6.5 pp over strongest prior proof-hint baseline; no task-specific hint retrieval. Gains concentrated in first ~5 iterations (Fig. 2). Permissive helper-lemma setting: GPT-5.5 med 92.0%.

MethodEasyMedHardOverall
AxDafny75.0%52.0%28.0%56.4%
GPT-5.5 pass@113.0%14.0%4.0%11.6%

Easy ablation: GPT-5.5 med 86/100; Gemini 3.1 Pro 77; GPT-5.5 low 75; Opus 4.5 52.

Among verified→Python compiles under LCB-Pro harness: easy 75 verified → 32 pass / 39 TLE / 4 MLE; med 52 verified → 6 pass / 44 TLE / 2 MLE. Specs enforce functional correctness, not asymptotic complexity — agents can verify inefficient code.

“On DafnyBench, AxDafny verifies 725/782 instances (92.7%), outperforming the strongest previously reported proof-hint baseline by 6.5 percentage points.”

“…most executable failures result from resource limits, because the Dafny specifications enforce functional correctness rather than asymptotic complexity.”

“verification success and runtime test performance measure different aspects of generated code.”

  • Copyable harness: spec-preservation check + bypass denylist + dafny verify CEx → repair with rolling reflection memory.
  • Prefer obligation-local feedback early (first 5 iters deliver most of DafnyBench lift).
  • Treat “verified” as functional vs trusted spec, not performance or product correctness.
  • Dafny compile-to-Python secondary eval is a useful reality check for agent fleets that only watch the verifier.
  • Contrast: AxDafny joint synth vs DafnyBench/AutoVerus-style annotation restore — hard split still open (28%).
  • Specs are trusted fixed objects; incomplete/wrong specs → false confidence (Impact Statement).
  • LCB-Pro-Dafny curation used model-assisted specs + human review; not pure auto-translate.
  • Hard split far from saturated; competition-style asymptotics outside FV gate.