arxiv-2606.32007-axdafny
Extract — AxDafny: Agentic Verified Code Generation in Dafny
Section titled “Extract — AxDafny: Agentic Verified Code Generation in Dafny”- URL: https://arxiv.org/html/2606.32007 (abs: https://arxiv.org/abs/2606.32007)
- Authors: Benjamin Breen, Austin Letson, Borja Requena Pozo, Leopoldo Sarra (Axiomatic AI, Boston)
- Code: https://github.com/Axiomatic-AI/ax-dafny
- Read via: defuddle parse —md
Claims
Section titled “Claims”- AxDafny = verifier-guided agentic loop for joint Dafny program + proof synthesis (impl, invariants, asserts, termination) — not proof-hint-only.
- Architecture adapted from Axiomatic’s minimal Lean ATP agent: proposer → deterministic+LLM review → reflection memory; budget 20 iterations/instance.
- Deterministic gates: (1) proposed
requires/ensuresmust superset original (no spec weakening); (2) regex ban on bypass:{:axiom},{:verify false},assume,{:extern}. Thendafny verify(+ optional--counterexamples). LLM reviewer catches vacuous predicate rewrites. - Memory: previous attempt verbatim + rolling scratchpad regenerated each iter (preserve lessons).
- New bench LCB-Pro-Dafny: 250 LiveCodeBench-Pro problems → Dafny (100 easy / 100 med / 50 hard) with NL + signature + formal spec; curated (LLM-as-judge flagged 14/250 = 5.6% semantic mismatches, all fixed). Dafny 4.11.0.
Key numbers
Section titled “Key numbers”DafnyBench (proof-hint restore)
Section titled “DafnyBench (proof-hint restore)”| Method | Verify % |
|---|---|
| AxDafny (Gemini-3.1-Pro) | 92.7% (725/782) |
| AxDafny (GPT-5.5 medium) | 88.9% (695/782) |
| AxDafny (GPT-5.5 low) | 85.3% |
| DafnyPro (full, w/ hint library) | 86.2% |
| DafnyPro w/o retrieval | 76.2% |
| Gemini-3.1-Pro pass@1 | 68.5% |
| DafnyBench paper baseline | 68.0% |
| GPT-5.5 med/low pass@1 | 54.6% / 54.1% |
+6.5 pp over strongest prior proof-hint baseline; no task-specific hint retrieval. Gains concentrated in first ~5 iterations (Fig. 2). Permissive helper-lemma setting: GPT-5.5 med 92.0%.
LCB-Pro-Dafny (program+proof synthesis)
Section titled “LCB-Pro-Dafny (program+proof synthesis)”| Method | Easy | Med | Hard | Overall |
|---|---|---|---|---|
| AxDafny | 75.0% | 52.0% | 28.0% | 56.4% |
| GPT-5.5 pass@1 | 13.0% | 14.0% | 4.0% | 11.6% |
Easy ablation: GPT-5.5 med 86/100; Gemini 3.1 Pro 77; GPT-5.5 low 75; Opus 4.5 52.
Verify ≠ runtime
Section titled “Verify ≠ runtime”Among verified→Python compiles under LCB-Pro harness: easy 75 verified → 32 pass / 39 TLE / 4 MLE; med 52 verified → 6 pass / 44 TLE / 2 MLE. Specs enforce functional correctness, not asymptotic complexity — agents can verify inefficient code.
Quotes
Section titled “Quotes”“On DafnyBench, AxDafny verifies 725/782 instances (92.7%), outperforming the strongest previously reported proof-hint baseline by 6.5 percentage points.”
“…most executable failures result from resource limits, because the Dafny specifications enforce functional correctness rather than asymptotic complexity.”
“verification success and runtime test performance measure different aspects of generated code.”
DIY / fleet relevance
Section titled “DIY / fleet relevance”- Copyable harness: spec-preservation check + bypass denylist + dafny verify CEx → repair with rolling reflection memory.
- Prefer obligation-local feedback early (first 5 iters deliver most of DafnyBench lift).
- Treat “verified” as functional vs trusted spec, not performance or product correctness.
- Dafny compile-to-Python secondary eval is a useful reality check for agent fleets that only watch the verifier.
- Contrast: AxDafny joint synth vs DafnyBench/AutoVerus-style annotation restore — hard split still open (28%).
Caveats
Section titled “Caveats”- Specs are trusted fixed objects; incomplete/wrong specs → false confidence (Impact Statement).
- LCB-Pro-Dafny curation used model-assisted specs + human review; not pure auto-translate.
- Hard split far from saturated; competition-style asymptotics outside FV gate.