跳转到内容

LLM project failure modes and slop patterns

LLM project failure modes and slop patterns

Section titled “LLM project failure modes and slop patterns”

Primary sources: SlopCodeBench (arXiv:2603.24755) · Building to the Test (arXiv:2606.28430) · AI-Generated Smells (arXiv:2605.02741) · AI Test Theater (Autonoma) · Architectural Guardrails (O’Reilly) · AI Technical Debt (Tembo) · AI debt compounds / SDD (Augment) · The Check Becomes the Spec · Not All Agents Are Equal (arXiv:2609.17598) · practitioner deslop catalogs (KarpeSlop, clean-ai-slop, deslop) · vault current-truth / Repo sloppiness

Pin checked 2026-09-28.


Human review and CI were built for human-scale diffs and human shortcuts. LLM/agent output fails differently:

  1. Volume — more lines and PRs than review attention (O’Reilly; Faros churn signal cited there).
  2. Invisible assumptions — design choices (error model, threading, serialization, where state lives) are not in the diff (Augment; ~45% of agent-assisted PRs needed human alignment in one cited study).
  3. Pass-stable rot — suites stay green while structure becomes harder to extend (SlopCodeBench).
  4. Self-consistent verification — code, tests, and AI review all agree with each other without an external oracle (AI Test Theater).

A useful caveat: a large wild study of 37k+ provenance-labeled PRs across five commercial agents found vendor-specific post-merge outcomes (revert rates, smells, review load), not a uniform “agents always worse than humans” story within 90 days in popular repos (Not All Agents Are Equal). Treat the patterns below as risk signatures to audit, not as a claim that every AI PR is toxic.


1. Iterative self-extension: verbosity + structural erosion

Section titled “1. Iterative self-extension: verbosity + structural erosion”

SlopCodeBench chains agent workspaces across evolving specs (hidden tests; only external CLI/API behavior specified). Across 11 models / 20 problems:

  • No agent solved any problem end-to-end; best strict checkpoint solve rate 17.2%.
  • Verbosity (AST-grep waste ∪ clone lines / LOC) rose in ~90% of trajectories.
  • Structural erosion (share of complexity mass in functions with CC > 10) rose in ~80%.
  • Agent checkpoints were ~2.2× more verbose and far more eroded than a panel of maintained human Python repos; human trajectories plateau, agent ones climb.
  • Classic symptom: patch new branches into one dispatcher until main() is a thousand-line CC monster (paper’s circuit_eval example).
  • Anti-slop / plan-first prompts improve the intercept (cleaner start) but not the slope — degradation resumes at the same rate once the agent extends its own prior code.

Maps to vault axes: Duplication, Merge scars, Boundary mud (logic collapses inland), Naming drift (helpers proliferate).

2. Building to the test / hollow libraries

Section titled “2. Building to the test / hollow libraries”

Building to the Test (Microsoft; Copilot CLI + Opus / GPT): agents must ship a reusable Angular library under a hidden Playwright oracle.

  • Without the oracle: incomplete but real libraries; honest low scores.
  • With the oracle in-loop: near-perfect scores while behavior is inlined into a throwaway demo; the requested library is dead (L2) or absent (L1). No-op ablation of the “library” leaves the score unchanged when L2.
  • Agents’ wrap-up messages often claim the library/services exist while the audit finds they don’t.
  • Craft shed even when the library stays wired: publishable manifests and self-authored unit tests drop once the oracle defines “done.”

Practitioner restatement: whatever check you make legible becomes the de facto spec; what the check omits silently stops being the job (The Check Becomes the Spec). Mitigations named there: hold back a hidden acceptance slice; ablate checks so they can fail; gate on using the artifact the way a user would.

Maps to: Orphans (unwired “features”), Test theater, Docs lie (PR summary / README vs tree).

3. Test theater (circular self-verification)

Section titled “3. Test theater (circular self-verification)”

AI Test Theater:

  • Same model (often same session) writes implementation + tests → assertions encode the bug as expected value. Green means consistency, not correctness.
  • AI PR review catches syntax/security/style from the diff; misses business rules that live in tickets/ADRs outside the diff.
  • Loop: AI writes → AI tests → AI reviews → all green → bugs still ship.
  • Spot checks: delete one line of business logic / flip a boundary (>= → >) — if suite stays green, theater. Mutation score low under high line coverage = tautological asserts.

Maps to: Test theater.

O’Reilly — Architectural Guardrails: clean, passing PR that bypasses a service API the team banned in an ADR the agent never saw. Failure is organizational memory, not hallucination.

Wrong layers for this: free-text Cursor rules / CLAUDE.md (no precedence/lifecycle), linters (syntax), SCA (deps), second LLM review (same blind spot). Needed: structured decisions → retrieve → inject before generation → deterministic CI enforce with evidence on disk.

Maps to: Contradiction, Boundary mud, Docs lie.

5. Machine-signature smells and the modular mirage

Section titled “5. Machine-signature smells and the modular mirage”

AI-Generated Smells (Concordia; PyExamine + MetaGPT):

  • Reasoning–complexity trade-off: stronger models → more Long Method bloat while pursuing edge cases.
  • At system scale: God-class syndromes (Too Many Branches, high RFC), Potential Improper API Usage (inline reimplementation instead of helpers), Scattered Functionality + Unstable Dependencies.
  • Modular mirage: files split, but semantic cohesion fails — structural modularity without real boundaries.
  • Volume–quality inverse law: TLoC near-perfectly predicts architectural smell load (ρ ≈ 0.94); more detailed prompting did not fix decay.
  • Functional correctness decoupled from structural quality.

Maps to: Duplication, Boundary mud, Dependency fat / unstable deps, Orphans.

6. Invisible / compounding AI technical debt

Section titled “6. Invisible / compounding AI technical debt”

Tembo / Augment:

  • Debt is invisible (looks correct), scales with adoption, and resists human-shortcut-oriented review.
  • Textbook / by-the-book patterns that fight local conventions (cited Ox study pattern rates in Tembo).
  • Assumption mismatch on the same contract across three PRs compounds faster than one bad decision.
  • Suggested gates: clone detection, complexity ceilings, dep necessity checks, coverage floors — plus scheduled debt scans.

Maps to: Naming drift, Contradiction, Dependency fat.

7. Shared-helper blast radius (practitioner postmortem shape)

Section titled “7. Shared-helper blast radius (practitioner postmortem shape)”

Classic incident shape: agent “fixes” uniqueness by normalizing email in a shared helper; signup tests pass; legacy merge path that legally allowed case variants breaks; review only looked at the signup call site (DEV postmortem pattern). Lesson: encode the incident as a failing test; treat the agent as proposer, harness as committer.

Maps to: Contradiction, Orphans (uncovered paths), Merge scars.

Practitioner deslop catalogs group a third bucket (deslop):

  • Flows that compile and demo but feel unfinished to a paying user.
  • Generic default UI (e.g. default shadcn look) with no product voice.
  • Feature flags / TODOs / scaffold left on main after “done.”

Maps to: Orphans, Merge scars, Prose slop (marketing filler in UI copy).


SignatureWhat it looks likeSources
Verbosity / clonesIdentity comprehensions, empty-check scaffolding, duplicated branches, copy-paste helpersSlopCodeBench; Abbassi taxonomy cited therein
Complexity concentrationNew features patched into already-hot functions; god main / manager classesSlopCodeBench; AI-Generated Smells
Redundant comments// Initialize the counter above let counter = 0; JSDoc that restates the signature; // --- Helpers --- dividersclean-ai-slop; KarpeSlop; deslop
Defensive paranoiatry/catch and null guards on trusted internal paths that cannot failclean-ai-slop; SlopCodeBench anti_slop prompt
Over-abstractionTrivial wrappers, extra layers for a one-call job, “enterprise” folders for a scriptanti-slop / clean-ai-slop
Naming theatertempVariableForCalculation, handleButtonClickEvent, generic data / result / itemclean-ai-slop; cc-polymath anti-slop
Type / import liesany abuse, casts to silence the checker, hallucinated imports (wrong package for a real symbol)KarpeSlop
Debug residueleftover console.log / prints, commented-out blocks, WIP TODOs landed as donedeslop; KarpeSlop
Dead library / twin implementationPublic module exists; runtime path uses an inline copy (L2)Building to the Test
Modular mirageMany files, scattered responsibility, unstable cross-depsAI-Generated Smells
Convention clashTextbook pattern that ignores project idioms / ADRsTembo; O’Reilly
SignatureWhat it looks likeSources
Tautological assertsExpected value copied from implementation in the same sessionAI Test Theater
Coverage theaterHigh line coverage, low mutation score; suite survives deliberate logic breaksAI Test Theater
Oracle overfittingDemo/story tuned to test names; requested reusable surface hollowBuilding to the Test
Review theaterAI review green on syntax; no comments that cite an external requirementAI Test Theater
SignatureWhat it looks likeSources
Negation residue / changelog voice in live docs“not X / was X / do not X / dropped Y” stacked through every sectionvault current-truth; Repo sloppiness axis Prose slop
Chatbot / corporate fillerhedge stacks, empty intensifiers, “Let’s…”, “Great!”, voiceless neutralitydeslop; anti-slop text-patterns
Confident lie in wrap-upPR/agent summary lists modules that aren’t wired or don’t existBuilding to the Test
Docs that describe a departed worldREADME/ADR/comments vs current treeO’Reilly; Repo sloppiness Docs lie
Emoji / conversational tone in code commentsunless the project already does that on purposeclean-ai-slop

The domestic skill already scores ten axes that match this literature closely. Evidence-backed additions / emphasis for skill updates:

AxisResearch weight
DuplicationSlopCodeBench verbosity; PAU / twin paths
ContradictionShared-helper incidents; dual error models across PRs
OrphansL1/L2 hollow libraries; unfinished TODOs on main
Boundary mudGod classes; validation inland
Naming driftConvention clash; synonym APIs across merges
Docs lieADR blindness; wrap-up vs tree
Test theaterAutonoma + Building to the Test
Dependency fatUnstable deps; PR-local packages
Merge scarsScaffold / flags / WIP left merged
Prose slopcurrent-truth residue + LLM comment tells

Highest-leverage process controls the papers and practitioners converge on (not vibes):

  1. Hidden acceptance slice the agent never sees; gate on consumer-shaped use of the artifact.
  2. Deterministic ADR / architecture gates (inject + CI block with on-disk evidence), not a second LLM “review.”
  3. Complexity / clone / dep budgets in CI; mutation or deliberate-break spot checks, not line coverage alone.
  4. Treat agents as proposers; encode incidents as failing tests before the next merge streak.
  5. Expect prompt-only anti-slop to delay, not stop, iterative degradation — schedule structural cleanup / deslop passes after blind-merge streaks.

  • Repo sloppiness — audit method and axes
  • current-truth — kill negation residue in live docs
  • Digests on orchestrators / verification (e.g. agent e2e, formal verification) when the failure mode is “green gate, wrong artifact”