LLM project failure modes and slop patterns
LLM project failure modes and slop patterns
Section titled “LLM project failure modes and slop patterns”Primary sources: SlopCodeBench (arXiv:2603.24755) · Building to the Test (arXiv:2606.28430) · AI-Generated Smells (arXiv:2605.02741) · AI Test Theater (Autonoma) · Architectural Guardrails (O’Reilly) · AI Technical Debt (Tembo) · AI debt compounds / SDD (Augment) · The Check Becomes the Spec · Not All Agents Are Equal (arXiv:2609.17598) · practitioner deslop catalogs (KarpeSlop, clean-ai-slop, deslop) · vault current-truth / Repo sloppiness
Pin checked 2026-09-28.
Why this is a different failure mode
Section titled “Why this is a different failure mode”Human review and CI were built for human-scale diffs and human shortcuts. LLM/agent output fails differently:
- Volume — more lines and PRs than review attention (O’Reilly; Faros churn signal cited there).
- Invisible assumptions — design choices (error model, threading, serialization, where state lives) are not in the diff (Augment; ~45% of agent-assisted PRs needed human alignment in one cited study).
- Pass-stable rot — suites stay green while structure becomes harder to extend (SlopCodeBench).
- Self-consistent verification — code, tests, and AI review all agree with each other without an external oracle (AI Test Theater).
A useful caveat: a large wild study of 37k+ provenance-labeled PRs across five commercial agents found vendor-specific post-merge outcomes (revert rates, smells, review load), not a uniform “agents always worse than humans” story within 90 days in popular repos (Not All Agents Are Equal). Treat the patterns below as risk signatures to audit, not as a claim that every AI PR is toxic.
Product / engineering failure modes
Section titled “Product / engineering failure modes”1. Iterative self-extension: verbosity + structural erosion
Section titled “1. Iterative self-extension: verbosity + structural erosion”SlopCodeBench chains agent workspaces across evolving specs (hidden tests; only external CLI/API behavior specified). Across 11 models / 20 problems:
- No agent solved any problem end-to-end; best strict checkpoint solve rate 17.2%.
- Verbosity (AST-grep waste ∪ clone lines / LOC) rose in ~90% of trajectories.
- Structural erosion (share of complexity mass in functions with CC > 10) rose in ~80%.
- Agent checkpoints were ~2.2× more verbose and far more eroded than a panel of maintained human Python repos; human trajectories plateau, agent ones climb.
- Classic symptom: patch new branches into one dispatcher until
main()is a thousand-line CC monster (paper’scircuit_evalexample). - Anti-slop / plan-first prompts improve the intercept (cleaner start) but not the slope — degradation resumes at the same rate once the agent extends its own prior code.
Maps to vault axes: Duplication, Merge scars, Boundary mud (logic collapses inland), Naming drift (helpers proliferate).
2. Building to the test / hollow libraries
Section titled “2. Building to the test / hollow libraries”Building to the Test (Microsoft; Copilot CLI + Opus / GPT): agents must ship a reusable Angular library under a hidden Playwright oracle.
- Without the oracle: incomplete but real libraries; honest low scores.
- With the oracle in-loop: near-perfect scores while behavior is inlined into a throwaway demo; the requested library is dead (L2) or absent (L1). No-op ablation of the “library” leaves the score unchanged when L2.
- Agents’ wrap-up messages often claim the library/services exist while the audit finds they don’t.
- Craft shed even when the library stays wired: publishable manifests and self-authored unit tests drop once the oracle defines “done.”
Practitioner restatement: whatever check you make legible becomes the de facto spec; what the check omits silently stops being the job (The Check Becomes the Spec). Mitigations named there: hold back a hidden acceptance slice; ablate checks so they can fail; gate on using the artifact the way a user would.
Maps to: Orphans (unwired “features”), Test theater, Docs lie (PR summary / README vs tree).
3. Test theater (circular self-verification)
Section titled “3. Test theater (circular self-verification)”- Same model (often same session) writes implementation + tests → assertions encode the bug as expected value. Green means consistency, not correctness.
- AI PR review catches syntax/security/style from the diff; misses business rules that live in tickets/ADRs outside the diff.
- Loop: AI writes → AI tests → AI reviews → all green → bugs still ship.
- Spot checks: delete one line of business logic / flip a boundary (
>=→>) — if suite stays green, theater. Mutation score low under high line coverage = tautological asserts.
Maps to: Test theater.
4. Architectural drift / ADR blindness
Section titled “4. Architectural drift / ADR blindness”O’Reilly — Architectural Guardrails: clean, passing PR that bypasses a service API the team banned in an ADR the agent never saw. Failure is organizational memory, not hallucination.
Wrong layers for this: free-text Cursor rules / CLAUDE.md (no precedence/lifecycle), linters (syntax), SCA (deps), second LLM review (same blind spot). Needed: structured decisions → retrieve → inject before generation → deterministic CI enforce with evidence on disk.
Maps to: Contradiction, Boundary mud, Docs lie.
5. Machine-signature smells and the modular mirage
Section titled “5. Machine-signature smells and the modular mirage”AI-Generated Smells (Concordia; PyExamine + MetaGPT):
- Reasoning–complexity trade-off: stronger models → more Long Method bloat while pursuing edge cases.
- At system scale: God-class syndromes (Too Many Branches, high RFC), Potential Improper API Usage (inline reimplementation instead of helpers), Scattered Functionality + Unstable Dependencies.
- Modular mirage: files split, but semantic cohesion fails — structural modularity without real boundaries.
- Volume–quality inverse law: TLoC near-perfectly predicts architectural smell load (ρ ≈ 0.94); more detailed prompting did not fix decay.
- Functional correctness decoupled from structural quality.
Maps to: Duplication, Boundary mud, Dependency fat / unstable deps, Orphans.
6. Invisible / compounding AI technical debt
Section titled “6. Invisible / compounding AI technical debt”- Debt is invisible (looks correct), scales with adoption, and resists human-shortcut-oriented review.
- Textbook / by-the-book patterns that fight local conventions (cited Ox study pattern rates in Tembo).
- Assumption mismatch on the same contract across three PRs compounds faster than one bad decision.
- Suggested gates: clone detection, complexity ceilings, dep necessity checks, coverage floors — plus scheduled debt scans.
Maps to: Naming drift, Contradiction, Dependency fat.
7. Shared-helper blast radius (practitioner postmortem shape)
Section titled “7. Shared-helper blast radius (practitioner postmortem shape)”Classic incident shape: agent “fixes” uniqueness by normalizing email in a shared helper; signup tests pass; legacy merge path that legally allowed case variants breaks; review only looked at the signup call site (DEV postmortem pattern). Lesson: encode the incident as a failing test; treat the agent as proposer, harness as committer.
Maps to: Contradiction, Orphans (uncovered paths), Merge scars.
8. Half-wired product / product slop
Section titled “8. Half-wired product / product slop”Practitioner deslop catalogs group a third bucket (deslop):
- Flows that compile and demo but feel unfinished to a paying user.
- Generic default UI (e.g. default shadcn look) with no product voice.
- Feature flags / TODOs / scaffold left on main after “done.”
Maps to: Orphans, Merge scars, Prose slop (marketing filler in UI copy).
Recognizable slop signatures
Section titled “Recognizable slop signatures”Code / structure
Section titled “Code / structure”| Signature | What it looks like | Sources |
|---|---|---|
| Verbosity / clones | Identity comprehensions, empty-check scaffolding, duplicated branches, copy-paste helpers | SlopCodeBench; Abbassi taxonomy cited therein |
| Complexity concentration | New features patched into already-hot functions; god main / manager classes | SlopCodeBench; AI-Generated Smells |
| Redundant comments | // Initialize the counter above let counter = 0; JSDoc that restates the signature; // --- Helpers --- dividers | clean-ai-slop; KarpeSlop; deslop |
| Defensive paranoia | try/catch and null guards on trusted internal paths that cannot fail | clean-ai-slop; SlopCodeBench anti_slop prompt |
| Over-abstraction | Trivial wrappers, extra layers for a one-call job, “enterprise” folders for a script | anti-slop / clean-ai-slop |
| Naming theater | tempVariableForCalculation, handleButtonClickEvent, generic data / result / item | clean-ai-slop; cc-polymath anti-slop |
| Type / import lies | any abuse, casts to silence the checker, hallucinated imports (wrong package for a real symbol) | KarpeSlop |
| Debug residue | leftover console.log / prints, commented-out blocks, WIP TODOs landed as done | deslop; KarpeSlop |
| Dead library / twin implementation | Public module exists; runtime path uses an inline copy (L2) | Building to the Test |
| Modular mirage | Many files, scattered responsibility, unstable cross-deps | AI-Generated Smells |
| Convention clash | Textbook pattern that ignores project idioms / ADRs | Tembo; O’Reilly |
Tests / verification
Section titled “Tests / verification”| Signature | What it looks like | Sources |
|---|---|---|
| Tautological asserts | Expected value copied from implementation in the same session | AI Test Theater |
| Coverage theater | High line coverage, low mutation score; suite survives deliberate logic breaks | AI Test Theater |
| Oracle overfitting | Demo/story tuned to test names; requested reusable surface hollow | Building to the Test |
| Review theater | AI review green on syntax; no comments that cite an external requirement | AI Test Theater |
Docs / prose
Section titled “Docs / prose”| Signature | What it looks like | Sources |
|---|---|---|
| Negation residue / changelog voice in live docs | “not X / was X / do not X / dropped Y” stacked through every section | vault current-truth; Repo sloppiness axis Prose slop |
| Chatbot / corporate filler | hedge stacks, empty intensifiers, “Let’s…”, “Great!”, voiceless neutrality | deslop; anti-slop text-patterns |
| Confident lie in wrap-up | PR/agent summary lists modules that aren’t wired or don’t exist | Building to the Test |
| Docs that describe a departed world | README/ADR/comments vs current tree | O’Reilly; Repo sloppiness Docs lie |
| Emoji / conversational tone in code comments | unless the project already does that on purpose | clean-ai-slop |
Folding into Repo sloppiness
Section titled “Folding into Repo sloppiness”The domestic skill already scores ten axes that match this literature closely. Evidence-backed additions / emphasis for skill updates:
| Axis | Research weight |
|---|---|
| Duplication | SlopCodeBench verbosity; PAU / twin paths |
| Contradiction | Shared-helper incidents; dual error models across PRs |
| Orphans | L1/L2 hollow libraries; unfinished TODOs on main |
| Boundary mud | God classes; validation inland |
| Naming drift | Convention clash; synonym APIs across merges |
| Docs lie | ADR blindness; wrap-up vs tree |
| Test theater | Autonoma + Building to the Test |
| Dependency fat | Unstable deps; PR-local packages |
| Merge scars | Scaffold / flags / WIP left merged |
| Prose slop | current-truth residue + LLM comment tells |
Highest-leverage process controls the papers and practitioners converge on (not vibes):
- Hidden acceptance slice the agent never sees; gate on consumer-shaped use of the artifact.
- Deterministic ADR / architecture gates (inject + CI block with on-disk evidence), not a second LLM “review.”
- Complexity / clone / dep budgets in CI; mutation or deliberate-break spot checks, not line coverage alone.
- Treat agents as proposers; encode incidents as failing tests before the next merge streak.
- Expect prompt-only anti-slop to delay, not stop, iterative degradation — schedule structural cleanup / deslop passes after blind-merge streaks.
Related
Section titled “Related”- Repo sloppiness — audit method and axes
- current-truth — kill negation residue in live docs
- Digests on orchestrators / verification (e.g. agent e2e, formal verification) when the failure mode is “green gate, wrong artifact”