跳转到内容

Agent lessons ledger lineage

Companion to 2026-10-08 yetone magpie agent workflow. Related: 2026-09-28 DeepSeek harness dsh system prompt. It also feeds my lessons-ledger skill, which already cites magpie as its source pattern.

Times are UTC+8 unless marked. Influence means there’s an explicit citation. Parallel means similar idea, no evidence of influence.

No named origin. magpie’s LESSONS.md and docs/code-standards.md cite no outside inspiration. The only explicit outside influence in the repo is contributor Junjie Zhou’s issue #883, which names DeepSeek Harness (dsh). It shaped the subsystem docs, the router and the “report checks not run” piece, not the ledger. Each piece has well-dated earlier parallels. The closest single cousin is Anthropic’s oncall-kit (Aug 2026). It keeps a lessons.md and moves a pattern into the reference docs after 3 entries.

2. When the magpie loop started, and does it name sources?

Section titled “2. When the magpie loop started, and does it name sources?”

In-repo sequence (commit times converted from UTC):

When (UTC+8)WhatSource
10-02 19:36AGENTS.md gets a map of moved built-ins, backed by a test (TestMovedBuiltinsSayTheirPlugin)commit fa9daa6
10-05 15:28Junjie Zhou opens #883 (subsystem design docs + PR semantic diff for agent work). He explicitly proposes checking how DeepSeek Harness (dsh) organizes architecture docs, agent context and dev flow. yetone keeps it as a discussion and invites a pilot PR.#883
10-05 19:10PR #889 merged: AGENTS.md “Subsystem design and review” pointer + docs/subsystems/README.md. It includes “Keep each fact in one reference and link to it from agent instructions” and “Report checks that were not run.” The PR body itself says “Go tests were not run for this documentation-only change”. yetone’s merge comment records what he ran and on which head.PR #889, 390db18
10-06 01:37docs/code-standards.md created: “code standards distilled from PR reviews… (Junjie Zhou on Discord)”. The first tables are Junjie’s six review rules. The rest come from PRs #901–#930.926bcaa
10-06 02:17Adds acceptance criteria, gh pr merge --match-head-commit, and review patterns “distilled from recent PR reviews (#706–#931)“a5df344
10-06 03:10LESSONS.md born: “keeps what the nightly review of merged work finds, and agents read it before changing code”. The first review covered 10-05: 220 commits, ~170 releases, ~40 fixes of earlier releases. Three lessons that recurred dozens of times went straight to the standards. Carries a Co-Authored-By: Claude Opus 5.5 trailer.303819c
10-06 03:49Follow-up edit4b03eda
10-07 00:23Second nightly review18d973b
10-07 23:46First threshold promotion: four lessons seen a third day moved to the standards8731dd9

Named sources:

  • LESSONS.md describes itself only in-house (“Each night the day’s commits are reviewed…”).
  • code-standards.md opens “These rules come from what PR reviews keep asking for.”
  • Grepping both files for inspir/compound/karpathy/postmortem/ADR/ralph found nothing.
  • The only credited humans are Junjie Zhou (Discord review rules) and the PR reviewers.

a) Lessons ledger: an agent-maintained file of mistakes, read before work

Section titled “a) Lessons ledger: an agent-maintained file of mistakes, read before work”

All parallel:

  • 2023-08: ExpeL extracts “insights” from experience and edits them with ADD/EDIT/UPVOTE/DOWNVOTE. arXiv 2308.10144
  • 2025-02-06: Cline Memory Bank, structured markdown the agent re-reads every session. blog
  • 2025-05-11: Karpathy, “system prompt learning”. X
  • 2025-06-04 / 07-02: Cursor Memories, beta in 1.0 then GA in 1.2 “with user approvals for background-generated memories”. 1.0 forum, 1.2 forum
  • 2025-07-14: Geoffrey Huntley, Ralph: “When you learn something new… update @AGENT.md using a subagent but keep it brief”. ghuntley.com/ralph
  • 2025-08-18: Kieran Klaassen / Every coins “compounding engineering”: “every bug becomes a permanent lesson, and every code review updates the defaults.” every.to. Formalized 2025-12-11 as plan/work/review/compound, with lessons in docs/solutions/. every.to, plugin
  • 2026-01-02/03: Boris Cherny: “Anytime we see Claude do something incorrectly we add it to the CLAUDE.md”. He calls the PR-tagging flow “our version of @danshipper’s Compounding Engineering”. post, thread
  • 2026-02-05: Mitchell Hashimoto, harness engineering: “anytime you find an agent makes a mistake, you take the time to engineer a solution such that the agent never makes that mistake again”. mitchellh.com
  • 2026-08-18: Anthropic oncall-kit Lessons.md, “a running log of every incident… Claude appends to it on its own”. blog, template
  • 2026-09: the viral “tasks/lessons.md self-improvement loop” CLAUDE.md template on X, e.g. 2026-09-29. Original author not established.

b) Evidence on each lesson (SHAs/PRs, “this is evidence, not reflection”)

Section titled “b) Evidence on each lesson (SHAs/PRs, “this is evidence, not reflection”)”

All parallel:

  • 2025-12-01: Lance Martin’s Claude Diary. /diary entries plus /reflect, which tracks processed entries in processed.log. post
  • 2026-01-15: GitHub Copilot Memory. Each memory carries citations to code locations, which are checked against the current branch before use. blog, docs
  • 2026-07-04: Lilian Weng surveys evidence-driven harness edits (AHE: failure evidence, root cause, fix, predicted impact per change). post
  • 2026-08: oncall-kit entries cite incidents, and the setup provenance reads (seen 3×: INC-…) / (seen 1×, unverified). repo
  • 2026-09-08: oriko9/tribunal docs/LESSONS.md, “This is evidence, not reflection”. 2026-09-03→19: drewlr/AI-spec-kit lessons.md, “each entry keeps the incident that produced it”.

c) Promote after N, from ledger into standards

Section titled “c) Promote after N, from ledger into standards”

All parallel. None counts separate days the way magpie does.

  • 2023-08: ExpeL importance counts. A new insight starts at 2 and is dropped at 0. arXiv
  • 2025-10: ACE, bullets with helpful/harmful counters plus a Curator that grows and refines them. arXiv 2510.04618
  • 2025-12-01: Claude Diary /reflect: “2+ occurrences = pattern, 3+ = strong pattern”, then patterns go into CLAUDE.md. post
  • 2026-03 (reported): Claude Code AutoDream background memory consolidation, gated on ≥24h and ≥5 sessions (orient/gather/consolidate/prune). Secondary sources only: levelup 2026-03-26, wmedia 2026-04-05
  • 2026-08 (closest): oncall-kit Graduation. “When a tag accumulates 3 entries with the same mechanism, or an entry stops being a story and becomes a procedure… promote it into that class’s reference file by PR… replace the entries with one pointer line. Record the promotion itself as an entry.” The blog (2026-08-18) says “If the same pattern shows up enough times, we promote it into the investigation skill itself.” repo, blog. The repo was created 08-04 06:18 UTC+8.
  • Human-practice analogues (postmortem → runbook, ADR status lifecycle): dsh’s Agent Notes uses proposed/implemented/rejected/archived folders and covers “a gap a postmortem surfaced”. Parallel for the ledger. dsh influenced magpie’s subsystem docs (see e), but no lessons file exists in dsh’s tree.

d) Verification record (“what was run / not run”; real FAIL before the fix; merge at the reviewed head)

Section titled “d) Verification record (“what was run / not run”; real FAIL before the fix; merge at the reviewed head)”
  • Influence (partial): magpie’s “Report checks that were not run” first appears in the dsh-inspired PR #889 (10-05). dsh’s AGENTS.md says “report only commands run”. Its dsh-pre-push-checks asks for “the narrowest available test… that would fail for its regression”. #883 explicitly names dsh as a model, so this is the one piece with a documented line. The exact “not run” wording is magpie’s own.
  • Parallel:
    • Anthropic, Effective harnesses for long-running agents (2025-11-26): verify end-to-end before marking a feature done; progress file across sessions. yetone linked this post on 2026-10-03 17:01 as the start of harness engineering (X). That shows he read it, but he never ties it to LESSONS.md.
    • oncall-kit’s “Could not verify: … the exact one-line command the owner should run” (2026-08).
    • riusprime/zero-depth docs/LESSONS.md, first committed 10-07 03:33 UTC+8 (~a day after magpie’s), with “NOT YET RUN is a legal status”. No link to magpie.
  • --match-head-commit gating appears to be magpie’s own addition (10-06 02:17). I found no earlier agent-workflow source for it.

e) Thin AGENTS.md router (“before X, read Y”; CLAUDE.md = @AGENTS.md @LESSONS.md)

Section titled “e) Thin AGENTS.md router (“before X, read Y”; CLAUDE.md = @AGENTS.md @LESSONS.md)”
  • Influence: dsh via #883/#889 (10-05). dsh’s AGENTS.md routes to docs/architecture.md / docs/AGENTS.md, its CLAUDE.md just points at AGENTS.md, and it has docs/subsystems/. magpie adopted docs/subsystems/ and the pointer style the same day.
  • Parallel / background:
    • AGENTS.md as a cross-tool standard: agentsmd/agents.md repo created 2025-08-20 01:22 UTC+8.
    • Huntley’s later how-to-ralph-wiggum: “Keep @AGENTS.md operational only… A bloated AGENTS.md pollutes every future loop’s context.”
    • OpenAI Harness engineering (2026-02-11): “give Codex a map, not a 1,000-page instruction manual”, AGENTS.md as table of contents (~100 lines), docs/ as system of record, a recurring doc-gardening agent. A Chinese X recap yetone replied to credits this post with starting harness engineering (X 10-03).
    • Anthropic best practices (undated): “Treat CLAUDE.md like code… prune it regularly”, plus @imports.

Nobody found.

  • X searches turned up only general magpie promo posts, e.g. 1, 2. Queries tried: “LESSONS.md magpie”, “magpie yetone lessons/nightly/code standards”, and “match-head-commit / reviewed-through / code-standards.md”.
  • HN Algolia had 0 hits for “LESSONS.md” and “yetone magpie”.
  • Exa found only the file itself and mirrors. A PyShine source tour (2026-09-30) and AICoder (2026-10-03) cover the product and predate LESSONS.md.
  • GitHub code search for magpie’s markers (“separate days” in LESSONS.md, reviewed-through) found no copies.
  • The only downstream artifact is my own lessons-ledger skill. The loop is two days old, so silence is expected.
  • 2026-09-13 01:23: what he learned building avante.nvim (“a year before claude code”, in the weak-model era) went into Alma. That’s about harness design, not a lessons file.
  • 2026-09-13 01:40: you can tell how long someone has been building harnesses from their Edit tool.
  • 2026-10-03: says he builds magpie with Claude Code. The founding LESSONS.md commit carries a Claude co-author trailer.
  • 2026-10-03 17:01: links Anthropic’s long-running harness post as where “harness engineering” took off “last year”.
  • His other repos have no lessons/standards files: avante.nvim has no AGENTS.md/CLAUDE.md/LESSONS.md; cumora (2026-08) has CONTRIBUTING/docs only.
  • His skill repos native-feel-skill (2026-05) and kill-ai-slop (2026-07) show a habit of turning research into agent-readable rule files, which is a related instinct.
  • from:yetone searches for lessons/教训 (lessons)/复盘 (retrospective)/nightly/AGENTS.md/CLAUDE.md found nothing earlier about this loop. No blog or talk found.

6. Close cousins that do parts better, and what to borrow for lessons-ledger

Section titled “6. Close cousins that do parts better, and what to borrow for lessons-ledger”
Cousin (date)Does betterBorrow into lessons-ledger
Anthropic oncall-kit (2026-08-04 repo / 08-18 blog)Graduation also fires when an entry “becomes a procedure”, not only on count. Promotion happens by PR, leaves a pointer line, and is itself logged. A status banner serves as a resume point. “Could not verify” entries name the owner’s command. Humans prune.Promote-by-PR with pointer line. Log each promotion as an entry. Add a could-not-verify entry type. Add a “procedure-shaped” promotion trigger next to ≥3 days.
GitHub Copilot Memory (2026-01-15)Citations are re-validated against the current branch before a memory is used. Unused memories expire (28 days).Before reading or bumping a lesson, check that its evidence SHAs/paths still exist. Use auto-expiry instead of my soft “~30 days” prune.
ACE (2025-10) / ExpeL (2023-08)Counters run both ways (helpful/harmful, upvote/downvote to zero)Track a “done right on” counter alongside Seen. A lesson repeatedly done right can retire.
Claude Diary (2025-12-01)processed.log stops double counting. Rules that keep being violated get strengthened.Keep a processed log next to the watermark. Escalate wording or add a test when a promoted standard is violated again.
AutoDream (reported 2026-03)Dual gate on time and activitySkip the nightly run when there are fewer than N merged commits since the watermark
OpenAI harness engineering (2026-02-11)~100-line map; doc-gardening agent; mechanical checks on docsLint the ledger (format, evidence present, Seen dates valid). Cap the router length.
Mitchell Hashimoto (2026-02-05) + magpie’s own TestMovedBuiltinsSayTheirPluginFixes the harness/tooling instead of only adding proseWhen promoting, prefer a test, lint or hook over a sentence, and record “Enforced in” (as zero-depth does)
dsh Agent NotesLifecycle folders plus gate scripts enforce formatLesson states: proposed → active → promoted / rejected / retired

What magpie already does that the others don’t:

  • It counts separate days, not raw occurrences, which damps one bad afternoon.
  • It uses a commit watermark.
  • It reviews all merged commits on a schedule, not per incident or session.
  • It has a “Done right on” positive example.
  • It ties verification to the --match-head-commit merge gate.
  • No statement from yetone or Junjie on where the ledger + 3-day promotion idea came from. A Discord or direct question would settle it. Junjie’s Discord rules aren’t public.
  • Can’t rule out that the Claude agent that co-authored 303819c drew on oncall-kit, Claude Diary or the viral tasks/lessons.md template. There’s no evidence either way.
  • Not verified with dated primary sources: Windsurf Memories, Devin Knowledge, Claude Code # memory shortcut launch date, Peter Steinberger, Armin Ronacher, Steve Yegge, Simon Willison. They’re omitted rather than guessed.
  • AutoDream details rest on secondary blogs, and the Anthropic best-practices and power-user-tips pages are undated.
  • Reddit (via Exa site:reddit.com) returned nothing relevant.
  • Origin of the viral “tasks/lessons.md” template unknown.
  • lee259/lessons, a “verified engineering lessons” plugin that only captures after a passing test/build, is a relevant parallel, but I couldn’t get its creation date.

magpie

yetone on X

DeepSeek Harness

Anthropic

Individuals and Every

Vendors

Papers

Parallel repos