Agent lessons ledger lineage
Agent lessons ledger lineage
Section titled “Agent lessons ledger lineage”Companion to 2026-10-08 yetone magpie agent workflow. Related: 2026-09-28 DeepSeek harness dsh system prompt. It also feeds my lessons-ledger skill, which already cites magpie as its source pattern.
Times are UTC+8 unless marked. Influence means there’s an explicit citation. Parallel means similar idea, no evidence of influence.
1. Short answer
Section titled “1. Short answer”No named origin. magpie’s LESSONS.md and docs/code-standards.md cite no outside inspiration. The only explicit outside influence in the repo is contributor Junjie Zhou’s issue #883, which names DeepSeek Harness (dsh). It shaped the subsystem docs, the router and the “report checks not run” piece, not the ledger. Each piece has well-dated earlier parallels. The closest single cousin is Anthropic’s oncall-kit (Aug 2026). It keeps a lessons.md and moves a pattern into the reference docs after 3 entries.
2. When the magpie loop started, and does it name sources?
Section titled “2. When the magpie loop started, and does it name sources?”In-repo sequence (commit times converted from UTC):
| When (UTC+8) | What | Source |
|---|---|---|
| 10-02 19:36 | AGENTS.md gets a map of moved built-ins, backed by a test (TestMovedBuiltinsSayTheirPlugin) | commit fa9daa6 |
| 10-05 15:28 | Junjie Zhou opens #883 (subsystem design docs + PR semantic diff for agent work). He explicitly proposes checking how DeepSeek Harness (dsh) organizes architecture docs, agent context and dev flow. yetone keeps it as a discussion and invites a pilot PR. | #883 |
| 10-05 19:10 | PR #889 merged: AGENTS.md “Subsystem design and review” pointer + docs/subsystems/README.md. It includes “Keep each fact in one reference and link to it from agent instructions” and “Report checks that were not run.” The PR body itself says “Go tests were not run for this documentation-only change”. yetone’s merge comment records what he ran and on which head. | PR #889, 390db18 |
| 10-06 01:37 | docs/code-standards.md created: “code standards distilled from PR reviews… (Junjie Zhou on Discord)”. The first tables are Junjie’s six review rules. The rest come from PRs #901–#930. | 926bcaa |
| 10-06 02:17 | Adds acceptance criteria, gh pr merge --match-head-commit, and review patterns “distilled from recent PR reviews (#706–#931)“ | a5df344 |
| 10-06 03:10 | LESSONS.md born: “keeps what the nightly review of merged work finds, and agents read it before changing code”. The first review covered 10-05: 220 commits, ~170 releases, ~40 fixes of earlier releases. Three lessons that recurred dozens of times went straight to the standards. Carries a Co-Authored-By: Claude Opus 5.5 trailer. | 303819c |
| 10-06 03:49 | Follow-up edit | 4b03eda |
| 10-07 00:23 | Second nightly review | 18d973b |
| 10-07 23:46 | First threshold promotion: four lessons seen a third day moved to the standards | 8731dd9 |
Named sources:
- LESSONS.md describes itself only in-house (“Each night the day’s commits are reviewed…”).
- code-standards.md opens “These rules come from what PR reviews keep asking for.”
- Grepping both files for inspir/compound/karpathy/postmortem/ADR/ralph found nothing.
- The only credited humans are Junjie Zhou (Discord review rules) and the PR reviewers.
3. Lineage timeline per piece
Section titled “3. Lineage timeline per piece”a) Lessons ledger: an agent-maintained file of mistakes, read before work
Section titled “a) Lessons ledger: an agent-maintained file of mistakes, read before work”All parallel:
- 2023-08: ExpeL extracts “insights” from experience and edits them with ADD/EDIT/UPVOTE/DOWNVOTE. arXiv 2308.10144
- 2025-02-06: Cline Memory Bank, structured markdown the agent re-reads every session. blog
- 2025-05-11: Karpathy, “system prompt learning”. X
- 2025-06-04 / 07-02: Cursor Memories, beta in 1.0 then GA in 1.2 “with user approvals for background-generated memories”. 1.0 forum, 1.2 forum
- 2025-07-14: Geoffrey Huntley, Ralph: “When you learn something new… update @AGENT.md using a subagent but keep it brief”. ghuntley.com/ralph
- 2025-08-18: Kieran Klaassen / Every coins “compounding engineering”: “every bug becomes a permanent lesson, and every code review updates the defaults.” every.to. Formalized 2025-12-11 as plan/work/review/compound, with lessons in
docs/solutions/. every.to, plugin - 2026-01-02/03: Boris Cherny: “Anytime we see Claude do something incorrectly we add it to the CLAUDE.md”. He calls the PR-tagging flow “our version of @danshipper’s Compounding Engineering”. post, thread
- 2026-02-05: Mitchell Hashimoto, harness engineering: “anytime you find an agent makes a mistake, you take the time to engineer a solution such that the agent never makes that mistake again”. mitchellh.com
- 2026-08-18: Anthropic oncall-kit
Lessons.md, “a running log of every incident… Claude appends to it on its own”. blog, template - 2026-09: the viral “tasks/lessons.md self-improvement loop” CLAUDE.md template on X, e.g. 2026-09-29. Original author not established.
b) Evidence on each lesson (SHAs/PRs, “this is evidence, not reflection”)
Section titled “b) Evidence on each lesson (SHAs/PRs, “this is evidence, not reflection”)”All parallel:
- 2025-12-01: Lance Martin’s Claude Diary.
/diaryentries plus/reflect, which tracks processed entries inprocessed.log. post - 2026-01-15: GitHub Copilot Memory. Each memory carries citations to code locations, which are checked against the current branch before use. blog, docs
- 2026-07-04: Lilian Weng surveys evidence-driven harness edits (AHE: failure evidence, root cause, fix, predicted impact per change). post
- 2026-08: oncall-kit entries cite incidents, and the setup provenance reads
(seen 3×: INC-…)/(seen 1×, unverified). repo - 2026-09-08: oriko9/tribunal docs/LESSONS.md, “This is evidence, not reflection”. 2026-09-03→19: drewlr/AI-spec-kit lessons.md, “each entry keeps the incident that produced it”.
c) Promote after N, from ledger into standards
Section titled “c) Promote after N, from ledger into standards”All parallel. None counts separate days the way magpie does.
- 2023-08: ExpeL importance counts. A new insight starts at 2 and is dropped at 0. arXiv
- 2025-10: ACE, bullets with helpful/harmful counters plus a Curator that grows and refines them. arXiv 2510.04618
- 2025-12-01: Claude Diary
/reflect: “2+ occurrences = pattern, 3+ = strong pattern”, then patterns go into CLAUDE.md. post - 2026-03 (reported): Claude Code AutoDream background memory consolidation, gated on ≥24h and ≥5 sessions (orient/gather/consolidate/prune). Secondary sources only: levelup 2026-03-26, wmedia 2026-04-05
- 2026-08 (closest): oncall-kit Graduation. “When a tag accumulates 3 entries with the same mechanism, or an entry stops being a story and becomes a procedure… promote it into that class’s reference file by PR… replace the entries with one pointer line. Record the promotion itself as an entry.” The blog (2026-08-18) says “If the same pattern shows up enough times, we promote it into the investigation skill itself.” repo, blog. The repo was created 08-04 06:18 UTC+8.
- Human-practice analogues (postmortem → runbook, ADR status lifecycle): dsh’s Agent Notes uses proposed/implemented/rejected/archived folders and covers “a gap a postmortem surfaced”. Parallel for the ledger. dsh influenced magpie’s subsystem docs (see e), but no lessons file exists in dsh’s tree.
d) Verification record (“what was run / not run”; real FAIL before the fix; merge at the reviewed head)
Section titled “d) Verification record (“what was run / not run”; real FAIL before the fix; merge at the reviewed head)”- Influence (partial): magpie’s “Report checks that were not run” first appears in the dsh-inspired PR #889 (10-05). dsh’s AGENTS.md says “report only commands run”. Its dsh-pre-push-checks asks for “the narrowest available test… that would fail for its regression”. #883 explicitly names dsh as a model, so this is the one piece with a documented line. The exact “not run” wording is magpie’s own.
- Parallel:
- Anthropic, Effective harnesses for long-running agents (2025-11-26): verify end-to-end before marking a feature done; progress file across sessions. yetone linked this post on 2026-10-03 17:01 as the start of harness engineering (X). That shows he read it, but he never ties it to LESSONS.md.
- oncall-kit’s “Could not verify: … the exact one-line command the owner should run” (2026-08).
- riusprime/zero-depth docs/LESSONS.md, first committed 10-07 03:33 UTC+8 (~a day after magpie’s), with “NOT YET RUN is a legal status”. No link to magpie.
--match-head-commitgating appears to be magpie’s own addition (10-06 02:17). I found no earlier agent-workflow source for it.
e) Thin AGENTS.md router (“before X, read Y”; CLAUDE.md = @AGENTS.md @LESSONS.md)
Section titled “e) Thin AGENTS.md router (“before X, read Y”; CLAUDE.md = @AGENTS.md @LESSONS.md)”- Influence: dsh via #883/#889 (10-05). dsh’s AGENTS.md routes to
docs/architecture.md/docs/AGENTS.md, its CLAUDE.md just points at AGENTS.md, and it hasdocs/subsystems/. magpie adopteddocs/subsystems/and the pointer style the same day. - Parallel / background:
- AGENTS.md as a cross-tool standard: agentsmd/agents.md repo created 2025-08-20 01:22 UTC+8.
- Huntley’s later how-to-ralph-wiggum: “Keep @AGENTS.md operational only… A bloated AGENTS.md pollutes every future loop’s context.”
- OpenAI Harness engineering (2026-02-11): “give Codex a map, not a 1,000-page instruction manual”, AGENTS.md as table of contents (~100 lines), docs/ as system of record, a recurring doc-gardening agent. A Chinese X recap yetone replied to credits this post with starting harness engineering (X 10-03).
- Anthropic best practices (undated): “Treat CLAUDE.md like code… prune it regularly”, plus
@imports.
4. Who cites magpie’s LESSONS.md
Section titled “4. Who cites magpie’s LESSONS.md”Nobody found.
- X searches turned up only general magpie promo posts, e.g. 1, 2. Queries tried: “LESSONS.md magpie”, “magpie yetone lessons/nightly/code standards”, and “match-head-commit / reviewed-through / code-standards.md”.
- HN Algolia had 0 hits for “LESSONS.md” and “yetone magpie”.
- Exa found only the file itself and mirrors. A PyShine source tour (2026-09-30) and AICoder (2026-10-03) cover the product and predate LESSONS.md.
- GitHub code search for magpie’s markers (“separate days” in LESSONS.md,
reviewed-through) found no copies. - The only downstream artifact is my own lessons-ledger skill. The loop is two days old, so silence is expected.
5. yetone’s earlier signal
Section titled “5. yetone’s earlier signal”- 2026-09-13 01:23: what he learned building avante.nvim (“a year before claude code”, in the weak-model era) went into Alma. That’s about harness design, not a lessons file.
- 2026-09-13 01:40: you can tell how long someone has been building harnesses from their Edit tool.
- 2026-10-03: says he builds magpie with Claude Code. The founding LESSONS.md commit carries a Claude co-author trailer.
- 2026-10-03 17:01: links Anthropic’s long-running harness post as where “harness engineering” took off “last year”.
- His other repos have no lessons/standards files: avante.nvim has no AGENTS.md/CLAUDE.md/LESSONS.md; cumora (2026-08) has CONTRIBUTING/docs only.
- His skill repos native-feel-skill (2026-05) and kill-ai-slop (2026-07) show a habit of turning research into agent-readable rule files, which is a related instinct.
from:yetonesearches for lessons/教训 (lessons)/复盘 (retrospective)/nightly/AGENTS.md/CLAUDE.md found nothing earlier about this loop. No blog or talk found.
6. Close cousins that do parts better, and what to borrow for lessons-ledger
Section titled “6. Close cousins that do parts better, and what to borrow for lessons-ledger”| Cousin (date) | Does better | Borrow into lessons-ledger |
|---|---|---|
| Anthropic oncall-kit (2026-08-04 repo / 08-18 blog) | Graduation also fires when an entry “becomes a procedure”, not only on count. Promotion happens by PR, leaves a pointer line, and is itself logged. A status banner serves as a resume point. “Could not verify” entries name the owner’s command. Humans prune. | Promote-by-PR with pointer line. Log each promotion as an entry. Add a could-not-verify entry type. Add a “procedure-shaped” promotion trigger next to ≥3 days. |
| GitHub Copilot Memory (2026-01-15) | Citations are re-validated against the current branch before a memory is used. Unused memories expire (28 days). | Before reading or bumping a lesson, check that its evidence SHAs/paths still exist. Use auto-expiry instead of my soft “~30 days” prune. |
| ACE (2025-10) / ExpeL (2023-08) | Counters run both ways (helpful/harmful, upvote/downvote to zero) | Track a “done right on” counter alongside Seen. A lesson repeatedly done right can retire. |
| Claude Diary (2025-12-01) | processed.log stops double counting. Rules that keep being violated get strengthened. | Keep a processed log next to the watermark. Escalate wording or add a test when a promoted standard is violated again. |
| AutoDream (reported 2026-03) | Dual gate on time and activity | Skip the nightly run when there are fewer than N merged commits since the watermark |
| OpenAI harness engineering (2026-02-11) | ~100-line map; doc-gardening agent; mechanical checks on docs | Lint the ledger (format, evidence present, Seen dates valid). Cap the router length. |
Mitchell Hashimoto (2026-02-05) + magpie’s own TestMovedBuiltinsSayTheirPlugin | Fixes the harness/tooling instead of only adding prose | When promoting, prefer a test, lint or hook over a sentence, and record “Enforced in” (as zero-depth does) |
| dsh Agent Notes | Lifecycle folders plus gate scripts enforce format | Lesson states: proposed → active → promoted / rejected / retired |
What magpie already does that the others don’t:
- It counts separate days, not raw occurrences, which damps one bad afternoon.
- It uses a commit watermark.
- It reviews all merged commits on a schedule, not per incident or session.
- It has a “Done right on” positive example.
- It ties verification to the
--match-head-commitmerge gate.
7. Gaps
Section titled “7. Gaps”- No statement from yetone or Junjie on where the ledger + 3-day promotion idea came from. A Discord or direct question would settle it. Junjie’s Discord rules aren’t public.
- Can’t rule out that the Claude agent that co-authored 303819c drew on oncall-kit, Claude Diary or the viral tasks/lessons.md template. There’s no evidence either way.
- Not verified with dated primary sources: Windsurf Memories, Devin Knowledge, Claude Code
#memory shortcut launch date, Peter Steinberger, Armin Ronacher, Steve Yegge, Simon Willison. They’re omitted rather than guessed. - AutoDream details rest on secondary blogs, and the Anthropic best-practices and power-user-tips pages are undated.
- Reddit (via Exa site:reddit.com) returned nothing relevant.
- Origin of the viral “tasks/lessons.md” template unknown.
- lee259/lessons, a “verified engineering lessons” plugin that only captures after a passing test/build, is a relevant parallel, but I couldn’t get its creation date.
8. Sources
Section titled “8. Sources”magpie
- LESSONS.md
- docs/code-standards.md
- Commits 926bcaa, a5df344, 303819c, 4b03eda, 18d973b, 8731dd9, 390db18
- Issue #883, PR #889
yetone on X
DeepSeek Harness
Anthropic
- oncall-kit, lessons template, on-call blog 2026-08-18, ClaudeDevs X 2026-09-09
- long-running harnesses 2025-11-26
- best practices, power user tips
Individuals and Every
- Karpathy: system prompt learning 2025-05-11, llm-wiki 2026-04-04
- Huntley: Ralph 2025-07-14, how-to-ralph-wiggum
- Every: 2025-07-16, 2025-08-18, 2025-12-11, 2026-02-09, plugin
- Boris Cherny: thread 2026-01-03
- Lance Martin: Claude Diary 2025-12-01
- Mitchell Hashimoto: 2026-02-05
- Lilian Weng: 2026-07-04
Vendors
- OpenAI: Harness engineering 2026-02-11, Codex customization
- GitHub Copilot Memory: blog, changelog 2026-01-15
- Cursor: 1.0 Memories 2025-06-04, 1.2 GA 2025-07-02
- Cline: Memory Bank 2025-02-06
- agentsmd/agents.md
Papers
Parallel repos