跳转到内容

Overnight and long-running agents

Field practice (X + Reddit mirrors + practitioner writeups, late 2025–Oct 2026) converges on one pattern: package work so an agent can finish without you, then isolate + verify + budget. People actually ship overnight batch refactors, test fills, dependency bumps, doc syncs, issue triage, PR farms, research crawls, eval harnesses, and CI self-heal — not “build the product while I sleep” as a first move. The recurring failure is not model quality; it is permission popups, laptop sleep, overlapping worktrees, self-reported green, token burn / silent spend, and auth expiry. Suitable jobs are bounded, verifiable, low blast-radius. Unsuitable: ambiguous UI taste, secrets/prod deploys, auth/payments, and tasks whose spec lives only in a human head. How-to stacks span Claude Code /schedule Routines + headless -p, Codex background/worktree/cloud, Cursor Automations/Projects, Hermes (Telegram + VPS orchestration), Herdr/amux/Nightcrawler DIY, worktrees, resume/checkpoints, budgets, gates, and notify-on-wake.

Independent of suggestive/UI digests; overlaps 2026-09-26 Cloud agent orchestrators in the wild on harness landscape only.


Job shapeField signalTypical artifact
Mechanical / batch refactorsMacRun overnight guide; Easton solo-company split; X “hand the overnight refactor to a background box” (@verbal72)Branch or draft PR per task
Missing tests / coverage fillsamux overnight playbook task examplesPR + CI green
Dependency / lint / type / deprecation sweepsCTO Scorecard “unattended” candidates; Codex full-auto sweet spot (Particula)Small PRs
Bug hunts with clear reproNightcrawler / amux board cards; Cursor Automations on Slack bugsFix PR
Doc gen / OpenAPI / README syncClaude Routines “weekly doc sync”; amux API-docs examplePR or committed docs
PR farms / backlog drainDevelopers Digest /loop night (59 PRs / 21 repos — with failures); Nightcrawler TASK_QUEUE; herdr-factory beltsMany draft PRs, human merge
Issue triage / labeling / stale-PR nudgeClaude Code Routines docs (weeknight backlog + Slack summary)Labels + comments, not merges
CI self-heal / autofix on redCursor Projects subscriptions on CI/PRs; Automations on PR pushFix commit or PR comment
Evals / research crawls / content batchesLoop Maker “venture / content” modes; LinkedIn overnight multi-agent rebuildReports, packages, drafts

Productized / always-on (not just “one night”)

Section titled “Productized / always-on (not just “one night”)”
  • Claude Code Routines — scheduled / GitHub / API triggers on Anthropic cloud; laptop closed (docs).
  • Cursor Automations + Projects — cron, Slack, Linear, PR events; Projects coordinator + cloud workers while laptop closed (Automations, Projects).
  • Codex cloud / background / worktrees — isolation + review; field: “Codex when I want a long unattended run” (@SilentDevTools).
  • Hermes on VPS — always-on Telegram/Discord gateway + cron; delegates to Claude Code/Codex; “agent I text from anywhere that ships code while I sleep” (@r_alx_z); EventBridge/systemd patterns (@c199benzene).
  • Cheap VPS / queue folder — Hetzner ~€6–€7/mo, Claude headless + budget wrapper (@AffiliateKaj); Azure VM so lid-close doesn’t kill the agent (@49agents).
  • Herdr / amux / Moadim — detach terminals / watchdog loop daemons so close-lid doesn’t stop workers (@jothantranston on Herdr; amux guide; Moadim local Rust loop (@Guilty0399887)).

Cautionary “overnight” stories (still useful)

Section titled “Cautionary “overnight” stories (still useful)”
  • Permission popup mid-run — left Claude overnight; stopped halfway waiting on a question (@embw_l0x).
  • Self-approved dangerous tool call — overnight Grok Build rewrote half a branch; one tool call slipped through (@0xMfox).
  • Silent wrong / hallucinated PR — Developers Digest night: agent claimed PR #51; branch existed, PR did not (writeup).
  • Unbounded spend — hung scheduled/persistent session burned Max allowance + ~$1k auto-recharges with no kill UI (Claude Code issue reports).
  • Non-coding crypto/trading “left Claude running” viral posts — treat as anti-pattern for software digests (unbounded external spend, no CI gate).

Reddit-adjacent field (via sentiment mirrors / LinkedIn): first overnight multi-agent attempts “followed instructions but missed relations/feel, hit rate limits, stopped midway”; successful rebuilds used checkpoint commits + handoff notes, HITL by day / AFK by night, and plan→delegate→review (Miguel Votre LinkedIn). Reddit threads repeatedly split tools: Cursor for IDE grind, Claude Code for long multi-file / overnight loops, Codex for cheap parallel / background.


  • Bounded — 2–4h human-equivalent; one paragraph with problem / done / proof (MacRun).
  • Verifiable — tests, typecheck, lint, build exit codes, screenshot/diff evidence; independent verifier ≠ writer (Loop Maker, Nightcrawler Codex audit, Leymish verifier subagent).
  • Low blast-radius — own branch/worktree; draft PR only; no prod credentials; path allowlists.
  • Reversible / idempotent — dependency bump, doc sync, generated client refresh, lint repair, flaky-test fix with clear expected behavior.
  • Pre-curated queue — human-accepted task class; issue text is data, not authority (CTO Scorecard Q15).
  • Ambiguous UI taste, brand feel, product judgment.
  • Auth, payments, migrations, prod config, secrets, irreversible external actions (deploy, publish, spend, force-push).
  • Broad “refactor everything” without file boundaries.
  • Spec that lives only in someone’s head.
  • Parallel agents on the same files without worktrees (morning merge hell — Easton).
  • Anything that needs a permission popup or open question at 03:00.

Rule of thumb from X: pair when fuzzy; overnight when seams are clear (@ChristianXCesar).


HarnessTriggerNotes
Claude Code Routines/schedule, cron, GitHub events, API POSTCloud session; no approval picker mid-run — scope repos/network/connectors tightly (docs)
Cursor AutomationsSchedule, GH/GL, Slack, Linear, webhookCloud agents; can open PRs; secrets via dashboard (docs)
Cursor ProjectsCoordinator + subscriptions (Slack/schedule/PR/CI)Long-lived shared context; cloud by default (blog)
Codex cloudApp/API/integrationsParallel sandboxed tasks → review → PR

B. Self-host / VPS (data residency, multi-harness)

Section titled “B. Self-host / VPS (data residency, multi-harness)”
PatternRole
HermesAlways-on orchestrator: Telegram/Discord, memory, cron; delegates Claude Code (-p print or tmux interactive) / Codex
HerdrTerminal multiplexer + agent lifecycle; close lid, keep running
amuxTask board + tmux watchdog + worktrees + mobile stream
Nightcrawler / CamusBash pipeline: plan (Opus) → audit (Codex) → implement → review → commit → verify; Telegram notify
herdr-factory / Open SWE / OpenHandsTicket→PR factories (see orchestrators digest)
GitHub Actions + claude-code-actionCron roles (CEO/builder/verifier) with shell guardrails after agent step (Leymish)
tmux + claude -p loopMinimal DIY (MacRun script)

C. Isolation, resume, budgets, gates, notify

Section titled “C. Isolation, resume, budgets, gates, notify”
  1. One task → one branch → one worktree (non-negotiable for parallel nights).
  2. Permission mode set before sleep — e.g. accept edits + tests + commit/push/PR; deny deploy/secrets/out-of-repo (MacRun). Or phone-gate writes via hooks (Pushary).
  3. Hard caps — --max-turns, job timeout-minutes, $ budget per job, concurrent task caps, diff-size limits.
  4. Resume / checkpoints — git commit + handoff note / PROGRESS.md / JSONL journal after each card; never rely on chat history alone across compact/crash.
  5. Verifier ≠ implementer — second model or fresh-context subagent; CI on the PR is the ground truth.
  6. Notify on wake — Telegram/Slack morning brief; count PRs vs task list; reconcile gh pr list against agent claims.
  7. Machine must stay awake — VPS/cloud preferred; laptop sleep is the #1 empty-morning cause.

Minimal headless loop (field-standard shape):

Terminal window
# run-overnight.sh — sequential, one worktree per task
for task in tasks/*.md; do
name=$(basename "$task" .md)
dir="../work/$name"
git worktree add "$dir" -b "agent/$name" 2>/dev/null
(
cd "$dir" || exit 1
claude -p "Complete the task in $task. Commit, push, gh pr create. Stop when tests pass or blocked; say which." \
--permission-mode acceptEdits \
--max-turns 80 \
> "../logs/$name.log" 2>&1
)
done

Run under tmux/systemd on a box that does not sleep.


ModeWhat it looks likeMitigation
Stuck on permission / questionPopup since 01:00; empty morningPre-set permission mode; never-ask + decision log; phone-gate only for risky tools
Infinite / revision loopsSame reject forever; TaskCreate explosionMax repair rounds (2–5); convergence detector; process-group watchdog kill
Burn tokens / silent spendSleep-polling; hung scheduled session; no spend capCost budget kill-switch; --max-turns; disable auto-recharge for overnight; alert before overage
Silent wrong“Tests passed” / fake PR number / forged ratificationCI as judge; reconcile claims vs gh pr list; independent reviewer model
Merge conflicts / index racesTwo agents, one checkoutWorktrees; non-overlapping paths; sequential if unsure
Laptop sleep / SSH dropZero outputVPS, Routines, Cursor cloud, Herdr detach, caffeinate only as last resort
Auth expiry--print workers die mid-night (OAuth wipe)Prefer cloud Routines / API patterns that don’t rely on fragile local OAuth; monitor login health
Billing / CI wallPRs open but Actions red for paymentPreflight gh auth + Actions billing + disk + rate limits
Compaction driftGreat plan → compact → garbage spiralCheckpoint _session_state.json / PROGRESS before compact; don’t trust prose memory
Scope creepAsked for one function, rewrote three modulesTask “Do not:” list; diff-size cap; morning close-with-reason → tomorrow’s task

  • Tasks are 2–4h, each with Problem / Done / Proof / Do not
  • One worktree (or cloud sandbox) per task; no shared dirty tree
  • Permission mode + deny list set; no prod secrets in env
  • Caps: max-turns / timeout / $ budget / max concurrent
  • Verifier path: tests + CI protected on main; optional second-model review
  • Host will not sleep (VPS/cloud/Routines)
  • Preflight: gh auth status, Actions billing OK, disk OK, main already green
  • Notify channel ready (Telegram/Slack); morning: reconcile PR list vs claims
  • Kill switch known (stop routine / kill tmux / Pushary / dashboard)
## Task: paginate the /orders endpoint
Problem: GET /orders returns every order; large accounts time out.
Done: accepts ?cursor and ?limit (max 100), returns next_cursor.
Proof: tests/orders_pagination.test.ts passes; existing tests pass.
Do not: change existing response field shapes; no deploy; no secrets.
Execute only the accepted task in an isolated checkout.
Stop after two failed repair attempts, any protected-path touch,
or unexpected external dependency.
Produce a draft PR with evidence (diff, commands, test output, residual risk).
Do not merge or deploy. Never ask questions — log assumptions in DECISIONS.md and continue,
or mark BLOCKED with exact reason after three attempts.
Validate whether this item matches the approved unattended task class.
Check scope, affected paths, data, reversibility, acceptance criteria, and required tests.
Reject if any field is missing or risk exceeds policy.
Triage each completed run: accept for review | refine with one bounded request | reject | escalate.
Cite evidence. Prefer closing a bad PR with one-line why over merging and cleaning later.
Copy reject reasons into tomorrow’s task file.
/schedule daily PR review at 9am
/schedule weeknight issue triage and Slack summary at 22:00
/schedule clean up feature flag in one week
/loop-maker Add 3 features from FEATURES.md — build all, test e2e, don't stop until feature_list.json all pass or budget hit. Draft PRs only.

Direct site:reddit.com Exa queries returned empty this run (likely index/block). Used Reddit sentiment mirrors (sentinel-team snapshots) for Cursor/Claude/Codex role-split threads, plus LinkedIn overnight multi-agent rebuild narrative. Treat Reddit URLs as secondary; re-pull via old.reddit JSON or browser if primary quotes needed.


  1. Hermes vs Herdr vs Cursor Projects for Noa’s default overnight lane — multi-harness on owned VPS vs productized cloud (see orchestrators digest).
  2. Auth reliability for local claude -p 8h+ runs still contested in issue trackers — prefer Routines/cloud when unattended length matters.
  3. Spend caps still uneven across vendors; DIY kill-switch + budget wrapper remains necessary.
  4. Fresh primary Reddit thread URLs for “left Claude running overnight” — gap this digest; X + practitioner blogs carried most field evidence.
  5. Whether eval farms / research crawls deserve a separate checklist (artifact schema, rate limits, dataset integrity) beyond coding PRs.
  • Collected 2026-10-02 via user-X search_posts_all + user-exa web_search_exa / web_fetch_exa.
  • Did not wait on suggestive/UI digest workers.
  • Confidence: high on patterns (worktrees, caps, verify, VPS/cloud); medium on exact product UI names that churn weekly; low on Reddit primary URLs this run.