resume · git:20260824.38e4555 · 2026-08-24 · sha256 88b36bd7d9caa604

resume git:20260824.38e4555A

Immutable. This exact content is served forever at /api/v1/blob/88b36bd7d9caa604.

---
name: resume
tier: D
category: orchestration
description: |
  Restore an arc that was killed mid-flight by something external — API 529 /
  500, ENOTFOUND, ConnectionRefused, an expired login, a laptop that slept, a
  Ctrl-C. The operator restarts, types `resume`, and means *continue as though
  nothing had stopped you*: same arc, same plan, subagents and workflows that
  died put back on their feet. It does NOT mean "wrap up what you can".
  Nothing was written down in advance — the session did not intend to stop —
  so state is reconstructed forensically from the transcript, from the dead
  agents' surviving output files on disk, and from the tree.
  Distinct from `handoff` (a doc written on purpose for the next agent),
  `handback` (a message written on purpose for the human), and `persist`
  (a loop that restarts contexts by design). All three assume a VOLUNTARY
  stop that left an artifact. `resume` is the involuntary case.
  USE WHEN: "resume", "continue", "keep going", "carry on", "pick up where you
  left off", "as you were", "continue /autonomous", "resume the loop", or any
  session whose previous turn ended in an API error, a connection failure, an
  expired login, or an interrupt.
  NOT FOR: starting new work; a deliberate fresh-session pickup that has a
  handoff doc (read the doc); a finished arc the user is asking about.
---

# resume — continue the arc, do not conclude it

**`resume` is a restoration command, not a summarization command.**

## Why this skill exists

**The table below** is regenerated by `scripts/resume_corpus_stats.py`, which
ships with the skill. Run it; do not trust the numbers here on their own.

```
python3 "$SKILL_DIR/scripts/resume_corpus_stats.py"
```

Over this author's transcript corpus it reported **91** resume-class turns
(**59** a bare `resume`, across 10 distinct surface forms):

| | | |
|---|---|---|
| Preceded by an API error | 48 | 53% |
| Preceded by a user interrupt | 7 | 8% |
| Had a worker **in flight** when the session died | 22 | 24% |
| Re-spawned a worker afterwards | 15 | 16% |
| **Queried the workers that were already running** | **2** | **2%** |
| **Produced no tool calls at all** — a text-only wrap-up | **37** | **41%** |
| Acknowledged a death **and did not restore it** | 9 | vs 6 restored |
| `resume` typed **again** — the first one didn't take | 16 | 18% |

Your corpus will differ, and the absolute counts drift upward as it grows.
The ratios are what the argument rests on.

The dominant failure is not that the agent forgets what it was doing. It is
that it treats `resume` as *"deliver what survived"*. Transcripts say it in
the agent's own words — **"three follow-on sub-agents died on the session
limit; I'll note the gap"** — while those agents' complete transcripts sat on
disk, unread.

## The mechanisms you will otherwise get wrong

These are properties of the harness, checkable by reading a transcript. Each
was reached by getting it wrong first.

**1. A dead worker leaves no orphaned tool call.** The intuitive detector — a
`tool_use` with no `tool_result` — finds nothing, because workers launch
**asynchronously**: the spawn's result attests to *launching*, never
*finishing*, and completion arrives later as a separate `<task-notification>`.
**A spawn whose id never appears in a completion is what died.**

**2. Completion notices arrive on more than one record shape** —
`queue-operation`, `attachment` and `user`. Reading only one of them reports
finished workers as dead.

**3. There are three launch receipts, not one.** An agent's says `Async agent
launched successfully` with an `agentId:`; a background shell's says `Command
running in background with ID:`; a workflow's says `Workflow launched in
background. Task ID:` with a `Transcript dir:`. A detector that knows one of
them silently drops the others.

**4. The obvious output path is the one the crash deletes.** The launch
receipt points into `/private/tmp/...`, wiped by exactly the reboot this skill
is about. The harness also writes a durable copy under
`<project>/<session-id>/subagents/`. The scan prefers it.

**5. Liveness is a property of the worker KIND, and for one kind it is
unknown.** `Agent`/`Task` run in-process and die with the parent. A
`run_in_background` shell is a separate process and can outlive it. A
**workflow's** receipt says "launched in background", but this author's corpus
did not settle whether one outlives the parent — so the scan reports
`unknown-workflow` rather than guessing. Calling a live worker dead invites
re-running work that is still going.

**6. Recovered transcripts are large and contain credentials.** See *Privacy*.
Never `cat`, `Read` or `tail` one.

## Procedure

### 1. Scan before you say anything

The script lives beside this file, not in your working directory:

```bash
SKILL_DIR=~/.claude/skills/resume        # or the base directory printed above

python3 "$SKILL_DIR/scripts/resume_scan.py"                    # newest session for cwd
python3 "$SKILL_DIR/scripts/resume_scan.py" --previous         # the session BEFORE it
python3 "$SKILL_DIR/scripts/resume_scan.py" --list-sessions    # choose explicitly
python3 "$SKILL_DIR/scripts/resume_scan.py" --session <path.jsonl>
```

**If the crash dropped you into a FRESH session, the newest transcript is the
one you are sitting in — not the one that died.** The scan says so when it
sees no workers and other sessions exist; `--previous` is then the answer.
This was a real defect: auto-selection returned the live, near-empty session
and reported "nothing died mid-flight".

### 2. Check the surface that killed you is actually back

Clear it *before* re-spawning, or you will re-die and the operator will type
`resume` again — see the retyped row in the table above.

| Cause | Clear it with |
|---|---|
| `auth_expired` | `/login`; check `gh auth status`, `railway whoami` |
| `usage_limit` | note the reset time; do not fan out into the ceiling |
| `rate_limited` | back off; one call before many |
| `api_overload` / `api_5xx` | retry one call before spawning ten |
| `network` | one cheap real call, not an assumption |
| `stalled` | the response stopped arriving; retry once before fanning out |
| `machine_slept` | the laptop suspended mid-response — nothing is wrong upstream |
| `context_too_long` | the prompt exceeded the window; shrink it before retrying |
| `user_interrupt` | ask what the operator wanted changed before continuing |
| `api_error_unclassified` | an error the scan does not recognise — **read the evidence line**; do not treat it as clean |
| `clean_or_unknown` | no termination evidence in the window; the session may have ended normally, or the record may be gone |

### 3. Re-snapshot the tree (P15) — an edit may have died half-written

A worker that died *during* an edit leaves a partial file. `git status`,
`git diff`, branch, ahead/behind, open PRs, CI. Confirm the tree is coherent
before building on it.

### 4. Triage each worker — three outcomes, not one

| Finding | Action |
|---|---|
| Digest shows the work **finished**, only the report was lost | Fold it into the arc. **Do not re-spawn.** |
| Digest shows **partial** progress | Re-spawn scoped to *the remainder*, handing over what its predecessor established. |
| No digest — died early, or nothing survived | Re-spawn from the original prompt. The text render shows the first 4,000 chars and says `TRUNCATED` when it cuts; **`--json` carries the prompt whole** — re-spawn from that, not from the render. If a secret shape was masked, `redacted_kinds` is set and the placeholder is what you get: restore the real value from your own environment before re-spawning. |
| **REPORTED FAILURE** | It came back, and came back broken. Read its summary; a failed worker is not a finished one. |
| `liveness: possibly-live` | **A separate process that wrote after the session died.** Check before re-running: re-running a live deploy or migration is worse than waiting. |
| `liveness: unknown-no-output-file` | Liveness cannot be determined. Verify before re-running anything with side effects. |

### 5. Continue the arc

Pick up where the plan stopped. If a `/loop`, `/autonomous` or workflow was
driving, **restart the driver** — the arc is not over because its engine died.

### 6. Report restoration, not conclusion

What died, what was recovered from disk, what was re-spawned, what you are
doing now. One short paragraph.

## Privacy — read before pasting any of this anywhere

Recovered text is other sessions' output. Surviving worker files in this author's corpus contained secret-shaped
strings (`sk-ant-`, `ghp_`, `github_pat_`, `AKIA…`, `xoxb-`, bearer headers).

The scan masks secret-shaped strings in the fields it prints — recovered
prose, termination evidence, and each worker's command, prompt, summary and
description — and names what it masked. That list is the guarantee; nothing
wider is claimed. An earlier version of this paragraph promised masking in
"everything it prints" while the command line and the prompt went out raw in
both the render and `--json`.

**Even within that list it is a blunt instrument and does not make the output
safe to publish**: it matches shapes, not secrets, so an unusual credential
format, customer data, or private source passes straight through. File paths
and tool names are not masked at all. Treat scan output as sensitive, keep it
in the session, and read it before pasting it anywhere.

## Anti-rationalization

| Excuse | Reality |
|---|---|
| "I'll summarize where things stand." | That is the measured failure — see the "no tool calls at all" row above. `resume` asks you to *act*. |
| "The workers are gone, I'll note the gap." | Their transcripts are on disk, and the durable copy survives the crash. "Noting the gap" discards recoverable work. |
| "I'll re-spawn everything to be safe." | Two of three outcomes are *not* re-spawn, and a live background process can be re-run into a double deploy. |
| "The scan says nothing died." | Check WHICH session it read. If you are in a fresh one, use `--previous`. |
| "It reported, so it's fine." | A worker can report **failed** or **killed**. Reported ≠ succeeded. |
| "The user typed one word, so they want something small." | They typed one word because they expect the arc intact. Scope is the arc. |
| "CI was green when it died, so it's done." | The last thing you *observed* is not the last thing that *happened*. |

## Composition

- **P15 Snapshot** — step 3 is the standard snapshot, not a lighter one.
- **P9 Wait** — a watcher killed with the session is not watching; restart it.
- **P5 Fanout** — re-spawned workers follow the same worktree isolation rules.
- **`handoff`** — if the arc genuinely cannot continue, *then* write a handoff.
- **`handback`** — if the block needs a human, end in an ask, never a silent stop.

## Validation

```bash
PYTHONDONTWRITEBYTECODE=1 python3 -m pytest tests/ -q
python3 scripts/resume_corpus_stats.py          # regenerates the table above
```

140 unit tests. Coverage is deliberately weighted toward the paths a green
suite hid: notification records in all three shapes, background-shell
detection, liveness in both polarities, failed/killed completions, non-dict
JSON lines (a bare number in raw stdout happens to parse as JSON, and crashed
the digest), truncated
tails, zero-record files, directories and FIFOs, and secret redaction.

Both false-positive classes in termination detection are pinned by regression
tests, and the signatures are copied from strings the harness actually emits —
two earlier sets were written to satisfy the regex rather than copied from
production, and classified almost nothing.

Numeric bounds live in the `LIMITS` table at the top of `resume_scan.py`. The
tests that pin them assert **hardcoded** expected values rather than comparing
a constant to itself — an earlier set sized its inputs from the constant it
was checking, so widening the constant widened the input and the assertion
could never fail.