deep-audit · diff

v1.0.0 to git:20260821.6600333

62 added, 188 removed. Audit A to A.

---
name: deep-audit
- description: |
- Deep consistency audit of the entire repository infrastructure.
- Launches 4 parallel specialist agents to find factual errors, code bugs,
- count mismatches, and cross-document inconsistencies. Then fixes all issues
- and loops until clean.
- Use when: after making broad changes, before releases, or when user says
- "audit", "find inconsistencies", "check everything".
- allowed-tools: ["Read", "Write", "Edit", "Bash", "Glob", "Grep", "Agent", "Task"]
- disable-model-invocation: true
+ description: Comprehensive adversarial audit of a theory, proof, math/econ paper, codebase, or set of claims — decompose into components, fan out independent skeptics that must return CONCRETE defects, adjudicate every finding with a separate judge, fix all confirmed defects, then re-verify. Use when correctness must be bulletproof and single-pass or round-by-round review is too slow and too shallow. Invoke for "audit this rigorously", "find ALL the bugs/gaps", "make this rock solid", "converge faster on correctness".
+ allowed-tools: ["Read", "Grep", "Glob", "Bash", "Write", "Edit", "Agent", "Task"]
metadata:
- author: Claude Code Academic Workflow
- version: 1.0.0
+ protocol: threat-prioritization
---
- # /deep-audit — Repository Infrastructure Audit
-
- Run a comprehensive consistency audit across the entire repository, fix all issues found, and loop until clean.
-
- ## When to Use
-
- - After broad changes (new skills, rules, hooks, guide edits)
- - Before releases or major commits
- - When the user asks to "find inconsistencies", "audit", or "check everything"
-
- ## Workflow
-
- ### PHASE 0: Mechanical checks (run FIRST, cheap, deterministic)
-
- Before spawning agents, run the mechanical parity checks:
-
- ```bash
- python3 scripts/check-skill-integrity.py --verbose
- ```
-
- This catches four classes of bug that agent-based audits have historically missed:
-
- 1. Frontmatter `allowed-tools` ↔ body tool-invocation parity (e.g. body spawns `Task` but `Task` not in `allowed-tools` — the v1.7.0 PR #92 miss).
- 2. `argument-hint` ↔ body flag parity (flags documented but not advertised, or vice versa).
- 3. Internal markdown anchors resolve (no broken `[text](path#anchor)` links — the `#category-11-numerical-discipline` miss on PR #87).
- 4. Rule `paths:` ↔ skill implementation parity (rule claims skill follows protocol but skill body has none of the protocol keywords — the `/interview-me` miss on PR #92).
-
- If Phase 0 reports P0 or P1 findings, fix them (or tune the regex if they are false positives) **before** launching the 4 agents. The mechanical layer is cheaper and more precise than agent prompts for these classes.
-
- ### PHASE 1: Launch 4 Parallel Audit Agents
-
- Launch these 4 agents simultaneously using `Agent` with `subagent_type=general-purpose`. Each agent's prompt **must** tell it to read `.claude/references/audit-pet-peeves.md` and explicitly check for each class of bug before reporting clean. The pet-peeves file is a living catalogue of drift patterns review bots have caught; it grows with each PR.
-
- #### Agent 1: Guide Content Accuracy
- Focus: `guide/workflow-guide.qmd`
- - All numeric claims match reality (skill count, agent count, rule count, hook count)
- - All file paths mentioned actually exist on disk
- - All skill/agent/rule names match actual directory names
- - Code examples are syntactically correct
- - Cross-references and anchors resolve
- - No stale counts from previous versions
-
- #### Agent 2: Executable Code Quality
- Focus: **all** executable code in the repo — `.claude/hooks/*.py`, `.claude/hooks/*.sh`, `scripts/*.py`, `scripts/*.sh`, `.claude/scripts/*.sh`. Not just `.claude/hooks/` — when PR #93 added new code under `scripts/`, the original narrow scope meant Copilot + Codex caught 5 bugs the audit missed.
-
- Hook-specific checks (Stop/PreToolUse/SessionStart protocols, `CLAUDE_PROJECT_DIR` usage, hash-length consistency) apply only to `.claude/hooks/`. Everything below applies to ALL executable code:
-
- - No remaining `/tmp/` usage in anything that manages state (should use `~/.claude/sessions/`)
- - Hash length consistency (`[:8]` across all hooks) [hooks only]
- - Proper error handling — **fail-open pattern** where the docstring promises it (top-level `try/except` with `sys.exit(0)`). Python `read_text()` must catch `UnicodeError` (not just `OSError`) if the script is promised fail-open for corrupt files. Bash `set -u` without `set -e` or explicit post-command checks does NOT catch command failures — verify.
- - **Docstring-claim ↔ implementation parity.** If a function's docstring describes "bidirectional parity" / "fail-open" / "exits 1 on X", the implementation must match. Common drift: one-directional implementation of a claimed-bidirectional contract; exit codes documented as one thing but returning another.
- - **Config-map entries point at live targets.** Keyword dicts, path maps, and rule registries should not contain dead entries (e.g. rule files that don't exist, fields the script doesn't actually read). Dead entries mislead maintainers.
- - JSON input/output correctness (stdin for input, stdout/stderr for output) [hooks only]
- - Exit code correctness. Two valid blocking protocols for Stop/PreToolUse hooks:
- (a) **exit 2 + reason on stderr** — legacy, still supported
- (b) **exit 0 + JSON `{"decision": "block", "reason": "..."}` on stdout** — modern; this is what `log-reminder.py` uses and it works correctly
- Non-blocking hooks always exit 0. PreCompact hooks MUST exit 0 (stdout is discarded by the harness — use stderr for diagnostics)
- - `from __future__ import annotations` for Python 3.8+ compatibility
- - Correct field names from hook input schema (`source` not `type` for SessionStart)
- - PreCompact hooks print to stderr (stdout is ignored)
-
- #### Agent 3: Skills and Rules Consistency
- Focus: `.claude/skills/*/SKILL.md` and `.claude/rules/*.md`
- - Valid YAML frontmatter in all files
- - No stale `disable-model-invocation: true`
- - `allowed-tools` values are sensible
- - **`allowed-tools` actually covers every tool the skill body invokes.** For every `Agent` spawn, `Bash` command, `Write`/`Edit` call mentioned in the skill's Steps / Phases / Workflow body, verify the tool appears in the `allowed-tools` array. Common miss: skill body says "spawn `agent-X` via the `Agent` tool with `context=fork`" but `Task` is absent from `allowed-tools` — runtime permission error or silent bypass. Caught this class of bug after Codex/Copilot flagged it on PR #92 (4 skills promised `Task` in their Post-Flight sections but 3 of 4 had no `Task` permission).
- - **Rule `paths:` scope matches skill implementation.** If rule X lists skill Y in `paths:`, verify skill Y actually implements the protocol rule X mandates. A rule claiming a skill follows a protocol is meaningless if the skill doesn't.
- - Rule `paths:` reference existing directories
- - No contradictions between rules
- - CLAUDE.md skills table matches actual skill directories 1:1
- - All templates referenced in `.claude/rules/*.md` and the guide (`guide/workflow-guide.qmd`) exist in `templates/`
-
- #### Agent 4: Cross-Document Consistency
- Focus: `README.md`, `docs/index.html`, `docs/workflow-guide.html`
- - All feature counts agree across all 3 documents
- - All links point to valid targets
- - License section matches LICENSE file
- - Directory tree matches actual structure
- - No stale counts from previous versions
-
- ### PHASE 2: Triage Findings
-
- Categorize each finding:
- - **Genuine bug**: Fix immediately
- - **False alarm**: Discard (document WHY it's false for future rounds)
-
- Common false alarms to watch for:
- - Quarto callout `## Title` inside `:::` divs — this is standard syntax, NOT a heading bug
- - `allowed-tools` linter warning — known linter bug (Claude Code issue #25380), field IS valid
- - Counts in old session logs — these are historical records, not user-facing docs
- - Counts in `CHANGELOG.md` under past version headings — those are snapshots; do NOT update
- - `log-reminder.py` outputting `{"decision": "block"}` with `sys.exit(0)` — this IS the modern Claude Code Stop-hook block protocol, NOT a bug
-
- **Count drift specifically: search for every phrasing variant.** A common failure mode is that `replace_all` on one phrasing (e.g., `"26 skills"`) misses sibling phrasings in the same repo. When checking counts, grep for ALL of:
- - `"N skills"`, `"N skill "` (with space)
- - `"N slash commands"`
- - `"N specialized"` (as in "N specialized agents")
- - `"template's N"` (informal count in prose)
- - Commas/conjunctions: `"skills,"` vs `"skills, and"` are treated as different strings by `replace_all`
- Verify zero matches for the OLD number across the whole tree before declaring clean.
-
- ### PHASE 3: Fix All Issues
-
- Apply fixes in parallel where possible. For each fix:
- 1. Read the file first (required by Edit tool)
- 2. Apply the fix
- 3. Verify the fix (grep for stale values, check syntax)
+ # Deep adversarial audit
- ### PHASE 4: Re-render if Guide Changed
+ A convergent alternative to slow round-by-round review. Instead of one reviewer finding one or two issues per pass, fan out many independent skeptics over the *whole* artifact at once, adjudicate what they find, fix everything confirmed, and re-verify. Modeled on the multi-agent methodology behind hard formal-proof efforts (diverse independent portfolio, adversarial throughout, concrete evidence only, synthesize-challenge-repeat).
- If `guide/workflow-guide.qmd` was modified:
- ```bash
- quarto render guide/workflow-guide.qmd
- cp guide/workflow-guide.html docs/workflow-guide.html
- ```
+ ## When to reach for this
+ - The artifact is dense enough that a single review keeps surfacing *new* issues each pass (the tell that round-by-round is the wrong tool).
+ - Correctness is the priority and the cost of a missed defect is high (a paper going to a top venue, a proof, a security-sensitive change, a migration).
+ - The user asked to "fix ALL of it", "be deeper", "converge faster", "100% rock solid".
- ### PHASE 5: Loop-until-dry or Declare Clean
+ Requires the user to have opted into multi-agent orchestration (they asked for a workflow / deep audit / to fan out agents, or ultracode is on). If they haven't, propose it and its rough cost first.
- After fixing, launch a fresh set of 4 agents to verify. This is the **loop-until-dry** primitive ([`orchestrator-protocol.md`](../../rules/orchestrator-protocol.md)):
- - **Converge** when a round surfaces **0 new genuine issues** (deduped on file+issue) — declare clean and report summary.
- - If new issues found → fix and loop again.
- - **Fallback cap: 5 loops** bounds a non-converging audit (prevents infinite cycling); a finding that survives rounds N and N+2 is escalated to the user rather than re-patched ([`summary-parity.md`](../../rules/summary-parity.md)).
+ ## The method
- ## Key Lessons from Past Audits
+ **1. Decompose (diverse portfolio).** Break the artifact into components by *idea*, not by section: each independent claim, lemma, estimator, subsystem, invariant. Add cross-cutting *failure-mode lenses* (see below). Aim for coverage such that every load-bearing claim is attacked by at least one agent that is looking straight at it. Don't tell the agents your favored reading — preserve independence so they don't all converge on the same attractive-but-wrong conclusion.
- These are real bugs found across 7 rounds — check for these specifically:
+ **2. Fan out adversarial finders (one per component).** Each finder is prompted to *refute*, defaulting to "there is a bug," and must ground every claim in the actual text/code (read it, don't paraphrase from memory). Hard rules, borrowed from what works:
+ - **Concrete findings only.** Every finding = exact location (file:line / label + quoted text) + one-sentence defect + a **failing case** (specific inputs/configuration → wrong output, or the exact missing hypothesis).
+ - **Reject** status reports, "looks fine", "this is standard/routine", vague optimism, and "the global step is straightforward."
+ - **A fix that re-imposes the same difficulty elsewhere, or assumes its own conclusion, is not a fix** — flag it.
+ - If, after genuinely attacking, nothing is found, the agent must state the *specific* attacks it ran and why each closed — not just "clean."
- | Bug Pattern | Where to Check | What Went Wrong |
- |-------------|---------------|-----------------|
- | Stale counts ("19 skills" → "21") | Guide, README, landing page | Added skills but didn't update all mentions |
- | Hook exit codes | All Python hooks | Exit 2 in PreCompact silently discards stdout |
- | Hook field names | post-compact-restore.py | SessionStart uses `source`, not `type` |
- | State in /tmp/ | All Python hooks | Should use `~/.claude/sessions/<hash>/` |
- | Hash length mismatch | All Python hooks | Some used `[:12]`, others `[:8]` |
- | Missing fail-open | Python hooks `__main__` | Unhandled exception → exit 1 → confusing behavior |
- | Python 3.10+ syntax | Type hints like `dict | None` | Need `from __future__ import annotations` |
- | Missing directories | quality_reports/specs/ | Referenced in rules but never created |
- | Always-on rule listing | Guide + README | meta-governance omitted from listings |
- | macOS-only commands | Skills, rules | `open` without `xdg-open` fallback |
- | Stale hook references | Rules, guide, CHANGELOG, settings.json | Removed hooks still mentioned somewhere |
+ **3. Adjudicate every finding (independent judge).** A separate judge re-opens each cited location and decides CONFIRMED / REFUTED / DOWNGRADED, skeptical of *both* the artifact and the finding. This kills false positives (misreads, hypotheses that are actually present elsewhere, failing cases that don't arise under the stated conditions) — the step that keeps the fix list honest.
- ## Output Format
+ **4. Synthesize.** Dedup by location, rank fatal > major > minor, and hand back one clean defect list. Nothing is accepted as an issue until it survives this.
- After each round, report:
+ **5. Fix all confirmed, then re-verify.** Apply every confirmed fix (you, in the main loop — fixing needs care and judgment). Then re-audit the touched spots and check that no fix created a new defect. Repeat waves until an audit pass comes back empty. Don't stop after the first wave.
- ```
- ## Round N Audit Results
+ ## Failure-mode lenses (adapt to domain)
+ Beyond per-component attacks, sweep these cross-cutting modes explicitly — they are where real defects hide:
+ - **Overclaim:** the headline/abstract claims more than the theorems/tests actually deliver.
+ - **Scope creep in a proof:** a *pointwise* result used where a *uniform* one is needed; a *both-correct* property stated unqualified; a special-case argument invoked generally.
+ - **Silent hypotheses:** a differentiability/density/continuity/boundedness/positivity condition used but never stated (in math), or an un-checked precondition/invariant (in code).
+ - **Edge cases:** atoms/ties, endpoints/unbounded support, empty/degenerate inputs, boundary of the parameter space.
+ - **Circularity:** an assumption that assumes its own conclusion; a result that cites itself; a citation that gives less than claimed (read the cited source).
+ - **Internal contradiction:** a definition/notation used two ways; a table cell contradicting a proposition; main text vs appendix disagreement; a dangling/wrong cross-reference.
- ### Issues Found: X genuine, Y false alarms
+ ## The fresh-eyes pass (final gate)
+ Every targeted wave inherits the blind spots of whoever wrote its prompts: focus hints, fix history, and expected failure modes all prime the auditors toward known territory. **After all targeted waves and fixes are done, run one cold audit with little to no context**: independent auditors given ONLY the artifact and a minimal instruction ("find concrete defects: location + failing case"), with no cluster assignments, no history, no special-focus lists. Diversify only the *entry point* (main-text-first as a journal referee would; appendix-first; tables/claims-first; a single deep dive of the auditor's own choosing). Adjudicate as usual. Clean fresh-eyes pass + clean targeted coverage + green mechanical battery is the closure standard; a fresh-eyes finding that targeted waves missed is also a diagnosis of the prompt set — add the missed failure mode to the lenses.
- | # | Severity | File | Issue | Status |
- |---|----------|------|-------|--------|
- | 1 | Critical | file.py:42 | Description | Fixed |
- | 2 | Medium | file.qmd:100 | Description | Fixed |
+ ## Full inventory — never sample
+ For a paper/proof artifact: **enumerate every formal statement first** (grep `\begin{theorem|proposition|lemma|corollary}` + labels) and assign each proof to a verifier — coverage must be 100% of load-bearing statements, not "a few proofs of the reviewer's choice." Sampling converges linearly and stochastically; inventories converge in one wave. Group tightly-coupled small lemmas into clusters; big proofs get their own verifier. Each verifier returns, besides findings, a **steps-verified list** and a **hypotheses ledger** (used-vs-stated; used-but-unstated is a finding).
- ### Verification
- - [ ] No stale counts (grep confirms)
- - [ ] All hooks have fail-open + future annotations
- - [ ] Guide renders successfully
- - [ ] docs/ updated
+ ## The mechanical battery (the highest-yield check)
+ Written arguments can read soundly while the object they define is wrong. For **every estimating equation, influence-function identity, identification claim, and population moment**, write an executable check that *computes the population object* on adversarial toy designs — truncation (censoring endpoint below the outcome endpoint), interior atoms, misspecified nuisances, boundary/overlap failure — and asserts the claimed centering/identity numerically (analytic or fine-grid/large-N with fixed seed). Keep the scripts as a permanent test directory in the repo with a README; rerun after any change to the corresponding formula. A 5-line population computation catches classes of defects (tail-renormalized roots, sign flips, mass-deficit weighting) that neither careful reading nor model consensus reliably finds.
- ### Result: [CLEAN | N issues remaining]
- ```
+ ## Fix hygiene
+ - **New math introduced by fixes is un-audited math**: every fix wave is followed by a verification wave over exactly the fixed spots before anything is declared closed.
+ - **In workflow synthesis, match findings to verdicts by INDEX** (require the judge to return verdicts in the findings' order), never by location string — judges paraphrase locations and silent drops follow. Verify the synthesized summary against the journal before acting on it.
- ## Findings are validated, not just written (v2.5)
+ ## Blocked routes are outcomes
+ If a component cannot be fixed under the stated assumptions, that is a *finding*, not a failure of the audit: report the exact remaining gap (the precise missing hypothesis or broken step) and the honest options (weaken the claim, add the hypothesis, restrict scope). Do not search for a favorable reading, and do not let an agent paper over a theorem-strength gap as "routine."
- This skill's reviewers emit findings under the machine-checked contract in
- [`finding-schema.json`](../../references/finding-schema.json). Reports are JSON **arrays**.
+ ## Orchestration
+ - Use a **Workflow** for the fan-out: `pipeline(components, finder, judge)` so each component's findings are judged the moment its finder returns (no barrier), then synthesize. Return the confirmed list; do the fixing yourself afterward.
+ - **Model division:** finders = the strong *execution/analysis* model (find and attack); judges = the strong *adjudication* model. Match to the local convention (here: Opus finds, Fable judges). Keep judges to one-per-component (adjudicating all that component's findings at once) to conserve the scarcer judging model.
+ - Set finder `effort` high; give each the exact labels/locations to read and its specific attack list.
+ - **Persist.** Don't return "best effort" or a list of why it's hard. Return the confirmed defects (and, once fixed, a clean re-audit) — or the single strongest remaining gap stated exactly.
- **Smoke-test the harness before spending review effort** — a run that fans out reviewers and
- then cannot write a valid report has wasted the whole pass:
+ ## Prompt skeletons
+ Finder: *"You are a HOSTILE referee auditing ONE component. Read the ACTUAL text at {locations}. Attack: {failure modes}. Return CONCRETE findings only (location + defect + failing case); no 'looks fine'/'routine'/vague. A fix that re-imposes the difficulty isn't a fix. If clean, list the specific attacks you ran and why each closed."*
- ```bash
- echo '[]' | python3 scripts/validate-findings.py
- ```
+ Judge: *"An adversarial referee returned these findings on component X. For each, open the cited location, verify against what the text ACTUALLY says and its proof, mark CONFIRMED/REFUTED/DOWNGRADED. Skeptical of both the artifact and the finding."*
- Then, before presenting any summary:
+ ## Shared doctrine lives in one place
- ```bash
- python3 scripts/validate-findings.py <report>.json # exit 0 required
- ```
+ Three things this audit depends on are **not** restated here, because they are the same rules
+ every other verification surface uses and a third copy would drift:
- What the contract forces, and why:
+ - **Seeded-fault calibration** — a check that has not caught a planted defect licenses nothing.
+ → [`/vaccinate`](../vaccinate/SKILL.md), and [`verification-ladder.md`](../../references/verification-ladder.md) rung 0.
+ - **Independence and correlated errors** — agreement between models is not confirmation; they
+ fail the same way. → [`verification-ladder.md`](../../references/verification-ladder.md) rung 3.
+ - **The five credibility questions** — evidence for one never clears another.
+ → [`verification-ladder.md`](../../references/verification-ladder.md) §6 and [`external-oracle-process.md`](../../references/external-oracle-process.md) §6.
- - **`rule`** — the documented rule or standard violated. A finding citing no rule is an
- opinion, and opinions do not gate a commit.
- - **`failing_case`** — a concrete configuration under which the claim breaks, or the exact
- missing hypothesis. *"This could be clearer"* does not validate.
- - **`id = sha1("<file>:<line>:<locus>:<lens>")`** — deterministic, so dedup across rounds is
- exact and the two-strikes rule is checkable rather than eyeballed.
- - **`mechanical`** — `true` only for fixes that cannot change a result (typo, cross-reference,
- formatting, label). **Never** for an estimand, assumption, specification, inference
- procedure, sample definition, or reporting language: those return to the researcher.
+ ## Auditing this repository itself
- Apply the **per-lens evidence burdens** and the **"does NOT count" filters** in
- [`orchestration-schemas.md` §7](../../references/orchestration-schemas.md) *before*
- verification, so known false alarms never reach the judge. The verifier pass is
- **refute-biased**: only `verdict: "confirmed"` findings ship; anything it cannot ground is
- dropped, not downgraded to a warning.
+ For the repo-infrastructure application — surface-sync, skill/agent/rule integrity, hook and
+ script review, doc-vs-reality drift — see
+ [`references/repo-infrastructure-audit.md`](references/repo-infrastructure-audit.md).
+ Start with `./scripts/backtest.sh`: the mechanical battery is already written, and an agent
+ should never hand-check what a script decides.