validate · diff

git:20260903.568e99d to git:20260903.10f0277

77 added, 163 removed. Audit A to A.

---
name: validate
- description: 'Freshly judge a finished change against its acceptance: PASS, FAIL, or NOT_PROVEN. Not for claim-vs-tree checks; that is reality-check. Triggers: "validate", "is this proven".'
+ description: 'Freshly judge a finished change against its acceptance: PASS, FAIL, or NOT_PROVEN. Not for claim-vs-tree checks; that is reality-check. Triggers: "validate", "is this proven", "check this change".'
practices:
- design-by-contract
- llm-eval-harness
- content-addressed-storage
hexagonal_role: driving-adapter
consumes:
- subject-manifest.v1
produces:
- subject-manifest.v1
- validation-result
- verdict.v2
context_rel:
- kind: customer-of
with: plan
- kind: customer-of
with: implement
skill_api_version: 1
user-invocable: true
metadata:
graph_root: true
tier: judgment
dependencies: []
capabilities: [compute_subject_identity, judge_acceptance, return_validation_result, persist_verdict]
effects: [write_verdict_artifact]
canonical_status: canonical
disposition: keep
output_contract: 'PASS | FAIL | NOT_PROVEN with criteria, evidence, checked/not_checked, identity, and freshness; optional schemas/verdict.v2.schema.json persistence'
---
# Validate
Independently judge one exact subject against the acceptance in its existing
bead or caller source, return one semantic result, and stop. Validate is the
- sole `verdict.v2` writer when persistence is requested. It never asks the model
- to reconstruct Plan or Candidate packets.
+ sole `verdict.v2` writer when persistence is requested. Before the verdict,
+ read `boundaries.md` in the rpi skill's `references` directory for the state
+ Validate leaves to the caller.
+ ## Prompt
+
+ ```text
+ Validate bead ag-1234 in this fresh context. Intent: the bead text and digest.
+ Subject: manifest.json from `python3 skills/validate/scripts/validate.py
+ manifest --root . --include cli/internal/gates`. Author context ctx-a1. Re-run `cd cli && go test
+ ./internal/gates/...`. Return PASS, FAIL, or NOT_PROVEN with evidence; stop.
+ ```
+
## Preconditions
- The subject is a nonempty implementation candidate: the manifest lists at
- least one entry, and `store-verdict` refuses an empty one. Plans, audits,
- reviews, and other control artifacts are not completion subjects unless the
+ least one entry. Plans, audits, and reviews are subjects only when the
caller explicitly requested document review.
- - The intent source is available as a caller-owned artifact or runtime-owned
+ - The intent source is a caller-owned artifact or a runtime-owned
content-addressed snapshot; its acceptance digest is derived automatically.
- - The subject manifest still matches the subject.
- - Author and validator context IDs are explicit.
- - Freshness is explicitly attested with `source: runtime | caller` and an
- attester identity.
-
- Missing, colliding, or unattested identities produce `NOT_PROVEN`. This is a
- declared trust fact, not cryptographic proof that contexts were isolated.
+ - Author and validator context IDs are explicit, and freshness is attested
+ with `source: runtime | caller` and an attester identity. Missing,
+ colliding, or unattested identities produce `NOT_PROVEN`: a declared trust
+ fact, not cryptographic proof of isolation.
## Cross-family fresh validator (default on risky surfaces)
- A second fresh validator from a different model family runs by default whenever
- the diff touches a risky surface:
-
- - `cli/internal/gates/**`
- - `scripts/check-*.sh`
- - `tests/**`
- - `skills/*/scripts/**`
- - `skills/cc-hooks/policies/**`
- - `lib/**`
- - `.github/workflows/**` and `scripts/security-gate.sh` (the gate scans the whole tree, so "anything it scans" is not a predicate)
-
- Outside that list the second validator is caller-elected and a single fresh
- validator remains the default shape.
-
- Dispatch is fixed by the runtime floor: never `claude -p` or `claude --print`,
- directly or indirectly.
-
- | Orchestrating runtime | Cross-family judge leg |
- |---|---|
- | Claude | a read-only `codex exec` judge leg |
- | Codex | a caller-selected interactive Claude session in an NTM pane |
-
- Probe the adapters at runtime through the `agent-native` model-dispatch recipe
- (`codex-exec` and/or `ntm`). The judge leg reads and judges; it never mutates
- the subject. Record author and validator `model_identity` in evidence refs and
- freshness attestation notes; do not change the `verdict.v2` schema.
-
- When no authorized live adapter is available, disclose the unsatisfied
- diversity request as `diversity_unsatisfied`. Off a risky surface that
- disclosure rides along with a same-model result. On a risky surface a
- single-family PASS is an unverified acceptance surface: the result is
- `NOT_PROVEN`, and same-family agreement never counts as convergence. A
- single-family FAIL stands as FAIL; a wrong subject needs no second judge.
+ A second fresh validator from a different model family runs by default when
+ the diff touches `cli/internal/gates/**`, `scripts/check-*.sh`, `tests/**`,
+ `skills/*/scripts/**`, `skills/cc-hooks/policies/**`, `lib/**`,
+ `.github/workflows/**`, or `scripts/security-gate.sh`; elsewhere it is
+ caller-elected. The runtime floor holds: never `claude -p` or
+ `claude --print`, directly or indirectly. Adapters and `model_identity`
+ recording: [references/mechanics.md](references/mechanics.md). With no
+ authorized live adapter, disclose `diversity_unsatisfied`: off a risky surface
+ it rides along with a same-model result; on a risky surface a single-family
+ PASS is `NOT_PROVEN`, and same-family agreement is not convergence. A
+ single-family FAIL stands.
## Mutating-check quarantine
- Before running any acceptance-listed command, classify it as read-only or
- subject-mutating. Regen scripts, sync scripts, formatters, and anything with
- `--force` are subject-mutating until proven otherwise. Never run a
- subject-mutating check against an uncommitted subject: on 2026-07-15,
- `scripts/test-ci-deterministic-gates.sh` regenerated `skills-codex/` from HEAD
- mid-validation and destroyed the uncommitted subject, forcing `NOT_PROVEN`
- (verdict `b6e759dd...cb6a`); only restoring the subject and revalidating in a
- fresh context produced the PASS (`e9b6cdb8...37b9`). If a mutating check is
- genuinely required by acceptance, run it against a disposable copy or a
- committed subject, never the judged working tree.
+ Classify every acceptance-listed command as read-only or subject-mutating
+ before running it: regen scripts, sync scripts, formatters, and anything with
+ `--force` are mutating until proven otherwise. Run a mutating check only
+ against a disposable copy or a committed subject, never the judged working
+ tree (the boundaries appendix records the regen that overwrote a subject).
## Scope disclosure
`not_checked` has exactly one meaning: **in-scope acceptance surface this
- validation did not verify**. PASS asserts that the whole declared acceptance
- surface was verified, so a PASS carries no `not_checked` entries; the helper
- refuses one and records a `validate.integrity` finding.
-
- That rule never pays for deleting an honest caveat, because every kind of scope
- limit has a home that survives inside a PASS:
-
- | Scope limit | Home | Example |
- |---|---|---|
- | A criterion proven by a bounded check | `criteria[].reason` on that criterion | "proven by the unit suite; the full integration matrix was not replayed" |
- | A declared non-goal or out-of-scope area | the intent source's non-goals, optionally restated as an evidence-backed boundary criterion in `criteria` | "`cli/**` is a declared non-goal; the diff proves it untouched" |
- | Residual risk or judgment caveat | the caller-facing report | "the migration path is untested against pre-3.0 stores" |
- | Acceptance that genuinely went unverified | `not_checked`, and the result is `NOT_PROVEN` rather than PASS | "criterion 3 needs hardware this context cannot reach" |
-
- Emptying `not_checked` to obtain PASS is a contract violation, not a
- workaround. If acceptance really went unverified, the honest result is
- `NOT_PROVEN`. If the entry was never acceptance in the first place, it belongs
- in one of the other homes, where it stays visible in the stored artifact
- instead of being deleted.
-
- ## Helper commands
-
- The helper ships beside this file. Invoke it through this skill's own
- directory rather than a checkout-relative path: `$SKILL_DIR` is the directory
- containing this `SKILL.md` — `skills/validate/` in a repository checkout,
- `.agents/skills/validate/` in an installed runtime.
-
- | Command | Required | Optional |
- |---|---|---|
- | `manifest` | `--root <dir>`, `--include <path>` (repeatable, at least one) | `--exclude <path-or-glob>` (repeatable), `--base-manifest <file>`, `--git-metadata-json <json>`, `--output <file>` |
- | `verify-manifest` | `--root <dir>`, `--manifest <file>` | `--base-manifest <file>` |
- | `snapshot-intent` | `--source <file>` (`-` reads stdin) | `--workspace <dir>`, `--intent-dir <dir>` |
- | `digest` | `<json-file>` positional | none |
- | `store-verdict` | `--draft`, `--intent-source`, `--subject-manifest`, `--author-context-id`, `--validator-context-id`, `--freshness-source <runtime\|caller>`, `--freshness-attester-id`, `--scope-result <PASS\|FAIL\|NOT_PROVEN>` | `--workspace <dir>`, `--verdict-dir <dir>` |
-
- ```sh
- python3 "$SKILL_DIR/scripts/validate.py" manifest \
- --root . --include skills/validate --exclude '**/*.log' --output manifest.json
- ```
+ validation did not verify**. PASS asserts the whole declared acceptance
+ surface was verified, so a PASS carries no `not_checked` entries; every other
+ scope limit has a home that survives inside a PASS: a bounded proof in
+ `criteria[].reason`, a declared non-goal in the intent source, residual risk
+ in the report (table in the mechanics reference). Emptying `not_checked` to
+ obtain PASS is a contract violation: unverified acceptance makes the honest
+ result `NOT_PROVEN`, and an entry that was never acceptance moves to its home
+ and stays visible.
## Workflow
- 1. Recompute and compare `subject-manifest.v1` with the `manifest` command
- above (`--root` plus at least one `--include`). The helper uses only
- filesystem content; Git commit/tree IDs are optional metadata. Derive the
- manifest at the start of validation and re-derive it at the end; any
- mismatch between the two is subject mutation and returns `NOT_PROVEN`.
- 2. Confirm the intent-source digest has not changed since implementation. If
- the subject changed or complete changed-path coverage cannot be derived,
- return `NOT_PROVEN`.
- 3. Adjudicate the actual diff, not a declared path list: compare
- runtime-derived actual changed paths against the intent's scope classes. A
- proven out-of-scope path returns `FAIL`; incomplete scope evidence returns
- `NOT_PROVEN`.
- 4. Inspect the exact subject and factual evidence. Reported exit codes are
- claims, not evidence: re-execute the claimed proofs that bear on acceptance
- (see the freshness rules below for when a digest-bound receipt suffices).
- If the subject changes a test, gate, fixture, golden, tolerance, suppression,
- or acceptance source, determine whether the original intent requires that
- change and whether green came from implemented behavior rather than a
- weakened oracle. Green obtained by weakening acceptance is `FAIL`, not
- evidence of completion.
- Judge every acceptance criterion and record criterion-level results,
- findings, evidence references, `checked`, and any acceptance surface that
- went unverified in `not_checked` (see Scope disclosure).
+ 1. Derive `subject-manifest.v1` with the helper's `manifest` command (flags
+ in the mechanics reference) at the start and again at the end; any
+ mismatch is subject mutation and returns `NOT_PROVEN`.
+ 2. Confirm the intent-source digest is unchanged since implementation, every
+ cited evidence digest matches the artifact it names, and complete
+ changed-path coverage can be derived; otherwise `NOT_PROVEN`.
+ 3. Adjudicate the actual diff: runtime-derived changed paths against the
+ intent's scope classes. A proven out-of-scope path is `FAIL`; incomplete
+ scope evidence is `NOT_PROVEN`.
+ 4. Inspect the exact subject and evidence. Reported exit codes are claims:
+ re-execute the proofs that bear on acceptance. A changed test, gate,
+ fixture, golden, tolerance, suppression, or acceptance source must be
+ required by the original intent, with green coming from implemented
+ behavior; green obtained by weakening acceptance is `FAIL`. Judge every
+ acceptance criterion against its own evidence reference; a criterion with
+ no evidence of its own is unverified, not passed.
5. Choose exactly one semantic result: `PASS`, `FAIL`, or `NOT_PROVEN`. Return
- it with criterion results, findings, evidence references, `checked`,
- `not_checked`, the acceptance and subject identities, distinct author and
- validator context IDs, and the freshness attestation. PASS requires distinct
- identities, explicit freshness, nonempty checked scope, top-level evidence,
- evidence for every criterion, and an empty `not_checked`; route bounded
- proofs, declared non-goals, and residual risk to the homes named in Scope
- disclosure rather than deleting them or downgrading a proven result.
+ it with criterion-level results, findings, evidence references, `checked`,
+ `not_checked`, both identities, both context IDs, and the freshness
+ attestation. PASS requires distinct identities, explicit freshness,
+ nonempty checked scope, nonempty top-level evidence, evidence for every
+ criterion, and an empty `not_checked`.
6. Only when the caller requests machine-readable evidence or a declared
downstream consumer requires it, persist canonical `verdict.v2` with the
- helper's
- `store-verdict --draft <draft.json> --intent-source <resolved-intent>
- --subject-manifest <manifest.json> --author-context-id <id>
- --validator-context-id <id> --freshness-source <runtime|caller>
- --freshness-attester-id <id> --scope-result <PASS|FAIL|NOT_PROVEN>`. The
- helper snapshots the exact resolved intent under
- `<workspace>/.agents/ao/intents/sha256/<digest>.intent`, then computes and
- injects intent and subject digests plus author, validator, and freshness
- facts. Identity and changed-path facts come from runtime-derived inputs and
- receipts, not model transcription. Storage defaults to
- `<workspace>/.agents/ao/verdicts/sha256/<digest>.json`; callers may provide
- `verdict_dir`.
- 7. Return the semantic result and, when persisted, the artifact path and digest.
- Stop.
-
- Routine rounds keep the bounded, receipt-driven freshness contract below, and
- the repository's full literal CI command set, as quoted in `AGENTS.md`, is
- required once, on the final integrated subject.
-
- The digest is SHA-256 over canonical JSON with `artifact_digest` omitted. Writes
- use a same-directory temporary file, flush, fsync, and atomic rename. Identical
- existing content is idempotent success; conflicting content is an integrity
- failure represented by `NOT_PROVEN`.
-
- ## Freshness without duplication
+ helper's `store-verdict` (mechanics reference), then return the artifact
+ path and digest with the result. Stop.
- Fresh validation means independent judgment over the exact subject. It does not
- require mechanically replaying every author command. Verify intent identity,
- scope, evidence digests, and every acceptance criterion; independently rerun
- the risk-critical, uncertain, or insufficiently evidenced checks. A
- digest-bound deterministic receipt may prove routine facts. Replay an expensive
- full suite only when acceptance requires that result or the supplied receipt
- cannot establish it.
+ Fresh validation is independent judgment over the exact subject, not a replay
+ of every author command: rerun the risk-critical, uncertain, or thinly
+ evidenced checks; a digest-bound deterministic receipt may prove routine
+ facts; replay an expensive full suite only when acceptance requires it. The
+ repository's full literal CI command set, as quoted in `AGENTS.md`, runs
+ once, on the final integrated subject.
## It's working if
- Observable in the trace, without reading the prose — and the rubric a fresh
+ Observable in the trace, without reading the prose, and the rubric a fresh
independent judge scores this skill against:
- A criterion whose evidence is a justification rather than a proof is named,
and the result is `NOT_PROVEN` rather than `PASS`.
- Green obtained by widening a tolerance, skipping a case, or re-baselining a
budget is reported as `FAIL`, never as completion.
- Every scope limit is placed in one of the Scope-disclosure homes; none was
deleted to reach `PASS`.
- - The subject manifest is derived twice — at the start and at the end — and the
+ - The subject manifest is derived twice, at the start and at the end, and the
two are compared.
## Boundary
- Validate emits no WARN, confidence, disposition, briefing learning, owner,
- next action, repair, retry, replan, helper, escalation, tracker, Git, release,
- closure, or delivery state. Generic provenance may record a verdict later, but
- ledger availability cannot change its validity.
+ Validate emits no next action, repair, retry, replan, or delivery state (full
+ list in the boundaries reference); ledger availability cannot change a
+ verdict's validity.