analyse-agentic-system · git:20260905.6241948 · 2026-09-05 · sha256 1903f84c57261dcd
analyse-agentic-system git:20260905.6241948A
Immutable. This exact content is served forever at /api/v1/blob/1903f84c57261dcd.
--- name: analyse-agentic-system description: "Use when asked to analyse, review, or refresh an external agent runtime, orchestration system, agent operating layer, agent memory/knowledge/context-engineering system, or narrower model-dependent operational mechanism from inspectable sources." type: kb/types/instruction.md user-invocable: true argument-hint: "<system identifier> plus source input (repository, checkout, snapshot/bundle, or documents) and optional public review path" allowed-tools: Read, Write, Grep, Glob, Bash, Task context: fork model: opus --- # Analyse an Agentic System Analyse one external agentic system at one frozen evidence boundary. Identify what the system wires, where its responsibilities end, and what its memory/context and epistemic routes support. Run a runtime baseline and both mandatory lenses, then write one exact result and publish a compact generated review. If the target is itself a memory, knowledge, or context-engineering system, also publish its legacy collection review from the same sources. Invocation authorizes the run directory under `kb/reports/state/agentic-system-analysis/`, one generated review under `kb/agentic-systems/reviews/`, its identical exact-result copy under `kb/reports/retained/agentic-system-analysis/<run-id>/result.md`, and, when applicable, one review under `kb/agent-memory-systems/` plus its semantic-QA state. It does not authorize changes to source worktrees, auxiliary indexes or surveys, transfer scans, landscape synthesis, other retained reports, or Git staging and commits. ## Failure rule Keep a correctable pre-publication failure in `running` state, fix the candidate or result, and repeat the failed check. Set `run-status: failed` with one concise reason only when abandoning the run or when a publication error leaves public state uncertain. Do not resume a failed run or maintain a phase ledger, packet, correction log, retry log, or validation receipt. Use a new run ID. Temporary candidates are non-canonical and may be overwritten or removed by the run owner. ## Steps ### 1. Open the run and resolve output paths 1. Allocate `AAS-<YYYY-MM-DD>-<system-slug>-<nn>`. Create `kb/reports/state/agentic-system-analysis/<run-id>/run-state.md` from the `agentic-system-analysis-run-state` template with `run-status: running`. The exact result path is always `<run-id>/result.md`. Run `commonplace-validate <run-state-path>` immediately. 2. Derive the public review path as `kb/agentic-systems/reviews/<system-slug>.md` unless the caller supplied one. A caller-supplied path must also be directly under `reviews/`. Inspect only an incumbent's frontmatter and source identity. Reuse the path only for a review generated by this skill from the same source identity. Otherwise choose an owner- or source-qualified slug. Never overwrite a hand-authored or different-source artifact. 3. Record separately any caller-authorized auxiliary paths and any separately commissioned transfer scan. Automatic review publication does not authorize those operations. 4. Confirm the target is in scope: an agent runtime, orchestration framework, agent operating layer, memory/knowledge/context-engineering system, or a narrower mechanism whose operation depends on model calls it issues or serves. An MCP server, tool, or returning computation may qualify without owning the enclosing runtime. If the target is outside this boundary, write and validate an `out-of-scope` result, complete the run without a public review, and stop. 5. Classify the target as an `enclosing runtime`, `embedded inner runtime`, `runtime client`, `returning computation`, `workflow`, `extension or tool mechanism`, `builder or improvement plane`, `host integration`, `memory/knowledge/context-engineering system`, or another defined class. State functional inclusions, exclusions, external dependencies, and one boundary kind: `whole-system`, `subsystem-only`, or `complete artifact, partial loop`. Do not assign responsibilities owned by an excluded host to the selected target. If no coherent boundary or reachable source can be established, write a typed `blocked` result with every required section and explicit unreached dispositions. Complete the run without publishing a review. ### 2. Freeze and inspect sources once 1. Before inspection, record a compact source allowlist: the exact repositories, captures, documents, and time boundary that may supply evidence. A supplied repository reference authorizes creating its missing ignored checkout and fetching the objects needed for the selected revision. It does not authorize changing an existing worktree, switching branches, merging, pulling, or resetting. 2. For GitHub, normalize the repository identity and use `related-systems/<owner>--<repo>/`. Require `git check-ignore -q related-systems` before creating it, verify an existing checkout's origin, and resolve the selected revision to a full commit. Inspect only with commit-addressed `git --no-replace-objects -C <absolute-root> ls-tree`, `show`, and `grep`; never read evidence from the worktree. 3. Turn every non-Git source set into one immutable capture or bundle with a stable identity, version or capture label, absolute path, and SHA-256. Do not analyse a moving live page as though it were frozen. 4. Put the Git commit or capture identity in `run-state.source` while the state remains `running`. Build one `SRC-*` register in the result with evidence layer, inspected scope, citation anchors, and access gaps. Keep implementation, doctrine/design, reported operation, observed runs, and causal experiments distinct. 5. Use a recorded search boundary only for a load-bearing absence claim. Name the searched roots or files, query, and revision; a casual search miss is a limitation, not an `ABS-*` record. 6. Select files and line ranges before reading content. Budget the aggregate output of parallel reads against the tool wrapper's delivery limit; an output cap alone does not bound the inspection. Treat truncated output as non-evidence. Narrow and repeat the read before citing it. For Git, cite the `SRC-*` ID plus a full commit-relative path and line range such as `packages/runtime/src/agent-run.ts:595-641`; use the full commit in GitHub blob links. If an exact result or review projection uses a quote-anchored blockquote, end it with either a matching full-commit GitHub blob URL or `` `commit-relative/path` @ `full-commit` ``. Publication checks the quote against the recorded Git blob, never the worktree. For a capture it checks the quote against the frozen captured bytes. This establishes occurrence, not whether the quote supports the surrounding claim. The source pin is an evidence boundary, not a recovery protocol. If it changes or cannot be verified, fail the run and start another one. ### 3. Use one vocabulary and one record set Use the result type's conclusion statuses exactly: `absent`, `inapplicable`, `uninspected`, `claimed`, `afforded`, `wired`, `observed`, and `causally supported`. Never upgrade context presence to activation, a claim to an affordance, an affordance to wiring, wiring to observation, observation to causality, or curation to warrant. Every negative or uncertain finding names the inspected boundary and conclusion prevented. Keep these distinctions: - **Memory read-back** means material accumulated or changed through use affects a later consumer invocation. Static shipped material and ordinary current-run state are not read-back. - **Activation** requires evidence that delivered material changed behavior. - **Behavioral authority** records consumer, channel, force, and horizon. Epistemic and operational authority remain separate. - **Guarantee strength** is separate from evidence status: invariant, protocol, policy, best effort, deployment guarantee, or no claimed guarantee. Describe every external mechanism in source-native terms before mapping it to Commonplace ontology. Explain the fit and mark partial or unresolved mappings. Do not turn omission of an open-ended mechanism into evidence of absence. Maintain one canonical register: `SRC-*` sources, `CMP-*` components, `OBJ-*` operative objects, `RTE-*` routes, `CLM-*` claims, `ABS-*` evidenced absences, and `BAP-*` behavioral-authority paths. The orchestrator owns IDs and generic identity. A lens annotates existing IDs and proposes new records under local tags that disappear when the orchestrator registers or merges them. An `uninspected` gap is a limitation, not an `ABS-*` record. Allocate canonical IDs monotonically. Never reuse an ID after merging or rejecting its record; gaps are harmless. An ID shared with a worker is canonical before final integration. Amend its evidence or status without changing its referent. Splitting a combined record requires new IDs for the parts and an explicit superseded disposition on the original record; do not assign its ID to one part. Use local labels for provisional seeds. When workers are useful and available, give each one the run ID, frozen source identity, reviewed boundary, source register, canonical register, lens scope, and sparse-return contract. A worker may inspect only that boundary, does not publish or delegate, and returns annotations, proposals, corrections, evidenced absences, and limitations. The orchestrator checks that every cited ID and source belongs to the supplied registers before merging. Discard and rerun a stale, malformed, or mismatched lens return; do not maintain packet or correction history. Running both lenses sequentially in the current context is also valid. ### 4. Run and challenge the runtime baseline 1. Begin with consequential claimed work and shipped entry paths. Trace one ordinary invocation end to end: principal, identity, context, state, model call, effects, runtime-client controls, coordination, terminal result, and retained or lost state. 2. Enumerate materially equivalent alternate paths before judging a guarantee: direct model calls, provider-native tools, host callbacks, shell access, extension code, subprocesses or remote workers, manual graph control, and durable variants where present. A guarantee covers only the paths its enforcement point covers. 3. Trace the smallest warranted set of forcing cases, ordinarily two to four for a full code-grounded pass. Prefer static inspection. Before any dynamic check, record the result type's execution-preflight fields and verify tools, packages, services, credentials, configuration, and authority. A check that never reaches the target remains `not run` and supports no negative finding. If no dynamic check is warranted, briefly list the checks considered and why static evidence was sufficient; keep the result's required disposition `no dynamic check planned`. 4. Record an executed check as a `SRC-*` probe evidence capsule. Use `causally supported` only for an actual intervention and comparison whose design supports the attribution. Exact output must remain inspectable in the one-file result. 5. For each material route record trigger, next-step owner, decision policy and form, context, state, executor and effect boundary, persistence, return, recovery, and terminal output. A load-bearing guarantee also names its owner, enforcement point, strength, covered and alternate paths, and required external contract. 6. Audit every `RTE-*` route for immediate return, later read-back, delegated visibility, selection predicate, invalidation or expiry, activation or effect, and evidence limits. Use explicit inapplicable or uninspected reasons instead of empty fields. 7. Distinguish the capability surface, current grant set, and deployed isolation envelope. Inspect permissions, approval, delegation, dynamic extension, reliability, observability, providers, packaging, and performance only where they change claimed work, a control path, evidence strength, or a lens result. ### 5. Run both lenses For memory/context and epistemic, first record trigger evidence, inspected boundary, pointed-to routes and objects, warranted `brief` or `full` depth, and rationale. Both lenses always run. A brief result still states what was inventoried, what was found, and which conclusions its thin evidence prevents. The memory/context lens looks for retained material, read-back into a later invocation, selection and framing, activation evidence, scope and invalidation, and behavioral authority. Separately detect whether the selected target's primary offered work is retaining, transforming, organizing, selecting, or delivering material for later agent work. Record the detection and evidence; publication of the legacy review happens only after the exact result validates. Invoke [`analyse-external-system-epistemic-architecture.md`](../analyse-external-system-epistemic-architecture.md) for the epistemic lens. Pass the frozen boundary, registers, statuses, scoping record, and classify-only routes. Require a sparse overlay on canonical IDs. Keep that procedure's architectural status and observed candidate state in their own vocabulary; never translate `implemented` into this workflow's conclusion-status field. ### 6. Reconcile and synthesize Resolve proposed records into canonical IDs, attach corrections and amendments to the affected records, preserve anchored conflicts, and report independent convergence only when the lenses reached it independently. Recheck shared-route ownership and the memory-review detection. If reconciliation exposes stale or unsupported lens work, rerun that lens before continuing. Write a system-organized synthesis: evidence basis and boundary, architectural characterization and claimed work, runtime map, discriminating mechanisms, scenario-relative assessment, limitations, and evidence or system changes that would alter the assessment. Do not concatenate lens reports or add a product ranking, generic adoption advice, system-wide epistemic grade, Commonplace delta, transfer recommendation, or universal maturity model. ### 7. Write and validate the exact result Write `<run-id>/result.md` using `kb/types/agentic-system-analysis-result.md`. Every disposition keeps all required headings. Its Run identity names the run state, generated review disposition, and legacy review disposition. Put probe evidence inline. Normalize the established memory/context findings into the result type's `memory-comparison` frontmatter. Name the memory boundary and fill every axis with its assessment, evidence basis, values, canonical records, and rationale. Keep uncertain, uninspected, inapplicable, and evidenced-absent cases explicit. Use the weakest evidence basis across an aggregated value set; never infer comparison values from legacy reviews or a missing tag or section. After reconciliation, check the integrated result against the type's comparison rules, not just the separate lens returns: - Match the profile scope to canonical objects, route branches and lens scope. Account for included alternatives and opaque parts in each known aggregate; distinguish excluded branches explicitly. - Check every scoped trace-fed write against the learning criterion, including compaction. Carry each qualifying route into the dependent assessments with its own source, task horizon, timing and form. - For each push signal, identify the consumer, trigger, selector input and selected retained part. Separate requested reads from automatic selection; a named object alone does not establish identifier-based push. - Check that source/status amendments still concern the same canonical IDs, and that overlays cite rather than duplicate their generic records. Record the checked routes and material dispositions in the existing Semantic verification section. A known assessment unsupported by its records blocks publication; properly scoped explicit uncertainty does not. Structural validation or the absence of a legacy review job does not perform this check. Run `commonplace-validate --full <result-path>` and verify every source anchor, canonical ID, evidence status, boundary, lens output, limitation, and blocker. Correct deterministic formatting errors before continuing. An unresolved evidence or semantic failure blocks publication. Correct it while the run stays `running`, or abandon the run under the failure rule. The result's `complete` disposition means its analysis content is complete; it does not claim that the review projections have been published. Do not persist a JSON validation receipt or review-job details in the result. ### 8. Publish validated candidates Skip publication for a blocked or out-of-scope result. For a complete result: 1. Generate the compact whole-system review solely from the validated result and its primary-source anchors. Write it first as a temporary candidate in the run directory. Use `kb/types/note.md` and exact frontmatter fields `generated-by: analyse-agentic-system`, `analysis-run`, `source-identity`, `reviewed-revision`, `analysis-result`, and `analysis-result-sha256`. The last two fields name `kb/reports/retained/agentic-system-analysis/<run-id>/result.md` and the SHA-256 of the validated exact result. Publication retains those identical bytes; do not draft a separate retained report or rewrite the result for the matrix. 2. If memory-review detection applies, invoke [Write an agent memory system review](../write-agent-memory-system-review/SKILL.md) with the frozen source register, reviewed revision, selected destination, and a candidate path inside this run directory. The child writes only that candidate. Require its workflow fields `generated-by: analyse-agentic-system`, `analysis-run`, `source-identity`, and `reviewed-revision`. 3. Run `commonplace-agentic-analysis-publication prepare` with the run state, generated candidate and destination, plus the legacy candidate, destination, and model partition when applicable. This validates candidate bytes as their intended public paths, verifies source anchors and any quote-anchored blocks against the frozen source, verifies workflow identity, checks incumbents, and creates one semantic review job for the legacy candidate. It changes no public artifact. 4. When prepare returns a review job, dispatch its prompt through the normal review worker path and run `commonplace-finalize-review-job` for that job. All applicable semantic gates must pass. Do not copy an incumbent baseline or substitute an informal review. 5. Run `commonplace-agentic-analysis-publication publish` with the same arguments. It requires current semantic pass baselines for the candidate revision, validates the prospective complete run state, replaces all review projections, retains the exact result, and writes the complete run state last. It rolls back ordinary in-process write or validation failures. A crash or power loss during the short replace sequence may still leave partial public writes; inspect them, mark the run `failed`, and use a new run ID. The publish command records the exact result and review hashes and the legacy review model partition in the run state. Candidate cleanup after success is best effort. A cleanup warning does not undo completion. Never patch generated prose independently of its source boundary and method. Never stage or commit unless the caller separately requested it. After a main review changes, report the comparison outputs under `kb/agentic-systems/comparisons/` stale unless rebuilt and validated under separate authority. The matrix, table, and numerical-analysis scripts read the retained result directly. Repeated `--review` arguments select a bounded corpus; without them every generated main review must meet the input contract. The old `kb/agent-memory-systems/systems.csv` and `systems-table.md` are historical snapshots and are not rebuilt by these scripts. Report a prior current landscape synthesis as historical unless it was refreshed under separate authority. Authorization alone does not establish that an operation completed. The new matrix records both evidence tiers; numerical claims use code-grounded rows, with doc-grounded findings kept separate. ### 9. Run an optional transfer scan after completion A transfer scan is separate, interest-conditioned state. Run [`scan-agentic-system-transfer`](../scan-agentic-system-transfer/SKILL.md) only when separately commissioned and only after the complete run state validates. Pass `result.md`, its SHA-256, the sibling `run-state.md` path, the interest brief, and permitted Commonplace read and output scope. The scan verifies the completed run and reads the exact result directly; a legacy review or compact public review is not a substitute. The scan never edits the analysis, published reviews, or comparison corpus. If it exposes an analysis defect, fail that conclusion and rerun the analysis before scanning again. ### 10. Report Run `commonplace-agentic-analysis-handoff <run-state-path>` and include its Markdown output unchanged in the final response. It validates the run state and current output bytes before rendering. Append the downstream freshness dispositions required by step 8. For a separately commissioned transfer scan, also return its output path and material findings, or its full output when no file was written. Report any blocker that prevented the scan from completing. These session outcomes supplement the checked handoff; they do not require additional run-state fields. A failed run reports its failure reason and does not use the handoff command. ## Verify - One run ID, frozen source boundary, and source register govern every finding. - Git reads use the full recorded commit; captures match their SHA-256; no truncated output supports a claim. - The runtime baseline covers ordinary, material alternate, and warranted forcing routes at the target's actual responsibility horizon. - Both lenses and both scoping records exist; thin evidence produces a bounded brief result, not a skipped lens. - Source-native mechanisms remain visible beneath Commonplace mappings, and no conclusion status is upgraded. - The exact result validates before publication. - Each public review has the SHA-256 and workflow identity recorded by the complete run state; a required legacy review also has current semantic pass baselines for the recorded model partition. - Correctable pre-publication failures keep the run `running`. A failed run was abandoned or has uncertain public state; its replacement is a new run. --- - [Agent-runtime analysis should separate scheduling, context assembly, and external state](../../notes/agent-runtime-analysis-should-separate-scheduling-context-state.md) — rests-on: the causal runtime responsibilities in step 4 - [Agent orchestration occupies a multi-dimensional design space](../../notes/agent-orchestration-occupies-a-multi-dimensional-design-space.md) — rests-on: why the runtime inventory remains open - [Agent memory is a crosscutting concern, not a separable niche](../../notes/agent-memory-is-a-crosscutting-concern-not-a-separable-niche.md) — rests-on: why memory is a mandatory lens - [Knowledge storage does not imply contextual activation](../../notes/knowledge-storage-does-not-imply-contextual-activation.md) — rests-on: the retention, read-back, presence, and activation distinctions - [Behavioral authority](../../notes/definitions/behavioral-authority.md) — rests-on: the consumer, channel, force, and horizon record