ai-agent-testing ยท diff
git:20260829.30ae153 to git:20260829.175e380
57 added, 1 removed. Audit A to A.
---
name: ai-agent-testing
description: Use this skill when you need to test AI agent goals, state, planning, recovery, and safety boundaries; triggers include ai agent testing.
---
+
# AI Agent Testing
- Read `prompts/ai-agent-testing.md`; test agent behavior and recovery without assuming authorization for external side effects.
+
+ ## When to Use
+
+ - Use this skill when you need to systematically verify agent planning, memory, tool use, recovery, safety boundaries, and task completion quality.
+ - Use it to review an existing plan, result, or evidence set and produce actionable improvements.
+ - Use it when context is incomplete but a bounded first pass is still valuable.
+
+ ## Output Format Options
+
+ - Default to Markdown for review, execution, and incremental refinement.
+ - When the user requests tables, CSV, JSON, or ticket fields, preserve risk, evidence, priority, and boundary information.
+ - For machine-consumed output, confirm the schema, enums, and required fields first.
+
+ ## How to Use
+
+ 1. Read and follow `prompts/ai-agent-testing.md`, including its input contract, execution rules, minimum coverage, and output order.
+ 2. Add only context that changes the decision: scope, environment, version, constraints, evidence, and success criteria.
+ 3. Audit the input, then separate confirmed facts, working assumptions, and open questions.
+ 4. Rank by risk and evidence strength, and produce an artifact that can be executed or reviewed directly.
+ 5. If information is missing, deliver a bounded first pass and state which conclusions remain unsupported.
+
+ ## Reference Files
+
+ - Always read `prompts/ai-agent-testing.md`; it is the complete execution specification for this skill.
+ - For evaluation or regression, read `evals/eval.yaml` and the relevant cases under `evals/cases/`.
+ - Load `references/`, `examples/`, `scripts/`, or `output-formats.md` only when those directories exist and the task needs them.
+
+ ## Core Constraints
+
+ - evaluate both final results and execution traces
+ - use repeated trials for stochastic behavior
+ - never treat one success as reliability evidence
+ - Never invent system behavior, fields, data, metrics, or root causes absent from the evidence.
+ - Link important conclusions to evidence; mark unsupported conclusions as hypotheses with a verification method.
+ - Explain priority using business impact, likelihood, or detectability.
+
+ ## Delivery Checklist
+
+ - [ ] Covered: task completion, planning quality, multi-step consistency, tool use, memory contamination, recovery, authorization boundaries, cost and latency.
+ - [ ] Separated facts, assumptions, gaps, and recommendations.
+ - [ ] Gave high-risk items a priority, evidence basis, owner or next action.
+ - [ ] Defined verifiable decision criteria instead of generic advice.
+ - [ ] Performed no unauthorized production writes or destructive actions.
+
+ ## Common Pitfalls
+
+ - Listing checks without preconditions, expected outcomes, or evidence.
+ - Marking everything high priority and avoiding tradeoffs.
+ - Substituting tool names or generic theory for domain reasoning.
+ - Refusing incomplete input, or pretending incomplete evidence supports certainty.
+
+ ## Best Practices
+
+ - Start with paths most likely to cause business loss, safety issues, or release blockage.
+ - Reduce uncertainty through the smallest verifiable experiment and record reproduction conditions.
+ - Make the artifact executable and independently reviewable by another engineer.