git:20260505.3801ee2 to git:20260515.919baf6

56 added, 18 removed. Audit A to A.

---
name: prompt-evaluation-runner
description: Use when evaluating prompts, LLM outputs, red-team suites, or model behavior with local eval configs and safe provider/cost controls.
---
# Prompt Evaluation Runner
## When to use
- Use this skill when you need to evaluate an LLM app, test a prompt, or run red-teaming/vulnerability scans against a target model or application.
+ Use when you need to evaluate an LLM app, test a prompt systematically, or run red-team/vulnerability scans against a target model or application.
## Requirements / Checks
- 1. Check if an evaluation tool is defined in project deps, scripts, lockfiles, or local toolchain.
- 2. Do not run external commands like `npx Prompt evaluation@latest` directly.
- 3. If missing, ask before adding a local dev dependency or using an ephemeral runner.
- 4. Confirm expected cost, provider, API keys, and network target before execution.
+ 1. Check if an evaluation tool is defined in project deps, scripts, lockfiles, or local toolchain (e.g., `promptfoo`, `evals`, `braintrust`).
+ 2. Do not run unvetted remote runners without checking the project's toolchain first (e.g., avoid `npx promptfoo@latest` if `promptfoo` is already installed locally).
+ 3. If no runner exists, ask before adding a dev dependency or using an ephemeral runner.
+ 4. Confirm expected cost, provider, API keys, and network target before any execution.
## Workflow
- 1. **Define risk**: State target behavior, failure mode, provider(s), and budget limits.
- 2. **Choose assertions**: Prefer deterministic checks first: exact/contains/regex/JSON/schema/javascript/python/cost/latency.
- 3. **Use model graders sparingly**: Pin grader provider/model and explain cost/non-determinism.
- 4. **Configure minimally**: Keep config to description, env refs, prompts, providers, default assertions, and tests.
- 5. **Handle env safely**: Use templated env references such as `{{env.NAME}}`; never hardcode keys.
- 6. **Execute locally**: Run smallest suite first. Ask before long, paid, red-team, or production-targeted runs.
- 7. **Analyze failures**: Separate prompt failures, provider variance, flaky graders, bad fixtures, and config mistakes.
+ 1. **Define risk** — state target behavior, failure mode, provider(s), and budget limits before writing any config.
+
+ 2. **Choose assertions** — prefer deterministic checks first:
+
+ | Assertion type | When to use |
+ |---|---|
+ | `contains` / `not-contains` | Output must include/exclude specific text |
+ | `regex` | Structured output pattern (e.g., JSON key present) |
+ | `json-schema` | Output must conform to a schema |
+ | `cost` | Must stay under a token/dollar budget |
+ | `latency` | Must respond within N ms |
+ | `javascript` / `python` | Custom logic when simpler types don't fit |
+ | Model grader | Last resort — only for subjective quality checks |
+
+ 3. **Use model graders sparingly** — pin the grader model and provider explicitly; document the cost and non-determinism risk.
+
+ 4. **Minimal config structure**:
+ ```yaml
+ description: "Test that the summarizer stays under 200 words"
+ providers:
+ - id: openai:gpt-4o-mini
+ config:
+ temperature: 0
+ prompts:
+ - "Summarize: {{input}}"
+ defaultTest:
+ assert:
+ - type: javascript
+ value: output.split(' ').length < 200
+ tests:
+ - vars:
+ input: "{{env.TEST_DOCUMENT}}"
+ ```
+
+ 5. **Handle env safely** — use `{{env.VAR_NAME}}` for all secrets and inputs. Never hardcode API keys or sensitive data in config files.
+
+ 6. **Execute locally** — run the smallest suite first. Ask before running long, paid, red-team, or production-targeted suites.
+
+ 7. **Analyze failures** — classify before fixing:
+ - Prompt failure (model output is wrong)
+ - Provider variance (non-deterministic model)
+ - Flaky grader (model grader is inconsistent)
+ - Bad fixture (test input is unrealistic)
+ - Config mistake (assertion logic error)
+
## Safety Constraints
- - Do NOT log, echo, or store secrets (API keys) in configuration files or chat output.
+ - Do NOT log, echo, or store API keys in configuration files or chat output.
- Do NOT run evaluations against production endpoints without user consent.
- - Avoid executing arbitrary remote code or unvetted plugins during evaluation.
+ - Do not execute arbitrary remote code or unvetted plugins during evaluation.
## Validation / Done Criteria
- - Evaluation config is valid, minimal, and uses safe env refs.
- - Deterministic assertions exist where possible.
- - Run scope, provider, and cost are reported.
- - Results are summarized without leaking sensitive data.
+ - Eval config is valid, minimal, and uses `{{env.VAR}}` references for secrets.
+ - Deterministic assertions exist where possible; model grader use is documented and justified.
+ - Run scope, provider, and estimated cost are reported before execution.
+ - Results are summarized without leaking sensitive input data.
## References
- `references/eval-config-patterns.md`