Home / nahid-sparktales / agent-dispatcher · skills/ai/agent-evals/SKILL.md · GitHub

agent-evals skillA

agent-evals is agent-read markdown (skill) from nahid-sparktales/agent-dispatcher: Build an eval suite that can actually detect a regression — cases pulled from real traffic, graders that check properties rather than vibes, a recorded baseline, and per-case diffs in both directions. Use before claiming a prompt, model or agent change is an improvement, when agent behaviour must not regress, or when someone reports "it seems better" after eyeballing a handful of outputs. Not for tracing what one run did (llm-observability), not for testing deterministic code, and never as evide.

Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.

What the file says

# Agent evals

"It seems better" is a sample of three, remembered favourably. An eval exists to turn a change in a
prompt, model, tool or retrieval step into a number you can defend — and to tell you honestly how
much of the system that number covers.

## When this fires

Before a prompt, model, tool definition or retrieval change is called an improvement; when agent
behaviour is depended on and must not silently regress; when a production failure needs to become
something that cannot come back. It does not fire for deterministic code paths, which are tests.

## Procedure

1. **Name the decision the eval serves** — "is this change an improvement", "can this ship", "which
   of these two prompts". A suite with no decision attached gets built once and read never. Write
   the decision above the suite.
2. **Take cases from real traffic, not imagination.** Pull them from logs, traces, support tickets
   and the failure that prompted this work. Invented cases test the behaviour you already thought
   of, which is the behaviour least likely to be broken. Freeze the set and version it with the
   code.
…

Read the whole file at its exact version.

How to install

Latest version
mdr add nahid-sparktales/agent-dispatcher/agent-evals@git:20260919.a0d4f55
Exact content
mdr add nahid-sparktales/agent-dispatcher/agent-evals@sha256:009248110b5152fe

Pin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.

Badge

mdr badge

[![mdr](https://markdownregistry.com/badge/art_e2mxvxoh4j5g4lg5.svg)](https://markdownregistry.com/a/art_e2mxvxoh4j5g4lg5)

1 badge views in 30 days

Versions

versioncommittedcommitsizeaudit
git:20260919.a0d4f55 latest2026-09-19 a0d4f55 7,454 BA view

Audit of the latest version

A  17 of 17 checks passed. Deterministic, no model, same answer every run.
  • pass: Frontmatter block present
  • pass: Frontmatter declares a name
  • pass: Frontmatter declares a description
  • pass: Size between 200 bytes and 200 KB (7454 bytes)
  • pass: No zero-width or bidi control characters
  • pass: No instruction hidden inside an HTML comment
  • pass: No link to an exfiltration or paste host
  • pass: No credential-shaped string
  • pass: No instruction to send local credentials anywhere
  • pass: No text hidden with inline styles
  • pass: No prompt-injection phrasing
  • pass: No curl or wget piped into a shell
  • pass: No recursive delete of root, home or parent
  • pass: No instruction to read or print local credentials
  • pass: No base64 blob over 200 characters
  • pass: No link to a raw IP address
  • pass: No script tag

Source

GitHub

nahid-sparktales/agent-dispatcher · 49 stars · license MIT · pushed 2026-09-23 · branch main

API

GET https://markdownregistry.com/api/v1/artifacts/art_e2mxvxoh4j5g4lg5
GET https://markdownregistry.com/api/v1/resolve?ref=nahid-sparktales/agent-dispatcher/agent-evals
GET https://markdownregistry.com/api/v1/blob/009248110b5152fef0ae33402f09bf3e12d339e867936a8ac31ff79749a10e23

Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.

More from nahid-sparktales/agent-dispatcher

agent-dispatcher skill
nahid-sparktales/agent-dispatcher · skills/agent-dispatcher/SKILL.md · Route work to a specialist role and load its task-specific guidance. Use when the user invokes /agent-dispatcher, names…
git:20260920.c183eff · audit A · 49 stars
agent-design skill
nahid-sparktales/agent-dispatcher · skills/ai/agent-design/SKILL.md · Scope an agent or subagent before it is built — the one job it owns, the smallest tool set that closes that job, what…
git:20260919.a0d4f55 · audit A · 49 stars
context-engineering skill
nahid-sparktales/agent-dispatcher · skills/ai/context-engineering/SKILL.md · Decide what actually occupies the model's window — progressive disclosure through an index, retrieval versus inlining…
git:20260919.a0d4f55 · audit A · 49 stars
llm-observability skill
nahid-sparktales/agent-dispatcher · skills/ai/llm-observability/SKILL.md · See what an agent actually did — one trace per run with nested model, tool and retrieval spans, token and latency…
git:20260919.a0d4f55 · audit A · 49 stars
mcp-design skill
nahid-sparktales/agent-dispatcher · skills/ai/mcp-design/SKILL.md · Build an MCP server, or bring an existing one into a project — choosing the transport, deciding which tools, resources…
git:20260919.a0d4f55 · audit A · 49 stars
memory-design skill
nahid-sparktales/agent-dispatcher · skills/ai/memory-design/SKILL.md · Decide what an agent should remember, which layer holds it, who it is scoped to, and how a stale or contradicted memory…
git:20260919.a0d4f55 · audit A · 49 stars
model-routing skill
nahid-sparktales/agent-dispatcher · skills/ai/model-routing/SKILL.md · Pick a model per job and degrade sensibly when one fails — a quality bar per call site, candidates compared on the same…
git:20260919.a0d4f55 · audit A · 49 stars
prompt-engineering skill
nahid-sparktales/agent-dispatcher · skills/ai/prompt-engineering/SKILL.md · Write or revise a prompt so it holds up — output contract, instruction placement, examples that earn their place, an…
git:20260919.a0d4f55 · audit A · 49 stars
prompt-injection-defense skill
nahid-sparktales/agent-dispatcher · skills/ai/prompt-injection-defense/SKILL.md · Treat everything an agent reads but did not author as data rather than instructions — an explicit trust boundary, a…
git:20260919.a0d4f55 · audit A · 49 stars
retrieval-rag skill
nahid-sparktales/agent-dispatcher · skills/ai/retrieval-rag/SKILL.md · Build and fix retrieval that actually returns the right passage — structure-aware chunking, one pinned embedding model…
git:20260919.a0d4f55 · audit A · 49 stars
structured-output skill
nahid-sparktales/agent-dispatcher · skills/ai/structured-output/SKILL.md · Get parseable, trustworthy structured results out of a model — schema design, the enforcement mechanism the provider…
git:20260919.a0d4f55 · audit A · 49 stars
tool-design skill
nahid-sparktales/agent-dispatcher · skills/ai/tool-design/SKILL.md · Design the tools a model calls — names, parameter shapes, what a result returns, and error text written as an…
git:20260919.a0d4f55 · audit A · 49 stars

Every file in nahid-sparktales/agent-dispatcher

Other files named agent-evals

agent-evals skill
sickn33/agentic-awesome-skills · plugins/agentic-awesome-skills-claude/skills/agent-evals/SKILL.md · Build automated evaluation suites for AI agents using golden datasets,
v1.0 · audit B · 46,848 stars

Browse by kind, by grade A, or by owner.