agent-evals skillA
agent-evals is agent-read markdown (skill) from nahid-sparktales/agent-dispatcher: Build an eval suite that can actually detect a regression — cases pulled from real traffic, graders that check properties rather than vibes, a recorded baseline, and per-case diffs in both directions. Use before claiming a prompt, model or agent change is an improvement, when agent behaviour must not regress, or when someone reports "it seems better" after eyeballing a handful of outputs. Not for tracing what one run did (llm-observability), not for testing deterministic code, and never as evide.
Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.
What the file says
# Agent evals "It seems better" is a sample of three, remembered favourably. An eval exists to turn a change in a prompt, model, tool or retrieval step into a number you can defend — and to tell you honestly how much of the system that number covers. ## When this fires Before a prompt, model, tool definition or retrieval change is called an improvement; when agent behaviour is depended on and must not silently regress; when a production failure needs to become something that cannot come back. It does not fire for deterministic code paths, which are tests. ## Procedure 1. **Name the decision the eval serves** — "is this change an improvement", "can this ship", "which of these two prompts". A suite with no decision attached gets built once and read never. Write the decision above the suite. 2. **Take cases from real traffic, not imagination.** Pull them from logs, traces, support tickets and the failure that prompted this work. Invented cases test the behaviour you already thought of, which is the behaviour least likely to be broken. Freeze the set and version it with the code. …
Read the whole file at its exact version.
How to install
mdr add nahid-sparktales/agent-dispatcher/agent-evals@git:20260919.a0d4f55mdr add nahid-sparktales/agent-dispatcher/agent-evals@sha256:009248110b5152fePin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.
[](https://markdownregistry.com/a/art_e2mxvxoh4j5g4lg5)
1 badge views in 30 days
Versions
Audit of the latest version
- pass: Frontmatter block present
- pass: Frontmatter declares a name
- pass: Frontmatter declares a description
- pass: Size between 200 bytes and 200 KB (7454 bytes)
- pass: No zero-width or bidi control characters
- pass: No instruction hidden inside an HTML comment
- pass: No link to an exfiltration or paste host
- pass: No credential-shaped string
- pass: No instruction to send local credentials anywhere
- pass: No text hidden with inline styles
- pass: No prompt-injection phrasing
- pass: No curl or wget piped into a shell
- pass: No recursive delete of root, home or parent
- pass: No instruction to read or print local credentials
- pass: No base64 blob over 200 characters
- pass: No link to a raw IP address
- pass: No script tag
Source
nahid-sparktales/agent-dispatcher · 49 stars · license MIT · pushed 2026-09-23 · branch main
API
GET https://markdownregistry.com/api/v1/artifacts/art_e2mxvxoh4j5g4lg5 GET https://markdownregistry.com/api/v1/resolve?ref=nahid-sparktales/agent-dispatcher/agent-evals GET https://markdownregistry.com/api/v1/blob/009248110b5152fef0ae33402f09bf3e12d339e867936a8ac31ff79749a10e23
Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.
More from nahid-sparktales/agent-dispatcher
Every file in nahid-sparktales/agent-dispatcher