monitor-evals skillA
monitor-evals is agent-read markdown (skill) from ai-analyst-lab/ai-analyst: Review evaluation history, distinguish system, data, suite, and evaluator changes, and apply an operating response. Use for regressions, eval monitoring, release gates, or trend questions..
Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.
What the file says
# Monitor evaluation history Read versioned run manifests with `helpers.evals.monitoring.load_history`. Use `classify_changes` before comparing scores. Separate: - system behavior changes; - data changes; - task-mix or suite-version changes; - evaluator changes; and - operational failures such as blocked tools or expired connections. Compare only compatible runs. Show per-case and slice movement, not only the aggregate. Run frozen sentinel examples for model graders so evaluator drift does not look like system drift. Apply the named operating rule: - continue when the intended change improved the target slice without a blocking regression; - investigate when the cause is unclear or several inputs changed; - rollback when a blocking regression follows a controlled system change; or - escalate when the evaluator, data, or authorization boundary may be invalid. Record the owner and next action. Monitoring is an operating practice, not a dashboard someone passively observes.
Read the whole file at its exact version.
How to install
mdr add ai-analyst-lab/ai-analyst/monitor-evals@git:20260910.c4b91efmdr add ai-analyst-lab/ai-analyst/monitor-evals@sha256:f0a20e96ae603f90Pin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.
[](https://markdownregistry.com/a/art_cptu4tbez5huomsb)
1 badge views in 30 days
Versions
Audit of the latest version
- pass: Frontmatter block present
- pass: Frontmatter declares a name
- pass: Frontmatter declares a description
- pass: Size between 200 bytes and 200 KB (1224 bytes)
- pass: No zero-width or bidi control characters
- pass: No instruction hidden inside an HTML comment
- pass: No link to an exfiltration or paste host
- pass: No credential-shaped string
- pass: No instruction to send local credentials anywhere
- pass: No text hidden with inline styles
- pass: No prompt-injection phrasing
- pass: No curl or wget piped into a shell
- pass: No recursive delete of root, home or parent
- pass: No instruction to read or print local credentials
- pass: No base64 blob over 200 characters
- pass: No link to a raw IP address
- pass: No script tag
Source
ai-analyst-lab/ai-analyst · 302 stars · license MIT · pushed 2026-09-22 · branch main
API
GET https://markdownregistry.com/api/v1/artifacts/art_cptu4tbez5huomsb GET https://markdownregistry.com/api/v1/resolve?ref=ai-analyst-lab/ai-analyst/monitor-evals GET https://markdownregistry.com/api/v1/blob/f0a20e96ae603f900686f87ee04b90c854fe5a3ff4ce10999aee32c995290425
Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.