Home / sickn33 / agentic-awesome-skills · plugins/agentic-awesome-skills-claude/skills/agent-evals/SKILL.md · GitHub

agent-evals skillB

agent-evals is agent-read markdown (skill) from sickn33/agentic-awesome-skills: Build automated evaluation suites for AI agents using golden datasets,.

Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.

What the file says

# Agent Evals

Create repeatable checks so agent behavior improves safely over time.

## When to Use This Skill

Use this skill when:
- Shipping new agent features or changing prompts
- Adding CI gates for agent quality and safety
- Building regression suites for tool-calling agents
- Measuring LLM output quality at scale
- Validating RAG retrieval accuracy

## Prerequisites

- Python 3.10+
- An LLM API key (OpenAI, Anthropic, etc.)
- pytest or a custom eval harness
- Optional: Braintrust, Promptfoo, or LangSmith account

## Evaluation Layers

### Unit Evals — Prompt-Level Correctness

Test individual prompt → response quality:

```python
# evals/test_unit.py
import json
import pytest
from agent import generate_response

CASES = json.load(open("evals/fixtures/unit_cases.json"))

@pytest.mark.parametrize("case", CASES, ids=lambda c: c["id"])
def test_prompt_correctness(case):
    result = generate_response(case["prompt"], model=case.get("model", "default"))
    # Exact match for structured output
    if case.get("expected_json"):
        assert json.loads(result) == case["expected_json"]
    # Substring match for free-text
    for keyword in case.get("must_contain", []):
…

Read the whole file at its exact version.

How to install

Latest version
mdr add sickn33/agentic-awesome-skills/agent-evals@v1.0
Exact content
mdr add sickn33/agentic-awesome-skills/agent-evals@sha256:f7b1131e7bd0c482

Pin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.

Badge

mdr badge

[![mdr](https://markdownregistry.com/badge/art_gdpryqxxvzj5pgw5.svg)](https://markdownregistry.com/a/art_gdpryqxxvzj5pgw5)

1 badge views in 30 days

Versions

versioncommittedcommitsizeaudit
v1.0 latest2026-09-20 d1b521d 12,537 BB view

Audit of the latest version

B  16 of 17 checks passed. Deterministic, no model, same answer every run.
  • fail: No prompt-injection phrasing (matched: Ignore all previous instructions)
  • pass: Frontmatter block present
  • pass: Frontmatter declares a name
  • pass: Frontmatter declares a description
  • pass: Size between 200 bytes and 200 KB (12537 bytes)
  • pass: No zero-width or bidi control characters
  • pass: No instruction hidden inside an HTML comment
  • pass: No link to an exfiltration or paste host
  • pass: No credential-shaped string
  • pass: No instruction to send local credentials anywhere
  • pass: No text hidden with inline styles
  • pass: No curl or wget piped into a shell
  • pass: No recursive delete of root, home or parent
  • pass: No instruction to read or print local credentials
  • pass: No base64 blob over 200 characters
  • pass: No link to a raw IP address
  • pass: No script tag

Source

GitHub

sickn33/agentic-awesome-skills · 46,848 stars · license MIT · pushed 2026-09-24 · branch main

API

GET https://markdownregistry.com/api/v1/artifacts/art_gdpryqxxvzj5pgw5
GET https://markdownregistry.com/api/v1/resolve?ref=sickn33/agentic-awesome-skills/agent-evals
GET https://markdownregistry.com/api/v1/blob/f7b1131e7bd0c482eee1edb16da79d9a8b9a3c49fa5a71af4e0737f26e0b4914

Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.

More from sickn33/agentic-awesome-skills

AGENTS.md agents
sickn33/agentic-awesome-skills · AGENTS.md
git:20260923.78093ac · audit A · 46,848 stars
llms.txt@apps/web-app/public llmstxt
sickn33/agentic-awesome-skills · apps/web-app/public/llms.txt
git:20260923.e2e8fa1 · audit A · 46,848 stars
00-andruia-consultant skill
sickn33/agentic-awesome-skills · plugins/agentic-awesome-skills-claude/skills/00-andruia-consultant/SKILL.md · Arquitecto de Soluciones Principal y Consultor Tecnológico de Andru.ia. Diagnostica y traza la hoja de ruta óptima para…
git:20260905.32fa521 · audit A · 46,848 stars
007 skill
sickn33/agentic-awesome-skills · plugins/agentic-awesome-skills-claude/skills/007/SKILL.md · Security audit, hardening, threat modeling (STRIDE/PASTA), Red/Blue Team, OWASP checks, code review, incident response…
git:20260905.f2e5074 · audit A · 46,848 stars
10-andruia-skill-smith skill
sickn33/agentic-awesome-skills · plugins/agentic-awesome-skills-claude/skills/10-andruia-skill-smith/SKILL.md · Ingeniero de Sistemas de Andru.ia. Diseña, redacta y despliega nuevas habilidades (skills) dentro del repositorio…
git:20260905.32fa521 · audit A · 46,848 stars
20-andruia-niche-intelligence skill
sickn33/agentic-awesome-skills · plugins/agentic-awesome-skills-claude/skills/20-andruia-niche-intelligence/SKILL.md · Estratega de Inteligencia de Dominio de Andru.ia. Analiza el nicho específico de un proyecto para inyectar…
git:20260905.32fa521 · audit A · 46,848 stars
2slides-ppt-generator skill
sickn33/agentic-awesome-skills · plugins/agentic-awesome-skills-claude/skills/2slides-ppt-generator/SKILL.md · AI-powered presentation generation via the 2slides API — create slides from text, match a reference image style…
git:20260905.f2e5074 · audit A · 46,848 stars
3d-web-experience skill
sickn33/agentic-awesome-skills · plugins/agentic-awesome-skills-claude/skills/3d-web-experience/SKILL.md · Expert in building 3D experiences for the web - Three.js, React
git:20260729.2d74b64 · audit A · 46,848 stars
ab-test-setup skill
sickn33/agentic-awesome-skills · plugins/agentic-awesome-skills-claude/skills/ab-test-setup/SKILL.md · Use when designing an A/B or split test: define the hypothesis, control and variants, estimate sample size, verify…
git:20260905.b0bf123 · audit A · 46,848 stars
ab-testing skill
sickn33/agentic-awesome-skills · plugins/agentic-awesome-skills-claude/skills/ab-testing/SKILL.md · When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.
git:20260905.f182d39 · audit A · 46,848 stars
acceptance-orchestrator skill
sickn33/agentic-awesome-skills · plugins/agentic-awesome-skills-claude/skills/acceptance-orchestrator/SKILL.md · Use when a coding task should be driven end-to-end from issue intake through implementation, review, deployment, and…
git:20260905.32fa521 · audit A · 46,848 stars
access-review skill
sickn33/agentic-awesome-skills · plugins/agentic-awesome-skills-claude/skills/access-review/SKILL.md · Conduct periodic access reviews and certifications. Implement access
v1.0 · audit A · 46,848 stars

Every file in sickn33/agentic-awesome-skills

Other files named agent-evals

agent-evals skill
nahid-sparktales/agent-dispatcher · skills/ai/agent-evals/SKILL.md · Build an eval suite that can actually detect a regression — cases pulled from real traffic, graders that check…
git:20260919.a0d4f55 · audit A · 49 stars

Browse by kind, by grade B, or by owner.