Home / miaodx / roboclaws · skills/eval-harness/SKILL.md · GitHub

eval-harness skillA

eval-harness is agent-read markdown (skill) from miaodx/roboclaws: Select and run Roboclaws validation, product, eval-suite, and live-agent eval rows from a plan, diff, or explicit capability request..

Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.

What the file says

# Eval Harness

Use this skill when a Roboclaws plan, diff, PR, or agent-facing change needs
one maintainer proof surface. The skill answers:

1. which deterministic gates, product runs, eval suites, and live-agent evals
   are relevant;
2. why each row was selected, skipped, run, failed, or blocked;
3. where the resulting reports and regression-promotion evidence live.

It is an orchestration skill, not a robot behavior skill. Keep household-world
task strategy in `household-world`; keep reusable robot capability semantics
in MCP tools and capability profiles.

Use Just only as the human-facing facade shown below. Frozen rows execute their
package owners directly: eval rows use `python -m roboclaws.evals.cli`, and
product rows use `python -m roboclaws.cli.main run surface`. Keep the subprocess
boundary for row timeout and isolation; do not route an executing row back
through Just or add a second command registry. The eval CLI grammar is an
optional documented kebab-case tool name followed by `key=value` arguments.

## Commands

Recommend rows without running them:

```bash
just agent::eval recommend plan=docs/plans/example.md budget=focused
```
…

Read the whole file at its exact version.

How to install

Latest version
mdr add miaodx/roboclaws/eval-harness@git:20260918.a03741f
Exact content
mdr add miaodx/roboclaws/eval-harness@sha256:d3a659bb525c8bad

Pin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.

Badge

mdr badge

[![mdr](https://markdownregistry.com/badge/art_rzp6gehqhr36wo2h.svg)](https://markdownregistry.com/a/art_rzp6gehqhr36wo2h)

1 badge views in 30 days

Versions

versioncommittedcommitsizeaudit
git:20260918.a03741f latest2026-09-18 a03741f 7,558 BA view · diff
git:20260918.67dc4c52026-09-18 67dc4c5 6,909 BA view · diff
git:20260918.85c8d522026-09-18 85c8d52 6,906 BA view · diff
git:20260826.c82bec72026-08-26 c82bec7 6,653 BA view · diff
git:20260816.e6e11f52026-08-16 e6e11f5 6,673 BA view

Audit of the latest version

A  17 of 17 checks passed. Deterministic, no model, same answer every run.
  • pass: Frontmatter block present
  • pass: Frontmatter declares a name
  • pass: Frontmatter declares a description
  • pass: Size between 200 bytes and 200 KB (7558 bytes)
  • pass: No zero-width or bidi control characters
  • pass: No instruction hidden inside an HTML comment
  • pass: No link to an exfiltration or paste host
  • pass: No credential-shaped string
  • pass: No instruction to send local credentials anywhere
  • pass: No text hidden with inline styles
  • pass: No prompt-injection phrasing
  • pass: No curl or wget piped into a shell
  • pass: No recursive delete of root, home or parent
  • pass: No instruction to read or print local credentials
  • pass: No base64 blob over 200 characters
  • pass: No link to a raw IP address
  • pass: No script tag

Source

GitHub

miaodx/roboclaws · 6 stars · license MIT · pushed 2026-09-24 · branch main

API

GET https://markdownregistry.com/api/v1/artifacts/art_rzp6gehqhr36wo2h
GET https://markdownregistry.com/api/v1/resolve?ref=miaodx/roboclaws/eval-harness
GET https://markdownregistry.com/api/v1/blob/d3a659bb525c8bad6f61f1fc2154f47598da5166241d5b5abb9993fd984a7860

Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.

More from miaodx/roboclaws

AGENTS.md agents
miaodx/roboclaws · AGENTS.md
git:20260816.ecc907e · audit A · 6 stars
CLAUDE.md claude
miaodx/roboclaws · CLAUDE.md
git:20260626.e56711f · audit A · 6 stars
cloudml-eval-ops skill
miaodx/roboclaws · skills/cloudml-eval-ops/SKILL.md · Run frozen Roboclaws Eval Harness rows on CloudML with bounded parallelism, official cml lifecycle commands…
git:20260826.4cd1c7a · audit A · 6 stars
eval-evolution skill
miaodx/roboclaws · skills/eval-evolution/SKILL.md · Prepare, run, review, and explicitly promote bounded Skill or existing-MCP Eval Evolution campaigns through the…
git:20260816.2788518 · audit A · 6 stars
household-world skill
miaodx/roboclaws · skills/household-world/SKILL.md · Complete household-world goals through public household MCP tools.
git:20260918.c6faf47 · audit A · 6 stars
raw-fpv-visual-labeler skill
miaodx/roboclaws · skills/raw-fpv-visual-labeler/SKILL.md · Label visible cleanup-relevant movable objects from grouped RAW-FPV frames without creating executable cleanup handles.
git:20260622.3dd53c6 · audit A · 6 stars
report-performance-analysis skill
miaodx/roboclaws · skills/report-performance-analysis/SKILL.md · Use this skill whenever the user asks whether a Roboclaws live-agent report run is faster, wants to compare Agent…
git:20260728.cda130b · audit A · 6 stars
scene-gaussian-map-alignment skill
miaodx/roboclaws · skills/scene-gaussian-map-alignment/SKILL.md · Align scene Gaussian/splat, USD/mesh, and robot map assets into an honest digital-twin evidence workflow. Use when a…
git:20260816.3e0c2f2 · audit A · 6 stars
visual-result-showcase skill
miaodx/roboclaws · skills/visual-result-showcase/SKILL.md · Create blog, README, social, or demo-review visual showcases from completed Roboclaws run artifacts, especially…
git:20260816.39a87d9 · audit A · 6 stars

Every file in miaodx/roboclaws

Other files named eval-harness

eval-harness skill
affaan-m/ecc · .agents/skills/eval-harness/SKILL.md · Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles. Use when a…
git:20260918.c752aac · audit A · 266,832 stars
eval-harness skill
affaan-m/ecc · docs/es/skills/eval-harness/SKILL.md · Framework formal de evaluación para sesiones de Claude Code que implementa principios de desarrollo orientado a evals…
git:20260607.ac0f11c · audit A · 266,832 stars
eval-harness skill
a5c-ai/babysitter · library/methodologies/everything-claude-code/skills/eval-harness/SKILL.md · Evaluation harness for testing agent and skill quality through structured benchmarks, regression tests, and quality…
git:20260601.da7723a · audit A · 1,809 stars
eval-harness skill
hashgraph-online/awesome-codex-plugins · plugins/Colin4k1024/tsp/skills/eval-harness/SKILL.md · 正式评估框架,实现 eval-driven development (EDD) 原则。 用于定义 pass/fail 标准、测量 pass@k 指标、创建回归测试套件。
git:20260512.6523792 · audit A · 1,074 stars
eval-harness skill
jikig-ai/soleur · plugins/soleur/skills/eval-harness/SKILL.md · This skill provides a promptfoo eval harness that measures whether a Soleur skill or agent edit actually improves…
git:20260916.6c1dbcb · audit A · 15 stars

Browse by kind, by grade A, or by owner.