haabe/mycelium · plugins/mycelium/skills/eval-runner/SKILL.md

eval-runner skillA

Run eval scenarios to benchmark Mycelium effectiveness. Execute tasks using reflexion loop, validate against success criteria, record metrics.

Latest version
mdr add haabe/mycelium/eval-runner@git:20260704.7639f18
Exact content
mdr add haabe/mycelium/eval-runner@sha256:f398c17e0461ce51
Badge

mdr badge

[![mdr](https://markdownregistry.com/badge/art_5vgmtxhbn3g5pz24.svg)](https://markdownregistry.com/a/art_5vgmtxhbn3g5pz24)

0 badge views in 30 days

Versions

versioncommittedcommitsizeaudit
git:20260704.7639f18 latest2026-07-04 7639f18 4,277 BA view · diff
git:20260525.9fac9422026-05-25 9fac942 4,277 BA view · diff
git:20260522.1825b222026-05-22 1825b22 3,967 BA view · diff
git:20260509.70c07fb2026-05-09 70c07fb 3,953 BA view · diff
git:20260509.46569102026-05-09 4656910 3,937 BA view

Audit of the latest version

A  17 of 17 checks passed. Deterministic, no model, same answer every run.

Source

GitHub

haabe/mycelium · 45 stars · license MIT · pushed 2026-09-04 · branch main

API

GET https://markdownregistry.com/api/v1/artifacts/art_5vgmtxhbn3g5pz24
GET https://markdownregistry.com/api/v1/resolve?ref=haabe/mycelium/eval-runner
GET https://markdownregistry.com/api/v1/blob/f398c17e0461ce517d24b9d61981a568500acd9ea89b6a6ccd8c49d31fdebfdf