Home / yogsoth-ai / de-anthropocentric-research-engine · skills/benchmark-audit/SKILL.md · GitHub

benchmark-audit skillA

benchmark-audit is agent-read markdown (skill) from yogsoth-ai/de-anthropocentric-research-engine: Systematic quality assessment using BetterBench 46-criterion framework.

Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.

What the file says

# Benchmark Audit Strategy

Systematic quality assessment of AI/ML benchmarks using the BetterBench 46-criterion framework, Datasheets for Datasets standards, and established psychometric evaluation principles.

## Purpose

Produce a structured quality report for each target benchmark covering: documentation completeness, construct validity indicators, statistical robustness, maintenance status, and known failure modes.

## Budget

| Resource | Floor | Target |
|----------|-------|--------|
| Benchmarks audited | 3 | 5 |
| Papers read | 20 | 30 |
| Web searches | 25 | 40 |

## State Ledger

```
<HARD-GATE>
| Metric | Current | Target | Status |
|--------|---------|--------|--------|
| Benchmarks audited | 0 | 5 | PENDING |
| Papers fetched | 0 | 30 | PENDING |
| Papers read | 0 | 20 | PENDING |
| Web searches | 0 | 40 | PENDING |
| Documentation audits complete | 0 | 5 | PENDING |
| Metric decompositions complete | 0 | 5 | PENDING |
| Contamination checks complete | 0 | 5 | PENDING |
| Synthesis reports produced | 0 | 5 | PENDING |
</HARD-GATE>
```

Cannot exit until 80% of all targets met.

## Available Tactics
…

Read the whole file at its exact version.

How to install

Latest version
mdr add yogsoth-ai/de-anthropocentric-research-engine/benchmark-audit@git:20260616.b2b0cde
Exact content
mdr add yogsoth-ai/de-anthropocentric-research-engine/benchmark-audit@sha256:cefbf097360f2f79

Pin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.

Badge

mdr badge

[![mdr](https://markdownregistry.com/badge/art_bury2dgpjvt4q7p7.svg)](https://markdownregistry.com/a/art_bury2dgpjvt4q7p7)

1 badge views in 30 days

Versions

versioncommittedcommitsizeaudit
git:20260616.b2b0cde latest2026-06-16 b2b0cde 3,941 BA view · diff
git:20260615.b734c352026-06-15 b734c35 3,919 BA view · diff
git:20260615.7a629362026-06-15 7a62936 3,059 BA view · diff
git:20260519.dc15fe22026-05-19 dc15fe2 2,889 BA view

Audit of the latest version

A  17 of 17 checks passed. Deterministic, no model, same answer every run.
  • pass: Frontmatter block present
  • pass: Frontmatter declares a name
  • pass: Frontmatter declares a description
  • pass: Size between 200 bytes and 200 KB (3941 bytes)
  • pass: No zero-width or bidi control characters
  • pass: No instruction hidden inside an HTML comment
  • pass: No link to an exfiltration or paste host
  • pass: No credential-shaped string
  • pass: No instruction to send local credentials anywhere
  • pass: No text hidden with inline styles
  • pass: No prompt-injection phrasing
  • pass: No curl or wget piped into a shell
  • pass: No recursive delete of root, home or parent
  • pass: No instruction to read or print local credentials
  • pass: No base64 blob over 200 characters
  • pass: No link to a raw IP address
  • pass: No script tag

Source

GitHub

yogsoth-ai/de-anthropocentric-research-engine · 497 stars · license Apache-2.0 · pushed 2026-09-23 · branch main

API

GET https://markdownregistry.com/api/v1/artifacts/art_bury2dgpjvt4q7p7
GET https://markdownregistry.com/api/v1/resolve?ref=yogsoth-ai/de-anthropocentric-research-engine/benchmark-audit
GET https://markdownregistry.com/api/v1/blob/cefbf097360f2f795709648abebf9079baf8cb1dc8cdd1c2604f98eb08e7220b

Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.

More from yogsoth-ai/de-anthropocentric-research-engine

AGENTS.md agents
yogsoth-ai/de-anthropocentric-research-engine · AGENTS.md
git:20260727.358b018 · audit A · 497 stars
CLAUDE.md@ladder-foundry claude
yogsoth-ai/de-anthropocentric-research-engine · ladder-foundry/CLAUDE.md
git:20260701.4f1b4d0 · audit A · 497 stars
formated-results skill
yogsoth-ai/de-anthropocentric-research-engine · ladder-foundry/skills/formated-results/SKILL.md · Closing skill for the research-executor, loaded as the last step of formated-specs. Summarize the design just produced…
git:20260916.f5399c9 · audit A · 497 stars
formated-specs skill
yogsoth-ai/de-anthropocentric-research-engine · ladder-foundry/skills/formated-specs/SKILL.md · Spec-slot skill for the research-executor. Emit the 4-layer DARE orchestration of the assigned topic as one…
git:20260916.f5399c9 · audit A · 497 stars
injection-fidelity skill
yogsoth-ai/de-anthropocentric-research-engine · ladder-foundry/skills/injection-fidelity/SKILL.md · Loss-1 judge (codex role). Given one sample's de-identified dialogue and its PolicyCard, decide axis-by-axis whether…
git:20260916.f5399c9 · audit A · 497 stars
ladder-quality-order skill
yogsoth-ai/de-anthropocentric-research-engine · ladder-foundry/skills/ladder-quality-order/SKILL.md · Loss-2 judge (codex role). Over one topic's 6 shuffled research-design samples, pairwise-rank by quality using the…
git:20260916.f5399c9 · audit A · 497 stars
optimization-loop skill
yogsoth-ai/de-anthropocentric-research-engine · ladder-foundry/skills/optimization-loop/SKILL.md · The optimizer brain for the ladder-foundry pretraining loop. Runs the two-level nested batch loop, delegates gating to…
git:20260916.f5399c9 · audit A · 497 stars
acu-nugget-recall skill
yogsoth-ai/de-anthropocentric-research-engine · paper-reading/skills/acu-nugget-recall/SKILL.md · Tactic: Extract atomic units from one paper and score how much of a caller-supplied summary covers. Use for ACU-style…
v1.0.0 · audit A · 497 stars
argumentative-zoning skill
yogsoth-ai/de-anthropocentric-research-engine · paper-reading/skills/argumentative-zoning/SKILL.md · Tactic: Label every sentence of one paper with its rhetorical role using Argumentative Zoning. Use when fixed…
v1.0.0 · audit A · 497 stars
atomic-unit-matching skill
yogsoth-ai/de-anthropocentric-research-engine · paper-reading/skills/atomic-unit-matching/SKILL.md · Judge, per atomic content unit, whether a target text (summary, abstract, or other candidate text) contains it — binary…
v1.0.0 · audit A · 497 stars
atomic-unit-recall-aggregate skill
yogsoth-ai/de-anthropocentric-research-engine · paper-reading/skills/atomic-unit-recall-aggregate/SKILL.md · Aggregate per-unit ACU/Nugget match judgments into a final recall score — normalized length-penalized recall for ACU…
v1.0.0 · audit A · 497 stars
atomic-unit-writing skill
yogsoth-ai/de-anthropocentric-research-engine · paper-reading/skills/atomic-unit-writing/SKILL.md · Extract (ACU-style) or freshly author (Nugget-style) a list of atomic content units from a paper, optionally tagged…
v1.0.0 · audit A · 497 stars

Every file in yogsoth-ai/de-anthropocentric-research-engine

Browse by kind, by grade A, or by owner.