Home / yogsoth-ai / de-anthropocentric-research-engine · skills/benchmark-archaeology/SKILL.md · GitHub

benchmark-archaeology skillA

benchmark-archaeology is agent-read markdown (skill) from yogsoth-ai/de-anthropocentric-research-engine: Evaluation Methodology Archaeology Campaign — 5 strategies for systematic.

Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.

What the file says

# Benchmark Archaeology

Systematic excavation and critical analysis of AI/ML evaluation methodology. Treats benchmarks as historical artifacts requiring forensic examination — uncovering hidden assumptions, methodological drift, validity decay, and coverage gaps that accumulate over time.

## Strategy Routing

| Signal | Route To |
|--------|----------|
| "audit this benchmark", "benchmark quality", "BetterBench" | benchmark-audit |
| "saturation", "ceiling", "score plateau", "when will X be solved" | saturation-analysis |
| "does it actually measure", "construct validity", "what does score mean" | validity-probing |
| "what's not tested", "coverage gaps", "missing capabilities" | coverage-mapping |
| "different papers get different scores", "protocol differences" | protocol-forensics |

## Manifest

### Strategies (5)

| Strategy | Purpose |
|----------|---------|
| benchmark-audit | Systematic quality assessment using BetterBench 46-criterion framework |
| saturation-analysis | Track score trajectories, detect saturation and failure points |
| validity-probing | Challenge construct validity — does benchmark measure claimed capability? |
…

Read the whole file at its exact version.

How to install

Latest version
mdr add yogsoth-ai/de-anthropocentric-research-engine/benchmark-archaeology@git:20260809.fa000d2
Exact content
mdr add yogsoth-ai/de-anthropocentric-research-engine/benchmark-archaeology@sha256:6e9f5d1fed037b69

Pin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.

Badge

mdr badge

[![mdr](https://markdownregistry.com/badge/art_xytgn2ezbllyimot.svg)](https://markdownregistry.com/a/art_xytgn2ezbllyimot)

1 badge views in 30 days

Versions

versioncommittedcommitsizeaudit
git:20260809.fa000d2 latest2026-08-09 fa000d2 5,903 BA view · diff
git:20260616.b2b0cde2026-06-16 b2b0cde 5,931 BA view · diff
git:20260615.b734c352026-06-15 b734c35 5,909 BA view · diff
git:20260615.7a629362026-06-15 7a62936 4,322 BA view · diff
git:20260519.dc15fe22026-05-19 dc15fe2 4,164 BA view

Audit of the latest version

A  17 of 17 checks passed. Deterministic, no model, same answer every run.
  • pass: Frontmatter block present
  • pass: Frontmatter declares a name
  • pass: Frontmatter declares a description
  • pass: Size between 200 bytes and 200 KB (5903 bytes)
  • pass: No zero-width or bidi control characters
  • pass: No instruction hidden inside an HTML comment
  • pass: No link to an exfiltration or paste host
  • pass: No credential-shaped string
  • pass: No instruction to send local credentials anywhere
  • pass: No text hidden with inline styles
  • pass: No prompt-injection phrasing
  • pass: No curl or wget piped into a shell
  • pass: No recursive delete of root, home or parent
  • pass: No instruction to read or print local credentials
  • pass: No base64 blob over 200 characters
  • pass: No link to a raw IP address
  • pass: No script tag

Source

GitHub

yogsoth-ai/de-anthropocentric-research-engine · 498 stars · license Apache-2.0 · pushed 2026-09-17 · branch main

API

GET https://markdownregistry.com/api/v1/artifacts/art_xytgn2ezbllyimot
GET https://markdownregistry.com/api/v1/resolve?ref=yogsoth-ai/de-anthropocentric-research-engine/benchmark-archaeology
GET https://markdownregistry.com/api/v1/blob/6e9f5d1fed037b6994e64a7388588cb70a451220eca1d6b82b22875b3c38f643

Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.

More from yogsoth-ai/de-anthropocentric-research-engine

AGENTS.md agents
yogsoth-ai/de-anthropocentric-research-engine · AGENTS.md
git:20260727.358b018 · audit A · 498 stars
CLAUDE.md@ladder-foundry claude
yogsoth-ai/de-anthropocentric-research-engine · ladder-foundry/CLAUDE.md
git:20260701.4f1b4d0 · audit A · 498 stars
formated-results skill
yogsoth-ai/de-anthropocentric-research-engine · ladder-foundry/skills/formated-results/SKILL.md · Closing skill for the research-executor, loaded as the last step of formated-specs. Summarize the design just produced…
git:20260916.f5399c9 · audit A · 498 stars
formated-specs skill
yogsoth-ai/de-anthropocentric-research-engine · ladder-foundry/skills/formated-specs/SKILL.md · Spec-slot skill for the research-executor. Emit the 4-layer DARE orchestration of the assigned topic as one…
git:20260916.f5399c9 · audit A · 498 stars
injection-fidelity skill
yogsoth-ai/de-anthropocentric-research-engine · ladder-foundry/skills/injection-fidelity/SKILL.md · Loss-1 judge (codex role). Given one sample's de-identified dialogue and its PolicyCard, decide axis-by-axis whether…
git:20260916.f5399c9 · audit A · 498 stars
ladder-quality-order skill
yogsoth-ai/de-anthropocentric-research-engine · ladder-foundry/skills/ladder-quality-order/SKILL.md · Loss-2 judge (codex role). Over one topic's 6 shuffled research-design samples, pairwise-rank by quality using the…
git:20260916.f5399c9 · audit A · 498 stars
optimization-loop skill
yogsoth-ai/de-anthropocentric-research-engine · ladder-foundry/skills/optimization-loop/SKILL.md · The optimizer brain for the ladder-foundry pretraining loop. Runs the two-level nested batch loop, delegates gating to…
git:20260916.f5399c9 · audit A · 498 stars
acu-nugget-recall skill
yogsoth-ai/de-anthropocentric-research-engine · paper-reading/skills/acu-nugget-recall/SKILL.md · Tactic: Extract atomic units from one paper and score how much of a caller-supplied summary covers. Use for ACU-style…
v1.0.0 · audit A · 498 stars
argumentative-zoning skill
yogsoth-ai/de-anthropocentric-research-engine · paper-reading/skills/argumentative-zoning/SKILL.md · Tactic: Label every sentence of one paper with its rhetorical role using Argumentative Zoning. Use when fixed…
v1.0.0 · audit A · 498 stars
atomic-unit-matching skill
yogsoth-ai/de-anthropocentric-research-engine · paper-reading/skills/atomic-unit-matching/SKILL.md · Judge, per atomic content unit, whether a target text (summary, abstract, or other candidate text) contains it — binary…
v1.0.0 · audit A · 498 stars
atomic-unit-recall-aggregate skill
yogsoth-ai/de-anthropocentric-research-engine · paper-reading/skills/atomic-unit-recall-aggregate/SKILL.md · Aggregate per-unit ACU/Nugget match judgments into a final recall score — normalized length-penalized recall for ACU…
v1.0.0 · audit A · 498 stars
atomic-unit-writing skill
yogsoth-ai/de-anthropocentric-research-engine · paper-reading/skills/atomic-unit-writing/SKILL.md · Extract (ACU-style) or freshly author (Nugget-style) a list of atomic content units from a paper, optionally tagged…
v1.0.0 · audit A · 498 stars

Every file in yogsoth-ai/de-anthropocentric-research-engine

Browse by kind, by grade A, or by owner.