Design and validity review for studies that benchmark one or more AI systems against a human-expert panel as the reference. Covers the evaluation question and arm definition, decoupled multi-dimensional rubrics with anchors, planted calibration probes, reviewer-panel construction, inter-rater reliability targets, LLM-as-judge versus human-as-judge adjudication, construct-independence guards, and a structured rating-export schema. Use before data collection on an AI-vs-expert evaluation.
mdr add aperivue/medsci-skills/design-ai-benchmarking@git:20260625.92ee046mdr add aperivue/medsci-skills/design-ai-benchmarking@sha256:b8f794a1f6c800d8[](https://markdownregistry.com/a/art_6jwjb34n77rpx3ix)
0 badge views in 30 days
| version | committed | commit | size | audit | |
|---|---|---|---|---|---|
| git:20260625.92ee046 latest | 2026-06-25 | 92ee046 | 12,094 B | A | view · diff |
| git:20260605.041597a | 2026-06-05 | 041597a | 10,820 B | A | view |
aperivue/medsci-skills · 283 stars · license MIT · pushed 2026-09-05 · branch main
GET https://markdownregistry.com/api/v1/artifacts/art_6jwjb34n77rpx3ix GET https://markdownregistry.com/api/v1/resolve?ref=aperivue/medsci-skills/design-ai-benchmarking GET https://markdownregistry.com/api/v1/blob/b8f794a1f6c800d821305a4df8a797bea61cf34a602e0dc0dbea8f2c0c458ca5