benchmark-reporting skillA
benchmark-reporting is agent-read markdown (skill) from khaledsaeed18/dotclaude: Report computational benchmarks and experimental comparisons the way examiners and reviewers expect: fair baselines run under the same conditions, multiple seeds with variance, hardware and configuration disclosed, tables generated from result files with consistent precision and the best result marked by rule, ablations that isolate each component, and honest treatment of losses and failure cases. Use when writing the results chapter of a systems or ML thesis, when a comparison table needs to be.
Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.
What the file says
A comparison is credible when the reader can see that every system had the same chance. Most rejected results chapters fail on fairness or on variance, not on the numbers themselves. ## Fairness - **Same conditions**: same hardware, same data splits, same preprocessing, same evaluation script, same budget (epochs, wall-clock, or tokens; state which). Numbers copied from other papers are labelled as such in the table and are not the headline comparison. - **Tuned baselines**: each baseline gets the hyperparameter search the proposed method got, or its authors' recommended settings, stated. Record the search space and budget per system. - **Current baselines**: the strongest published method for the task at thesis time, not a convenient old one. If the strongest is out of reach (compute, code), say so and compare with what is feasible. - **Same metric implementation**: one evaluation script for all systems; metric definitions and any thresholds stated. ## Variance - Every number is a mean over ≥ 3 seeds (or runs, or folds) with SD or a 95% CI, and n stated in the caption. A single run is reported as such and not compared. …
Read the whole file at its exact version.
How to install
mdr add khaledsaeed18/dotclaude/benchmark-reporting@git:20260921.c8f5dc3mdr add khaledsaeed18/dotclaude/benchmark-reporting@sha256:ce406c701b9ee422Pin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.
[](https://markdownregistry.com/a/art_eqdr6cepwgzq5j3m)
1 badge views in 30 days
Versions
Audit of the latest version
- pass: Frontmatter block present
- pass: Frontmatter declares a name
- pass: Frontmatter declares a description
- pass: Size between 200 bytes and 200 KB (5115 bytes)
- pass: No zero-width or bidi control characters
- pass: No instruction hidden inside an HTML comment
- pass: No link to an exfiltration or paste host
- pass: No credential-shaped string
- pass: No instruction to send local credentials anywhere
- pass: No text hidden with inline styles
- pass: No prompt-injection phrasing
- pass: No curl or wget piped into a shell
- pass: No recursive delete of root, home or parent
- pass: No instruction to read or print local credentials
- pass: No base64 blob over 200 characters
- pass: No link to a raw IP address
- pass: No script tag
Source
khaledsaeed18/dotclaude · 5 stars · license MIT · pushed 2026-09-22 · branch main
API
GET https://markdownregistry.com/api/v1/artifacts/art_eqdr6cepwgzq5j3m GET https://markdownregistry.com/api/v1/resolve?ref=khaledsaeed18/dotclaude/benchmark-reporting GET https://markdownregistry.com/api/v1/blob/ce406c701b9ee42221b0d38ff3f29e2f89616bfcc1007e923aef45666289618e
Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.