bioprobench skillA
bioprobench is agent-read markdown (skill) from pku-yuangroup/openai4s: Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses..
Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.
What the file says
# BioProBench — protocol understanding and reasoning Biological protocols are where a plausible-sounding answer becomes a failed experiment: a wrong dosage, a swapped step, an unflagged hazard. BioProBench scores a model on five tasks over real wet-lab protocols, roughly 5,000 instances in the full release. | Task | What it measures | Metrics returned | | --- | --- | --- | | `PQA` | Protocol question answering — reagents, dosages, parameters | `Accuracy`, `Brier_Score`, `Failed_Rate` | | `ORD` | Step ordering — reconstructing procedural sequence | `Exact_Match`, `Kendall_Tau`, `Failed_Rate` | | `ERR` | Error correction — is this modified step valid | `accuracy`, `precision`, `recall`, `f1`, `failed_rate` | | `GEN` | Protocol generation — synthesising steps | `BLEU`, `METEOR`, `ROUGE-L`, `KW_F1`, `Step_Recall`, `Redundancy_Penalty`, `Failed_Rate` | | `REA-ERR` | Error reasoning, graded by an LLM judge | `Consistency`, `Failure_Rate`, `Total`, `Failed`, `Total_Items`, `Judged`, `Unjudged`, `Coverage` | Metric key casing differs per task — `ERR` returns lowercase keys, the rest are capitalised. Read them off the table above rather than guessing. …
Read the whole file at its exact version.
How to install
mdr add pku-yuangroup/openai4s/bioprobench@git:20260828.e9753bemdr add pku-yuangroup/openai4s/bioprobench@sha256:56ddbe9da8c2cd8aPin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.
[](https://markdownregistry.com/a/art_e253qbf6z7a34ikj)
1 badge views in 30 days
Versions
| version | committed | commit | size | audit | |
|---|---|---|---|---|---|
| git:20260828.e9753be latest | 2026-08-28 | e9753be | 9,298 B | A | view · diff |
| git:20260721.9d92317 | 2026-07-21 | 9d92317 | 9,234 B | A | view |
Audit of the latest version
- pass: Frontmatter block present
- pass: Frontmatter declares a name
- pass: Frontmatter declares a description
- pass: Size between 200 bytes and 200 KB (9298 bytes)
- pass: No zero-width or bidi control characters
- pass: No instruction hidden inside an HTML comment
- pass: No link to an exfiltration or paste host
- pass: No credential-shaped string
- pass: No instruction to send local credentials anywhere
- pass: No text hidden with inline styles
- pass: No prompt-injection phrasing
- pass: No curl or wget piped into a shell
- pass: No recursive delete of root, home or parent
- pass: No instruction to read or print local credentials
- pass: No base64 blob over 200 characters
- pass: No link to a raw IP address
- pass: No script tag
Source
pku-yuangroup/openai4s · 586 stars · license MIT · pushed 2026-09-23 · branch main
API
GET https://markdownregistry.com/api/v1/artifacts/art_e253qbf6z7a34ikj GET https://markdownregistry.com/api/v1/resolve?ref=pku-yuangroup/openai4s/bioprobench GET https://markdownregistry.com/api/v1/blob/56ddbe9da8c2cd8a041ffc984f59d24372fc74457d0183aabafb4c681c1bda91
Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.