scdenney/open-science-skills · codex/llm-calibration-logprobs/SKILL.md

llm-calibration-logprobs skillA

Reads a model's own uncertainty off its token log-probabilities — collecting logprobs and aggregating multi-token labels, confidence tiers and margins for triage, calibration assessment with ECE, Brier scores, and reliability diagrams, using confidence downstream without laundering it into evidence, and what to archive for reproducibility. Use when the user asks how confident a classifier was on each item, asks about logprobs, top-k tokens, calibration, ECE, Brier, or reliability diagrams, or wa

Latest version
mdr add scdenney/open-science-skills/llm-calibration-logprobs@git:20260905.39e8276
Exact content
mdr add scdenney/open-science-skills/llm-calibration-logprobs@sha256:fd1573e1b07b4657
Badge

mdr badge

[![mdr](https://markdownregistry.com/badge/art_p3cil5c5uvveqmos.svg)](https://markdownregistry.com/a/art_p3cil5c5uvveqmos)

0 badge views in 30 days

Versions

versioncommittedcommitsizeaudit
git:20260905.39e8276 latest2026-09-05 39e8276 19,621 BA view · diff
git:20260902.da5a2632026-09-02 da5a263 19,154 BA view · diff
git:20260808.c81b3502026-08-08 c81b350 18,573 BA view · diff
git:20260702.0d142fe2026-07-02 0d142fe 18,741 BA view

Audit of the latest version

A  17 of 17 checks passed. Deterministic, no model, same answer every run.

Source

GitHub

scdenney/open-science-skills · 54 stars · license NOASSERTION · pushed 2026-09-05 · branch main

API

GET https://markdownregistry.com/api/v1/artifacts/art_p3cil5c5uvveqmos
GET https://markdownregistry.com/api/v1/resolve?ref=scdenney/open-science-skills/llm-calibration-logprobs
GET https://markdownregistry.com/api/v1/blob/fd1573e1b07b4657609aeda446ef7e811001f9f32f86174641cb7bfa9dc6a118