Home / andersonlimahw / lemon-ai-hub · plugins/karpathy-graph/skills/karpathy-graph-evaluate/SKILL.md · GitHub

karpathy-graph-evaluate skillA

karpathy-graph-evaluate is agent-read markdown (skill) from andersonlimahw/lemon-ai-hub: Stage 5 of the karpathy-graph pipeline. Score graph quality across extraction, resolution, structure, query, and operations, and run a ratchet loop that keeps only changes which improve the metric. Use when the user asks how good the graph is, wants to tune extraction prompts or the ontology, asks to measure or improve graph quality, or wants to detect resolution regressions..

Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.

What the file says

# Evaluate

The evaluation harness has the same shape as Karpathy's autoresearch loop. The
artifact being optimized is not `train.py` — it is the extraction prompt, the
ontology, the resolution policy, or the query serializer.

```
Read current extraction prompt and score history
  -> Propose ONE motivated change
  -> Run extraction on the gold set
  -> Compute precision, recall, F1, cost, latency
  -> If improved: keep.  If worse: revert.
  -> Record and continue.
```

One change at a time. A batch of five changes that improves the metric teaches
you nothing about which of the five worked.

## Build a gold set first

Without one you are guessing. Take 10–30 representative units of context and
hand-label the entities and relations that should be extracted. Include
adversarial cases deliberately:

- misleading aliases (two different things with near-identical names)
- contradictory dates across sources
- entities that *should not* merge despite high string similarity
- claims with no supporting source
- disconnected paths that look connected

A gold set of only easy cases certifies a system that fails on the hard ones.

## Metrics by layer
…

Read the whole file at its exact version.

How to install

Latest version
mdr add andersonlimahw/lemon-ai-hub/karpathy-graph-evaluate@v1.0.0
Exact content
mdr add andersonlimahw/lemon-ai-hub/karpathy-graph-evaluate@sha256:52976a80ad448a27

Pin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.

Badge

mdr badge

[![mdr](https://markdownregistry.com/badge/art_kxbaiemhksulspzr.svg)](https://markdownregistry.com/a/art_kxbaiemhksulspzr)

1 badge views in 30 days

Versions

versioncommittedcommitsizeaudit
v1.0.0 latest2026-08-23 5f57f3b 5,437 BA view

Audit of the latest version

A  17 of 17 checks passed. Deterministic, no model, same answer every run.
  • pass: Frontmatter block present
  • pass: Frontmatter declares a name
  • pass: Frontmatter declares a description
  • pass: Size between 200 bytes and 200 KB (5437 bytes)
  • pass: No zero-width or bidi control characters
  • pass: No instruction hidden inside an HTML comment
  • pass: No link to an exfiltration or paste host
  • pass: No credential-shaped string
  • pass: No instruction to send local credentials anywhere
  • pass: No text hidden with inline styles
  • pass: No prompt-injection phrasing
  • pass: No curl or wget piped into a shell
  • pass: No recursive delete of root, home or parent
  • pass: No instruction to read or print local credentials
  • pass: No base64 blob over 200 characters
  • pass: No link to a raw IP address
  • pass: No script tag

Source

GitHub

andersonlimahw/lemon-ai-hub · 14 stars · license none · pushed 2026-09-20 · branch main

API

GET https://markdownregistry.com/api/v1/artifacts/art_kxbaiemhksulspzr
GET https://markdownregistry.com/api/v1/resolve?ref=andersonlimahw/lemon-ai-hub/karpathy-graph-evaluate
GET https://markdownregistry.com/api/v1/blob/52976a80ad448a2767a83fccae9fffadc408d7b2ca8cd7fb2c34df9119b5970c

Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.

More from andersonlimahw/lemon-ai-hub

AGENTS.md agents
andersonlimahw/lemon-ai-hub · AGENTS.md
git:20260905.88fc672 · audit A · 14 stars
CLAUDE.md claude
andersonlimahw/lemon-ai-hub · CLAUDE.md
git:20260905.88fc672 · audit A · 14 stars
llms.txt@docs llmstxt
andersonlimahw/lemon-ai-hub · docs/llms.txt
git:20260614.5160142 · audit A · 14 stars
a11y-audit skill
andersonlimahw/lemon-ai-hub · plugins/a11y-audit/SKILL.md · WCAG 2.1 AA/AAA accessibility audit for web components, pages, and apps. Detects contrast failures, missing ARIA…
git:20260614.2f67b30 · audit A · 14 stars
a11y-guardian skill
andersonlimahw/lemon-ai-hub · plugins/a11y-guardian/SKILL.md · Continuous accessibility monitoring. Runs WCAG 2.1 audit on every PR, blocks merge on CRITICAL violations, and posts…
git:20260801.d400e34 · audit A · 14 stars
a11y-guardian skill
andersonlimahw/lemon-ai-hub · plugins/a11y-guardian/skills/a11y-guardian/SKILL.md · Continuous accessibility monitoring. Runs WCAG 2.1 audit on every PR, blocks merge on CRITICAL violations, and posts…
git:20260801.d400e34 · audit A · 14 stars
ads skill
andersonlimahw/lemon-ai-hub · plugins/ads/SKILL.md · When the user wants help with paid advertising campaigns on Google Ads, Meta (Facebook/Instagram), LinkedIn, Twitter/X…
v2.2.0 · audit A · 14 stars
agent-sdk-dev skill
andersonlimahw/lemon-ai-hub · plugins/agent-sdk-dev/SKILL.md · agent-sdk-dev plugin. Routes to its sub-skills: agent-sdk-dev, agent-sdk-verifier-py. Use when the task matches any of…
git:20260801.d400e34 · audit A · 14 stars
agent-sdk-dev skill
andersonlimahw/lemon-ai-hub · plugins/agent-sdk-dev/skills/agent-sdk-dev/SKILL.md · A comprehensive skill for creating new Agent SDK applications. Guides through creating TypeScript or Python SDK apps…
git:20260709.a9f43d7 · audit A · 14 stars
agent-sdk-verifier-py skill
andersonlimahw/lemon-ai-hub · plugins/agent-sdk-dev/skills/agent-sdk-verifier-py/SKILL.md · Thoroughly verifies Python Agent SDK applications for correct setup and best practices.
git:20260709.a9f43d7 · audit A · 14 stars
AGENTS.md@plugins/agentic-value-loops agents
andersonlimahw/lemon-ai-hub · plugins/agentic-value-loops/AGENTS.md
git:20260614.f27dfea · audit A · 14 stars
agentic-value-loops skill
andersonlimahw/lemon-ai-hub · plugins/agentic-value-loops/SKILL.md · agentic-value-loops plugin. Routes to its sub-skills: ai-tuning-loop, documentation-sync-loop…
git:20260801.d400e34 · audit A · 14 stars

Every file in andersonlimahw/lemon-ai-hub

Browse by kind, by grade A, or by owner.