Build an automated evaluation pipeline that tests a Python LLM application end-to-end with pixie test — real code paths, real LLM calls, instrumented external data — and scores outputs with evaluators instead of assertions. Use when adding evals to a Python AI app.
mdr add marielynneblock/arcanum-artifex/eval-driven-dev@v0.8.4mdr add marielynneblock/arcanum-artifex/eval-driven-dev@sha256:6a6b1ee1d1bf0a29[](https://markdownregistry.com/a/art_vljfatud6idrf4pk)
0 badge views in 30 days
| version | committed | commit | size | audit | |
|---|---|---|---|---|---|
| v0.8.4 latest | 2026-07-10 | 0115f49 | 17,688 B | A | view · diff |
| v0.8.4 | 2026-07-10 | ffb3400 | 17,688 B | A | view · diff |
| v0.8.4 | 2026-05-15 | b861afb | 17,411 B | B | view · diff |
| v0.8.4 | 2026-05-10 | f26bcec | 17,830 B | A | view |
marielynneblock/arcanum-artifex · 4 stars · license none · pushed 2026-09-05 · branch main
GET https://markdownregistry.com/api/v1/artifacts/art_vljfatud6idrf4pk GET https://markdownregistry.com/api/v1/resolve?ref=marielynneblock/arcanum-artifex/eval-driven-dev GET https://markdownregistry.com/api/v1/blob/6a6b1ee1d1bf0a29f8e2034982e33e9426d5941fda02dca6afd1b8dd64f7e598