moonlight-lupin/agent-skills · mlops/model-compare/SKILL.md

model-compare skillA

model-compare is agent-read markdown (skill) from moonlight-lupin/agent-skills: Blind side-by-side multi-model comparison. Send one prompt to 2-4 models simultaneously, present responses anonymously (Model A / B / C / D), let the user pick a winner, then reveal identities and show which model won. Supports custom evaluation criteria, synthesis of responses, and vote history logging. Trigger when the user says "compare models", "test these models", "which model is better for", "A/B test", "blind comparison", "model evaluation", or wants to see how different AI models handle .

Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.

How to install

Latest version
mdr add moonlight-lupin/agent-skills/model-compare@v1.2.0
Exact content
mdr add moonlight-lupin/agent-skills/model-compare@sha256:497568503a08d50d

Pin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.

Badge

mdr badge

[![mdr](https://markdownregistry.com/badge/art_nstvi4lokwxhbg65.svg)](https://markdownregistry.com/a/art_nstvi4lokwxhbg65)

0 badge views in 30 days

Versions

versioncommittedcommitsizeaudit
v1.2.0 latest2026-08-29 63d3bc1 36,972 BA view · diff
v1.2.02026-08-29 8edd63a 36,562 BA view · diff
v1.1.02026-08-19 46eea66 34,547 BA view · diff
v1.1.02026-08-14 2439073 34,541 BA view · diff
v1.0.02026-08-13 73a0ca7 23,537 BA view

Audit of the latest version

A  17 of 17 checks passed. Deterministic, no model, same answer every run.
  • pass: Frontmatter block present
  • pass: Frontmatter declares a name
  • pass: Frontmatter declares a description
  • pass: Size between 200 bytes and 200 KB (36972 bytes)
  • pass: No zero-width or bidi control characters
  • pass: No instruction hidden inside an HTML comment
  • pass: No link to an exfiltration or paste host
  • pass: No credential-shaped string
  • pass: No instruction to send local credentials anywhere
  • pass: No text hidden with inline styles
  • pass: No prompt-injection phrasing
  • pass: No curl or wget piped into a shell
  • pass: No recursive delete of root, home or parent
  • pass: No instruction to read or print local credentials
  • pass: No base64 blob over 200 characters
  • pass: No link to a raw IP address
  • pass: No script tag

Source

GitHub

moonlight-lupin/agent-skills · 56 stars · license MIT · pushed 2026-09-07 · branch main

API

GET https://markdownregistry.com/api/v1/artifacts/art_nstvi4lokwxhbg65
GET https://markdownregistry.com/api/v1/resolve?ref=moonlight-lupin/agent-skills/model-compare
GET https://markdownregistry.com/api/v1/blob/497568503a08d50d5ae247d8eb308c45865c1aaeb3694a348855c84ff74b4205