llm-eval skillA
llm-eval is agent-read markdown (skill) from yuecao365/offercome: 大模型与 Agent 评测深挖:评测集与真值、LLM-as-judge 校准、轨迹指标、回归门禁、去污染、红队。.
Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.
What the file says
## 面试官在意什么 评测是 2026 年 Agent 与大模型岗面试的硬门:没有评测闭环的项目一律按 demo 处理。面试官问的不是"用了什么框架",而是评测集从哪来、真值怎么定、judge 怎么校准、噪声底线多少、失败怎么回灌到下一版、门禁怎么进 CI。Agent 任务还要多一层:同样的结果可能走了 3 步或 30 步,只看成功率不够,要看轨迹。 ## 项目 / 实习怎么深挖 简历上出现下面这类经历时从哪里切、追什么。追到候选人能说出机制、数字的来源与一次真实的故障或取舍才算实;只有框架名与结论、说不出自己那一段的,记为危险信号。通用的追问方法见 project-deep-dive。 - 简历出现评测集 / benchmark → 追来源、大小、真值怎么定、有没有与训练或提示示例重叠、怎么维护 - 简历出现 LLM-as-judge → 追 judge 用哪个模型、和人工标注的一致性多少、位置与长度偏差怎么控 - 简历出现"成功率 / 准确率提升 X%" → 追跑了几次、噪声底线多少、样本数、有没有对照 - 简历出现回归 / 门禁 → 追门禁指标与阈值、拦过什么坏改动、误拦多少 - 简历出现 agent 评测 → 追轨迹级指标有哪些、评测环境怎么搭、模拟对手怎么做 - 简历出现红队 / 安全评测 → 追攻击集从哪来、覆盖哪些注入路径、拦截率与误拦率 ## 常见失守与危险信号 - 评测集构建与真值:从来没有评测集;真值就是"我觉得对";评测集与提示示例重叠 - LLM-as-judge 校准:judge 用的就是被评估的模型且没校验过一致性;不知道位置偏差与长度偏好 - 轨迹级指标:只有一个成功率数字;不知道步数、无效工具调用率、预算触顶率 - 噪声与统计:没跑过两次基线看噪声;一场翻转就当结论 - 回归门禁:改完不回归;门禁只有人工看 - 去污染:分数涨了不查是不是污染;不知道 n-gram 重叠与时间切分 - 失败回灌:失败修完没有对应的回归用例;同一类失败反复出现 - 红队与安全评测:只测正常输入;把"系统提示里写了不要"当防护 ## 常考主题清单 只列名字、阶梯与答实的标志,作"问到哪一层算实"的参考;问哪些、问几道由这份 JD 与这份简历定,不是配额。 ### 评测:轨迹、评测集与 judge …
Read the whole file at its exact version.
How to install
mdr add yuecao365/offercome/llm-eval@git:20260920.a4843f5mdr add yuecao365/offercome/llm-eval@sha256:78d4ec56496c71a9Pin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.
[](https://markdownregistry.com/a/art_tbk6xmwybzok6c5t)
1 badge views in 30 days
Versions
Audit of the latest version
- pass: Frontmatter block present
- pass: Frontmatter declares a name
- pass: Frontmatter declares a description
- pass: Size between 200 bytes and 200 KB (7112 bytes)
- pass: No zero-width or bidi control characters
- pass: No instruction hidden inside an HTML comment
- pass: No link to an exfiltration or paste host
- pass: No credential-shaped string
- pass: No instruction to send local credentials anywhere
- pass: No text hidden with inline styles
- pass: No prompt-injection phrasing
- pass: No curl or wget piped into a shell
- pass: No recursive delete of root, home or parent
- pass: No instruction to read or print local credentials
- pass: No base64 blob over 200 characters
- pass: No link to a raw IP address
- pass: No script tag
Source
yuecao365/offercome · 23 stars · license MIT · pushed 2026-09-23 · branch main
API
GET https://markdownregistry.com/api/v1/artifacts/art_tbk6xmwybzok6c5t GET https://markdownregistry.com/api/v1/resolve?ref=yuecao365/offercome/llm-eval GET https://markdownregistry.com/api/v1/blob/78d4ec56496c71a923aafb5b5550fb2cb0f2d7c4a5a49d1bbb861d832eb10933
Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.