Home / yuecao365 / offercome · src/lib/mock-interviews/skills/llm-eval/SKILL.md · GitHub

llm-eval skillA

llm-eval is agent-read markdown (skill) from yuecao365/offercome: 大模型与 Agent 评测深挖:评测集与真值、LLM-as-judge 校准、轨迹指标、回归门禁、去污染、红队。.

Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.

What the file says

## 面试官在意什么

评测是 2026 年 Agent 与大模型岗面试的硬门:没有评测闭环的项目一律按 demo 处理。面试官问的不是"用了什么框架",而是评测集从哪来、真值怎么定、judge 怎么校准、噪声底线多少、失败怎么回灌到下一版、门禁怎么进 CI。Agent 任务还要多一层:同样的结果可能走了 3 步或 30 步,只看成功率不够,要看轨迹。

## 项目 / 实习怎么深挖

简历上出现下面这类经历时从哪里切、追什么。追到候选人能说出机制、数字的来源与一次真实的故障或取舍才算实;只有框架名与结论、说不出自己那一段的,记为危险信号。通用的追问方法见 project-deep-dive。

- 简历出现评测集 / benchmark → 追来源、大小、真值怎么定、有没有与训练或提示示例重叠、怎么维护
- 简历出现 LLM-as-judge → 追 judge 用哪个模型、和人工标注的一致性多少、位置与长度偏差怎么控
- 简历出现"成功率 / 准确率提升 X%" → 追跑了几次、噪声底线多少、样本数、有没有对照
- 简历出现回归 / 门禁 → 追门禁指标与阈值、拦过什么坏改动、误拦多少
- 简历出现 agent 评测 → 追轨迹级指标有哪些、评测环境怎么搭、模拟对手怎么做
- 简历出现红队 / 安全评测 → 追攻击集从哪来、覆盖哪些注入路径、拦截率与误拦率

## 常见失守与危险信号

- 评测集构建与真值:从来没有评测集;真值就是"我觉得对";评测集与提示示例重叠
- LLM-as-judge 校准:judge 用的就是被评估的模型且没校验过一致性;不知道位置偏差与长度偏好
- 轨迹级指标:只有一个成功率数字;不知道步数、无效工具调用率、预算触顶率
- 噪声与统计:没跑过两次基线看噪声;一场翻转就当结论
- 回归门禁:改完不回归;门禁只有人工看
- 去污染:分数涨了不查是不是污染;不知道 n-gram 重叠与时间切分
- 失败回灌:失败修完没有对应的回归用例;同一类失败反复出现
- 红队与安全评测:只测正常输入;把"系统提示里写了不要"当防护

## 常考主题清单

只列名字、阶梯与答实的标志,作"问到哪一层算实"的参考;问哪些、问几道由这份 JD 与这份简历定,不是配额。

### 评测:轨迹、评测集与 judge
…

Read the whole file at its exact version.

How to install

Latest version
mdr add yuecao365/offercome/llm-eval@git:20260920.a4843f5
Exact content
mdr add yuecao365/offercome/llm-eval@sha256:78d4ec56496c71a9

Pin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.

Badge

mdr badge

[![mdr](https://markdownregistry.com/badge/art_tbk6xmwybzok6c5t.svg)](https://markdownregistry.com/a/art_tbk6xmwybzok6c5t)

1 badge views in 30 days

Versions

versioncommittedcommitsizeaudit
git:20260920.a4843f5 latest2026-09-20 a4843f5 7,112 BA view

Audit of the latest version

A  17 of 17 checks passed. Deterministic, no model, same answer every run.
  • pass: Frontmatter block present
  • pass: Frontmatter declares a name
  • pass: Frontmatter declares a description
  • pass: Size between 200 bytes and 200 KB (7112 bytes)
  • pass: No zero-width or bidi control characters
  • pass: No instruction hidden inside an HTML comment
  • pass: No link to an exfiltration or paste host
  • pass: No credential-shaped string
  • pass: No instruction to send local credentials anywhere
  • pass: No text hidden with inline styles
  • pass: No prompt-injection phrasing
  • pass: No curl or wget piped into a shell
  • pass: No recursive delete of root, home or parent
  • pass: No instruction to read or print local credentials
  • pass: No base64 blob over 200 characters
  • pass: No link to a raw IP address
  • pass: No script tag

Source

GitHub

yuecao365/offercome · 23 stars · license MIT · pushed 2026-09-23 · branch main

API

GET https://markdownregistry.com/api/v1/artifacts/art_tbk6xmwybzok6c5t
GET https://markdownregistry.com/api/v1/resolve?ref=yuecao365/offercome/llm-eval
GET https://markdownregistry.com/api/v1/blob/78d4ec56496c71a923aafb5b5550fb2cb0f2d7c4a5a49d1bbb861d832eb10933

Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.

More from yuecao365/offercome

AGENTS.md agents
yuecao365/offercome · AGENTS.md
git:20260921.abf64ee · audit A · 23 stars
CLAUDE.md claude
yuecao365/offercome · CLAUDE.md
git:20260803.e553cde · audit B · 23 stars
agent-runtime skill
yuecao365/offercome · src/lib/mock-interviews/skills/agent-runtime/SKILL.md · Agent 运行时深挖:循环与事件、工具协议与沙箱、子 agent、预算终止、恢复、输出契约。
git:20260920.a4843f5 · audit A · 23 stars
ai-agent skill
yuecao365/offercome · src/lib/mock-interviews/skills/ai-agent/SKILL.md · Agent 开发与运行时怎么面:循环与工具、RAG、上下文与记忆、评测、安全、成本。Agent 与 LLM 应用岗读。
git:20260920.a4843f5 · audit A · 23 stars
ai-algorithm skill
yuecao365/offercome · src/lib/mock-interviews/skills/ai-algorithm/SKILL.md · 大模型算法怎么面:Transformer、预训练、后训练与对齐、RL、微调、Embedding、评测。LLM 算法岗读。
git:20260920.a4843f5 · audit A · 23 stars
ai-app-testing skill
yuecao365/offercome · src/lib/mock-interviews/skills/ai-app-testing/SKILL.md · AI 应用测试深挖:非确定性输出、幻觉与 RAG 评估、Agent 链路审计、安全对抗、回归门禁、AI 生成用例。
git:20260920.a4843f5 · audit A · 23 stars
ai-infra skill
yuecao365/offercome · src/lib/mock-interviews/skills/ai-infra/SKILL.md · 大模型推理与训练基础设施怎么面:KV cache、调度、并行、量化算子、服务指标、训练集群。AI Infra 岗读。
git:20260920.a4843f5 · audit A · 23 stars
algorithm skill
yuecao365/offercome · src/lib/mock-interviews/skills/algorithm/SKILL.md · 算法与机器学习怎么面:ML 基础、深度学习、特征、AB 实验、部署监控、落地。算法与 ML 工程岗读。
git:20260920.a4843f5 · audit A · 23 stars
android skill
yuecao365/offercome · src/lib/mock-interviews/skills/android/SKILL.md · Android 怎么面:生命周期、Compose、协程与 Flow、性能、启动。JD 点名 Android 时读。
git:20260920.a4843f5 · audit A · 23 stars
backend skill
yuecao365/offercome · src/lib/mock-interviews/skills/backend/SKILL.md · 后端怎么面(栈无关):缓存、消息队列、接口、可靠性、可观测、容量、发布。服务端岗读。
git:20260920.a4843f5 · audit A · 23 stars
cpp skill
yuecao365/offercome · src/lib/mock-interviews/skills/cpp/SKILL.md · C++ 后端怎么面:内存模型、RAII、STL 性能、多线程与原子、IO 模型、现代 C++。JD 点名 C++ 时读。
git:20260920.a4843f5 · audit A · 23 stars
cs-fundamentals skill
yuecao365/offercome · src/lib/mock-interviews/skills/cs-fundamentals/SKILL.md · 计算机基础怎么面:操作系统、网络、数据结构与算法、数据库原理。技术岗校招兜底。
git:20260920.a4843f5 · audit A · 23 stars

Every file in yuecao365/offercome

Browse by kind, by grade A, or by owner.