Home / yuecao365 / offercome · src/lib/mock-interviews/skills/ai-app-testing/SKILL.md · GitHub

ai-app-testing skillA

ai-app-testing is agent-read markdown (skill) from yuecao365/offercome: AI 应用测试深挖:非确定性输出、幻觉与 RAG 评估、Agent 链路审计、安全对抗、回归门禁、AI 生成用例。.

Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.

What the file says

## 面试官在意什么

这本是测试开发在 2026 年的新分水岭:被测对象是大模型应用与 Agent,输出不确定、错误不报错、行为随上下文变。面试官要候选人说清怎么给非确定性输出定"对"、怎么测幻觉与检索质量、怎么审计 Agent 的决策链路、怎么做安全对抗、怎么把这些做成回归门禁进流水线;同时会问候选人自己怎么用 AI 生成用例、做意图驱动测试,以及 MCP 工具的测法。评测方法论的深挖在 llm-eval 包里,这里是测试岗视角的落地。

## 项目 / 实习怎么深挖

简历上出现下面这类经历时从哪里切、追什么。追到候选人能说出机制、数字的来源与一次真实的故障或取舍才算实;只有框架名与结论、说不出自己那一段的,记为危险信号。通用的追问方法见 project-deep-dive。

- 简历出现大模型应用测试 → 追怎么定义"对"、评测集从哪来、多大、通过率怎么算、非确定性怎么处理
- 简历出现幻觉 / RAG 评估 → 追怎么判幻觉、检索召回率怎么测、失败案例归因到哪一层
- 简历出现 Agent 测试 → 追测了哪些工具调用路径、决策链路怎么审计、死循环与越权怎么发现
- 简历出现安全测试 / 红队 → 追攻击集从哪来、注入路径覆盖、拦截率与误拦
- 简历出现评测流水线 / 回归门禁 → 追门禁指标与阈值、拦过什么、误拦多少、跑一次多久多贵
- 简历出现 AI 生成用例 / 意图驱动测试 → 追生成的用例怎么验、覆盖了什么人工没覆盖的、误报率
- 简历出现 MCP / Skill 工程化 → 追工具 schema 怎么测、权限怎么验、和普通接口测试的差别

## 常见失守与危险信号

- 非确定性输出的测法:还在用精确匹配;不知道多次采样与阈值;把偶发失败当 flaky 忽略
- 幻觉与 RAG 评估:判幻觉靠人看;分不清检索失败与生成失败;没有检索评测集
- Agent 链路审计:只测最终结果;不看工具调用序列;死循环与越权靶不到
- 安全对抗测试:只测正常输入;把"系统提示里写了不要"当防护;没有间接注入用例
- 回归门禁与评测流水线:改完不回归;门禁只有人工看;跑一次的成本不知道
- 用 AI 做测试:生成的用例不验直接用;不知道生成用例的误报率;意图驱动只是口号
- MCP 与工具的测法:把工具当普通接口测;不测权限与错误回传;不测描述对模型选择的影响
- 数据集与标注治理:评测集没版本;标注一致性没量过;测试数据里有生产隐私

## 常考主题清单

只列名字、阶梯与答实的标志,作"问到哪一层算实"的参考;问哪些、问几道由这份 JD 与这份简历定,不是配额。

### AI 时代的测试
…

Read the whole file at its exact version.

How to install

Latest version
mdr add yuecao365/offercome/ai-app-testing@git:20260920.a4843f5
Exact content
mdr add yuecao365/offercome/ai-app-testing@sha256:e358f40ff9f38cf3

Pin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.

Badge

mdr badge

[![mdr](https://markdownregistry.com/badge/art_ah2nnxo4cdsjq4l3.svg)](https://markdownregistry.com/a/art_ah2nnxo4cdsjq4l3)

1 badge views in 30 days

Versions

versioncommittedcommitsizeaudit
git:20260920.a4843f5 latest2026-09-20 a4843f5 7,225 BA view

Audit of the latest version

A  17 of 17 checks passed. Deterministic, no model, same answer every run.
  • pass: Frontmatter block present
  • pass: Frontmatter declares a name
  • pass: Frontmatter declares a description
  • pass: Size between 200 bytes and 200 KB (7225 bytes)
  • pass: No zero-width or bidi control characters
  • pass: No instruction hidden inside an HTML comment
  • pass: No link to an exfiltration or paste host
  • pass: No credential-shaped string
  • pass: No instruction to send local credentials anywhere
  • pass: No text hidden with inline styles
  • pass: No prompt-injection phrasing
  • pass: No curl or wget piped into a shell
  • pass: No recursive delete of root, home or parent
  • pass: No instruction to read or print local credentials
  • pass: No base64 blob over 200 characters
  • pass: No link to a raw IP address
  • pass: No script tag

Source

GitHub

yuecao365/offercome · 23 stars · license MIT · pushed 2026-09-23 · branch main

API

GET https://markdownregistry.com/api/v1/artifacts/art_ah2nnxo4cdsjq4l3
GET https://markdownregistry.com/api/v1/resolve?ref=yuecao365/offercome/ai-app-testing
GET https://markdownregistry.com/api/v1/blob/e358f40ff9f38cf31628c95f1ffca2b743bff710e7a8a0a6f3d6e4606e5a84c2

Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.

More from yuecao365/offercome

AGENTS.md agents
yuecao365/offercome · AGENTS.md
git:20260921.abf64ee · audit A · 23 stars
CLAUDE.md claude
yuecao365/offercome · CLAUDE.md
git:20260803.e553cde · audit B · 23 stars
agent-runtime skill
yuecao365/offercome · src/lib/mock-interviews/skills/agent-runtime/SKILL.md · Agent 运行时深挖:循环与事件、工具协议与沙箱、子 agent、预算终止、恢复、输出契约。
git:20260920.a4843f5 · audit A · 23 stars
ai-agent skill
yuecao365/offercome · src/lib/mock-interviews/skills/ai-agent/SKILL.md · Agent 开发与运行时怎么面:循环与工具、RAG、上下文与记忆、评测、安全、成本。Agent 与 LLM 应用岗读。
git:20260920.a4843f5 · audit A · 23 stars
ai-algorithm skill
yuecao365/offercome · src/lib/mock-interviews/skills/ai-algorithm/SKILL.md · 大模型算法怎么面:Transformer、预训练、后训练与对齐、RL、微调、Embedding、评测。LLM 算法岗读。
git:20260920.a4843f5 · audit A · 23 stars
ai-infra skill
yuecao365/offercome · src/lib/mock-interviews/skills/ai-infra/SKILL.md · 大模型推理与训练基础设施怎么面:KV cache、调度、并行、量化算子、服务指标、训练集群。AI Infra 岗读。
git:20260920.a4843f5 · audit A · 23 stars
algorithm skill
yuecao365/offercome · src/lib/mock-interviews/skills/algorithm/SKILL.md · 算法与机器学习怎么面:ML 基础、深度学习、特征、AB 实验、部署监控、落地。算法与 ML 工程岗读。
git:20260920.a4843f5 · audit A · 23 stars
android skill
yuecao365/offercome · src/lib/mock-interviews/skills/android/SKILL.md · Android 怎么面:生命周期、Compose、协程与 Flow、性能、启动。JD 点名 Android 时读。
git:20260920.a4843f5 · audit A · 23 stars
backend skill
yuecao365/offercome · src/lib/mock-interviews/skills/backend/SKILL.md · 后端怎么面(栈无关):缓存、消息队列、接口、可靠性、可观测、容量、发布。服务端岗读。
git:20260920.a4843f5 · audit A · 23 stars
cpp skill
yuecao365/offercome · src/lib/mock-interviews/skills/cpp/SKILL.md · C++ 后端怎么面:内存模型、RAII、STL 性能、多线程与原子、IO 模型、现代 C++。JD 点名 C++ 时读。
git:20260920.a4843f5 · audit A · 23 stars
cs-fundamentals skill
yuecao365/offercome · src/lib/mock-interviews/skills/cs-fundamentals/SKILL.md · 计算机基础怎么面:操作系统、网络、数据结构与算法、数据库原理。技术岗校招兜底。
git:20260920.a4843f5 · audit A · 23 stars
data-engineering skill
yuecao365/offercome · src/lib/mock-interviews/skills/data-engineering/SKILL.md · 数据工程怎么面:数仓建模、离线与实时链路、调度、数据质量、湖仓。数据开发与数仓岗读。
git:20260920.a4843f5 · audit A · 23 stars

Every file in yuecao365/offercome

Browse by kind, by grade A, or by owner.