ai-app-testing skillA
ai-app-testing is agent-read markdown (skill) from yuecao365/offercome: AI 应用测试深挖:非确定性输出、幻觉与 RAG 评估、Agent 链路审计、安全对抗、回归门禁、AI 生成用例。.
Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.
What the file says
## 面试官在意什么 这本是测试开发在 2026 年的新分水岭:被测对象是大模型应用与 Agent,输出不确定、错误不报错、行为随上下文变。面试官要候选人说清怎么给非确定性输出定"对"、怎么测幻觉与检索质量、怎么审计 Agent 的决策链路、怎么做安全对抗、怎么把这些做成回归门禁进流水线;同时会问候选人自己怎么用 AI 生成用例、做意图驱动测试,以及 MCP 工具的测法。评测方法论的深挖在 llm-eval 包里,这里是测试岗视角的落地。 ## 项目 / 实习怎么深挖 简历上出现下面这类经历时从哪里切、追什么。追到候选人能说出机制、数字的来源与一次真实的故障或取舍才算实;只有框架名与结论、说不出自己那一段的,记为危险信号。通用的追问方法见 project-deep-dive。 - 简历出现大模型应用测试 → 追怎么定义"对"、评测集从哪来、多大、通过率怎么算、非确定性怎么处理 - 简历出现幻觉 / RAG 评估 → 追怎么判幻觉、检索召回率怎么测、失败案例归因到哪一层 - 简历出现 Agent 测试 → 追测了哪些工具调用路径、决策链路怎么审计、死循环与越权怎么发现 - 简历出现安全测试 / 红队 → 追攻击集从哪来、注入路径覆盖、拦截率与误拦 - 简历出现评测流水线 / 回归门禁 → 追门禁指标与阈值、拦过什么、误拦多少、跑一次多久多贵 - 简历出现 AI 生成用例 / 意图驱动测试 → 追生成的用例怎么验、覆盖了什么人工没覆盖的、误报率 - 简历出现 MCP / Skill 工程化 → 追工具 schema 怎么测、权限怎么验、和普通接口测试的差别 ## 常见失守与危险信号 - 非确定性输出的测法:还在用精确匹配;不知道多次采样与阈值;把偶发失败当 flaky 忽略 - 幻觉与 RAG 评估:判幻觉靠人看;分不清检索失败与生成失败;没有检索评测集 - Agent 链路审计:只测最终结果;不看工具调用序列;死循环与越权靶不到 - 安全对抗测试:只测正常输入;把"系统提示里写了不要"当防护;没有间接注入用例 - 回归门禁与评测流水线:改完不回归;门禁只有人工看;跑一次的成本不知道 - 用 AI 做测试:生成的用例不验直接用;不知道生成用例的误报率;意图驱动只是口号 - MCP 与工具的测法:把工具当普通接口测;不测权限与错误回传;不测描述对模型选择的影响 - 数据集与标注治理:评测集没版本;标注一致性没量过;测试数据里有生产隐私 ## 常考主题清单 只列名字、阶梯与答实的标志,作"问到哪一层算实"的参考;问哪些、问几道由这份 JD 与这份简历定,不是配额。 ### AI 时代的测试 …
Read the whole file at its exact version.
How to install
mdr add yuecao365/offercome/ai-app-testing@git:20260920.a4843f5mdr add yuecao365/offercome/ai-app-testing@sha256:e358f40ff9f38cf3Pin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.
[](https://markdownregistry.com/a/art_ah2nnxo4cdsjq4l3)
1 badge views in 30 days
Versions
Audit of the latest version
- pass: Frontmatter block present
- pass: Frontmatter declares a name
- pass: Frontmatter declares a description
- pass: Size between 200 bytes and 200 KB (7225 bytes)
- pass: No zero-width or bidi control characters
- pass: No instruction hidden inside an HTML comment
- pass: No link to an exfiltration or paste host
- pass: No credential-shaped string
- pass: No instruction to send local credentials anywhere
- pass: No text hidden with inline styles
- pass: No prompt-injection phrasing
- pass: No curl or wget piped into a shell
- pass: No recursive delete of root, home or parent
- pass: No instruction to read or print local credentials
- pass: No base64 blob over 200 characters
- pass: No link to a raw IP address
- pass: No script tag
Source
yuecao365/offercome · 23 stars · license MIT · pushed 2026-09-23 · branch main
API
GET https://markdownregistry.com/api/v1/artifacts/art_ah2nnxo4cdsjq4l3 GET https://markdownregistry.com/api/v1/resolve?ref=yuecao365/offercome/ai-app-testing GET https://markdownregistry.com/api/v1/blob/e358f40ff9f38cf31628c95f1ffca2b743bff710e7a8a0a6f3d6e4606e5a84c2
Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.