Home / yuecao365 / offercome · src/lib/mock-interviews/skills/llm-inference-serving/SKILL.md · GitHub

llm-inference-serving skillA

llm-inference-serving is agent-read markdown (skill) from yuecao365/offercome: 推理服务深挖:投机解码、量化与算子、分布式推理与 PD 分离、MoE 服务、前缀缓存、约束解码、容量。.

Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified against the hash recorded here before it reaches your agent. The deterministic audit below grades the latest version, and the same file always earns the same grade.

What the file says

## 面试官在意什么

这本是 ai-infra 的推理那一段往下挖:把一个模型的吞吐与延迟再抬一档要动哪些机制。面试官要算账(接受率、并行通信量、KV 传输带宽、量化后的精度差)和要故障(PD 分离后 TTFT 反而变差、MoE 专家不均、量化后某类请求崩)。应用岗读它是为了知道自己的 agent 每回合的首字延迟和缓存命中从哪来。

## 项目 / 实习怎么深挖

简历上出现下面这类经历时从哪里切、追什么。追到候选人能说出机制、数字的来源与一次真实的故障或取舍才算实;只有框架名与结论、说不出自己那一段的,记为危险信号。通用的追问方法见 project-deep-dive。

- 简历出现投机解码 → 追草稿怎么产生、接受率、什么分布下加速消失、和 batch 的冲突
- 简历出现量化 → 追方法与量化对象、精度回归、kernel 来源、实测吞吐
- 简历出现多卡 / 多机推理 → 追并行方式、通信占比、PD 分离有没有做、掉卡表现
- 简历出现 MoE 推理 → 追专家并行与 all-to-all、负载不均、显存放法
- 简历出现前缀缓存 / 多轮 → 追命中率、驱逐策略、什么请求模式命中低
- 简历出现结构化输出 / 约束解码 → 追怎么实现、对吞吐的影响、和采样的冲突

## 常见失守与危险信号

- 投机解码:只会"小模型猜";说不出接受率决定加速比;不知道 batch 大时收益消失
- 量化与算子优化:量化只知道"变小变快";不做精度回归;写 kernel 不看 profiler
- 分布式推理与并行:张量并行说不出通信在哪;不知道 PD 分离解决什么、代价是什么
- 长上下文与显存:只答"加显存";不知道 offload 与 KV 压缩的取舍
- PD 分离与 KV 传输:不知道 KV 怎么在实例间传、带宽要多少;不知道什么规模才值得分
- MoE 服务:只答"选专家";不谈 all-to-all 与负载不均;不知道专家怎么放显存
- 前缀缓存与多轮:只答"缓存前缀";说不出命中条件;不知道系统提示放前面的意义
- 约束解码与采样:不知道结构化输出在推理侧怎么实现;不知道它对吞吐的影响
- 容量规划:不会从 SLO 反推卡数;没有峰谷弹性方案

## 常考主题清单

只列名字、阶梯与答实的标志,作"问到哪一层算实"的参考;问哪些、问几道由这份 JD 与这份简历定,不是配额。

### 投机解码
- 阶梯:为什么 decode 阶段有"白算"的空间 → 草稿模型 / n-gram / 自草稿(Medusa、EAGLE 类)各怎么产生候选,验证怎么一次前向完成 → 加速比由接受率与草稿长度决定,什么分布下接受率低;batch 大时为什么收益消失 → 验证的显存与算力成本、和 continuous batching 的冲突、什么业务值得开
…

Read the whole file at its exact version.

How to install

Latest version
mdr add yuecao365/offercome/llm-inference-serving@git:20260920.a4843f5
Exact content
mdr add yuecao365/offercome/llm-inference-serving@sha256:4e99f4872715f3ce

Pin to a label to follow the author's releases, or to a sha256 to freeze the exact bytes forever. Either way the resolved hash is written to mdr.lock, and mdr install reproduces it on any machine.

Badge

mdr badge

[![mdr](https://markdownregistry.com/badge/art_fsj2mvvggrkutje2.svg)](https://markdownregistry.com/a/art_fsj2mvvggrkutje2)

1 badge views in 30 days

Versions

versioncommittedcommitsizeaudit
git:20260920.a4843f5 latest2026-09-20 a4843f5 7,843 BA view

Audit of the latest version

A  17 of 17 checks passed. Deterministic, no model, same answer every run.
  • pass: Frontmatter block present
  • pass: Frontmatter declares a name
  • pass: Frontmatter declares a description
  • pass: Size between 200 bytes and 200 KB (7843 bytes)
  • pass: No zero-width or bidi control characters
  • pass: No instruction hidden inside an HTML comment
  • pass: No link to an exfiltration or paste host
  • pass: No credential-shaped string
  • pass: No instruction to send local credentials anywhere
  • pass: No text hidden with inline styles
  • pass: No prompt-injection phrasing
  • pass: No curl or wget piped into a shell
  • pass: No recursive delete of root, home or parent
  • pass: No instruction to read or print local credentials
  • pass: No base64 blob over 200 characters
  • pass: No link to a raw IP address
  • pass: No script tag

Source

GitHub

yuecao365/offercome · 23 stars · license MIT · pushed 2026-09-23 · branch main

API

GET https://markdownregistry.com/api/v1/artifacts/art_fsj2mvvggrkutje2
GET https://markdownregistry.com/api/v1/resolve?ref=yuecao365/offercome/llm-inference-serving
GET https://markdownregistry.com/api/v1/blob/4e99f4872715f3cef6a3d5d6a9482c9522dfce3e0aa35a305bb21d69e289266d

Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.

More from yuecao365/offercome

AGENTS.md agents
yuecao365/offercome · AGENTS.md
git:20260921.abf64ee · audit A · 23 stars
CLAUDE.md claude
yuecao365/offercome · CLAUDE.md
git:20260803.e553cde · audit B · 23 stars
agent-runtime skill
yuecao365/offercome · src/lib/mock-interviews/skills/agent-runtime/SKILL.md · Agent 运行时深挖:循环与事件、工具协议与沙箱、子 agent、预算终止、恢复、输出契约。
git:20260920.a4843f5 · audit A · 23 stars
ai-agent skill
yuecao365/offercome · src/lib/mock-interviews/skills/ai-agent/SKILL.md · Agent 开发与运行时怎么面:循环与工具、RAG、上下文与记忆、评测、安全、成本。Agent 与 LLM 应用岗读。
git:20260920.a4843f5 · audit A · 23 stars
ai-algorithm skill
yuecao365/offercome · src/lib/mock-interviews/skills/ai-algorithm/SKILL.md · 大模型算法怎么面:Transformer、预训练、后训练与对齐、RL、微调、Embedding、评测。LLM 算法岗读。
git:20260920.a4843f5 · audit A · 23 stars
ai-app-testing skill
yuecao365/offercome · src/lib/mock-interviews/skills/ai-app-testing/SKILL.md · AI 应用测试深挖:非确定性输出、幻觉与 RAG 评估、Agent 链路审计、安全对抗、回归门禁、AI 生成用例。
git:20260920.a4843f5 · audit A · 23 stars
ai-infra skill
yuecao365/offercome · src/lib/mock-interviews/skills/ai-infra/SKILL.md · 大模型推理与训练基础设施怎么面:KV cache、调度、并行、量化算子、服务指标、训练集群。AI Infra 岗读。
git:20260920.a4843f5 · audit A · 23 stars
algorithm skill
yuecao365/offercome · src/lib/mock-interviews/skills/algorithm/SKILL.md · 算法与机器学习怎么面:ML 基础、深度学习、特征、AB 实验、部署监控、落地。算法与 ML 工程岗读。
git:20260920.a4843f5 · audit A · 23 stars
android skill
yuecao365/offercome · src/lib/mock-interviews/skills/android/SKILL.md · Android 怎么面:生命周期、Compose、协程与 Flow、性能、启动。JD 点名 Android 时读。
git:20260920.a4843f5 · audit A · 23 stars
backend skill
yuecao365/offercome · src/lib/mock-interviews/skills/backend/SKILL.md · 后端怎么面(栈无关):缓存、消息队列、接口、可靠性、可观测、容量、发布。服务端岗读。
git:20260920.a4843f5 · audit A · 23 stars
cpp skill
yuecao365/offercome · src/lib/mock-interviews/skills/cpp/SKILL.md · C++ 后端怎么面:内存模型、RAII、STL 性能、多线程与原子、IO 模型、现代 C++。JD 点名 C++ 时读。
git:20260920.a4843f5 · audit A · 23 stars
cs-fundamentals skill
yuecao365/offercome · src/lib/mock-interviews/skills/cs-fundamentals/SKILL.md · 计算机基础怎么面:操作系统、网络、数据结构与算法、数据库原理。技术岗校招兜底。
git:20260920.a4843f5 · audit A · 23 stars

Every file in yuecao365/offercome

Browse by kind, by grade A, or by owner.