feat: deepen DeepSeek technical lineage
This commit is contained in:
+10
-3
@@ -14,7 +14,7 @@
|
|||||||
| 表示、位置与残差高速公路 | 完成首版 | 81% | 真实 hidden-state / norm traces、长上下文位置外推与深层稳定性消融 |
|
| 表示、位置与残差高速公路 | 完成首版 | 81% | 真实 hidden-state / norm traces、长上下文位置外推与深层稳定性消融 |
|
||||||
| Scaling Laws | 完成首版 | 74% | 真实拟合复现、置信区间与更多模型族对照 |
|
| Scaling Laws | 完成首版 | 74% | 真实拟合复现、置信区间与更多模型族对照 |
|
||||||
| 数据工程与预训练配方 | 完成首版 | 73% | FineWeb / DCLM 逐图精读、真实去重误伤与 mixture traces |
|
| 数据工程与预训练配方 | 完成首版 | 73% | FineWeb / DCLM 逐图精读、真实去重误伤与 mixture traces |
|
||||||
| DeepSeek 专题 | 进行中 | 71% | 补 R1 / DAPO 的逐图训练轨迹与复现对照 |
|
| DeepSeek 专题 | 完成二轮 | 83% | 真实专家负载 / MLA kernel / RL 训练 traces 与独立复现 |
|
||||||
| 指令微调与人类偏好 | 完成首版 | 75% | 真实偏好分歧、RM 长度偏置与 PPO/DPO 小模型复现 |
|
| 指令微调与人类偏好 | 完成首版 | 75% | 真实偏好分歧、RM 长度偏置与 PPO/DPO 小模型复现 |
|
||||||
| 推理与测试时扩展 | 完成首版 | 76% | 真实模型采样曲线、PRM 案例与逐篇图表精读 |
|
| 推理与测试时扩展 | 完成首版 | 76% | 真实模型采样曲线、PRM 案例与逐篇图表精读 |
|
||||||
| 工具使用与长程 Agent | 完成首版 | 74% | 真实环境 traces、cross-harness 对照、Agent RL 曲线与安全案例 |
|
| 工具使用与长程 Agent | 完成首版 | 74% | 真实环境 traces、cross-harness 对照、Agent RL 曲线与安全案例 |
|
||||||
@@ -38,10 +38,10 @@
|
|||||||
- [x] 提炼参考网站的编辑设计语言。
|
- [x] 提炼参考网站的编辑设计语言。
|
||||||
- [x] 确认 `git.k1412.top` 为 Gitea/Forgejo 兼容服务且本机 HTTPS 凭据可用于既有仓库。
|
- [x] 确认 `git.k1412.top` 为 Gitea/Forgejo 兼容服务且本机 HTTPS 凭据可用于既有仓库。
|
||||||
- [x] 使用 Grok CLI 检索并形成约 95 篇一手论文的补充路线,主代理已回查关键来源。
|
- [x] 使用 Grok CLI 检索并形成约 95 篇一手论文的补充路线,主代理已回查关键来源。
|
||||||
- [x] 完成 480 篇关键论文索引,覆盖 16 个标签专题与 Kimi/DeepSeek 聚光主线。
|
- [x] 完成 486 篇关键论文索引,覆盖 16 个标签专题与 Kimi/DeepSeek 聚光主线。
|
||||||
- [x] 完成可检索、可按专题筛选的论文库页面。
|
- [x] 完成可检索、可按专题筛选的论文库页面。
|
||||||
- [x] 完成 K3、语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全十七篇首版长文。
|
- [x] 完成 K3、语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全十七篇首版长文。
|
||||||
- [x] 完成 K3 三轴架构、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 谱系、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等五十五个原创交互视图。
|
- [x] 完成 K3 三轴架构、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 四联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等五十九个原创交互视图。
|
||||||
- [x] 完成长上下文首版:五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。
|
- [x] 完成长上下文首版:五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。
|
||||||
- [x] 核验 FlashAttention、DeepSeek-V2/V3.2/V4、Kimi Linear/K3 等六份论文原文,并建立长上下文研究账本。
|
- [x] 核验 FlashAttention、DeepSeek-V2/V3.2/V4、Kimi Linear/K3 等六份论文原文,并建立长上下文研究账本。
|
||||||
- [x] 核验 Switch、ST-MoE、DeepSeekMoE、Loss-Free、V3、LatentMoE 与 K3 原文,并建立 MoE 研究账本。
|
- [x] 核验 Switch、ST-MoE、DeepSeekMoE、Loss-Free、V3、LatentMoE 与 K3 原文,并建立 MoE 研究账本。
|
||||||
@@ -142,9 +142,16 @@
|
|||||||
- [x] Astro 类型检查、生产构建、21 个页面、1145 个站内引用和 14 个跨页锚点通过;桌面 / 移动端无文档级横向溢出。
|
- [x] Astro 类型检查、生产构建、21 个页面、1145 个站内引用和 14 个跨页锚点通过;桌面 / 移动端无文档级横向溢出。
|
||||||
- [x] 表示深度与既有十三个专题共十四套本地真实 Chrome 回归全部通过。
|
- [x] 表示深度与既有十三个专题共十四套本地真实 Chrome 回归全部通过。
|
||||||
- [x] 表示深度首版以源提交 `a2c9298`、不可变镜像 `20260729T023329Z-a2c9298` 发布;NAS、VPS/Tailscale、NPM、DNS、HTTPS、证书、门户、公开 Forgejo 与十四套生产 Chrome 回归全链路通过。
|
- [x] 表示深度首版以源提交 `a2c9298`、不可变镜像 `20260729T023329Z-a2c9298` 发布;NAS、VPS/Tailscale、NPM、DNS、HTTPS、证书、门户、公开 Forgejo 与十四套生产 Chrome 回归全链路通过。
|
||||||
|
- [x] 启动 DeepSeek 二轮深读:Grok Headless 只用于发散候选问题,正式事实逐项回到 DeepSeek 论文、官方代码仓库与 K3 报告。
|
||||||
|
- [x] 建立 24 张正式问题账、10 次历史转向、4 个实验合同和 60 个一手 / 官方节点;明确 R1-Zero / R1、DAPO / Dr.GRPO、MLA RoPE cache、V3 FP8 角色与 V4/K3 状态边界。
|
||||||
|
- [x] 完成 DeepSeek 二轮正文:24 个目录、Dense→MoE→MLA→V3→R1→V3.2→V4 因果链、五条旁支、证据审计与稀疏容量—MLA 缓存—V3 协同—RL 偏差四联实验。
|
||||||
|
- [x] 论文库新增 DeepSeek-Coder/Coder-V2、ESFT、Prover-V1.5/V2 与 Engram 6 个节点,从 480 篇扩充至 486 篇。
|
||||||
|
- [x] DeepSeek 专属 Chrome 断言通过:24 张账、10 次转向、60 节点、MLA 576 元素、FP8 角色、DualPipe、MTP、GRPO/DAPO/Dr.GRPO、R1 身份、键盘 tabs 与 390px 移动端均正确响应。
|
||||||
|
- [x] Astro 类型检查、生产构建、21 个页面、1145 个站内引用和 14 个跨页锚点通过;DeepSeek 与既有十四专题共十五套本地真实 Chrome 回归全部通过。
|
||||||
|
|
||||||
## 正在进行
|
## 正在进行
|
||||||
|
|
||||||
|
- [ ] DeepSeek 三轮:真实专家负载、MLA kernel、FP8 / pipeline 与 R1-like RL traces,外加独立小模型复现。
|
||||||
- [ ] 表示、位置与残差二轮:真实 hidden-state / norm traces、长上下文位置外推复现与 mHC / AttnRes 深层稳定性消融。
|
- [ ] 表示、位置与残差二轮:真实 hidden-state / norm traces、长上下文位置外推复现与 mHC / AttnRes 深层稳定性消融。
|
||||||
- [ ] 评测安全二轮:真实 cross-harness / pass@k 复跑、Judge 元评测、动态污染与过拒案例。
|
- [ ] 评测安全二轮:真实 cross-harness / pass@k 复跑、Judge 元评测、动态污染与过拒案例。
|
||||||
- [ ] 推理服务二轮:真实 GPU kernel / workload traces、功耗与成本、跨 vLLM / SGLang / TensorRT-LLM 复现。
|
- [ ] 推理服务二轮:真实 GPU kernel / workload traces、功耗与成本、跨 vLLM / SGLang / TensorRT-LLM 复现。
|
||||||
|
|||||||
@@ -17,9 +17,10 @@
|
|||||||
- 持续进度:[PROGRESS.md](./PROGRESS.md)
|
- 持续进度:[PROGRESS.md](./PROGRESS.md)
|
||||||
- 证据与写作规范:[research/METHODOLOGY.md](./research/METHODOLOGY.md)
|
- 证据与写作规范:[research/METHODOLOGY.md](./research/METHODOLOGY.md)
|
||||||
|
|
||||||
当前里程碑包含 17 专题学习地图、480 篇关键论文索引、Kimi K3 完整导读,
|
当前里程碑包含 17 专题学习地图、486 篇关键论文索引、Kimi K3 完整导读,
|
||||||
语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 技术谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、工具使用与长程 Agent、原生多模态、训练系统、推理服务、数值优化,以及评测与安全深度专题,
|
语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 技术谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、工具使用与长程 Agent、原生多模态、训练系统、推理服务、数值优化,以及评测与安全深度专题,
|
||||||
以及 55 个覆盖核心机制的原创交互视图。
|
以及 59 个覆盖核心机制的原创交互视图。DeepSeek 二轮专题以 24 张问题账、10 次技术转向、
|
||||||
|
4 个交互实验和 60 个一手/官方节点,串起 Dense、MoE、MLA、V3 协同、R1 与 V4。
|
||||||
其余专题按进度账本持续扩建。
|
其余专题按进度账本持续扩建。
|
||||||
|
|
||||||
## 本地开发
|
## 本地开发
|
||||||
|
|||||||
@@ -100,6 +100,11 @@ SFT / RM / PPO / DPO / RLAIF / RLVR 的角色合同,重点追踪 DeepSeek LLM
|
|||||||
|
|
||||||
CoT、自洽性、搜索、验证器、过程奖励、GRPO、DeepSeekMath、DeepSeek-R1/R1-Zero、Kimi k1.5、multi-effort RL 与 on-policy distillation。
|
CoT、自洽性、搜索、验证器、过程奖励、GRPO、DeepSeekMath、DeepSeek-R1/R1-Zero、Kimi k1.5、multi-effort RL 与 on-policy distillation。
|
||||||
|
|
||||||
|
DeepSeek 聚光二轮已完成:以 24 张问题账和 10 次问题转向,串起 DeepSeek LLM、DeepSeekMoE、
|
||||||
|
DeepSeekMath、V2、V3、R1-Zero/R1、V3.2 与 V4;DAPO / Dr.GRPO 明确作为公开后续反查,
|
||||||
|
Coder/Prover/VL/OCR/系统实现/Engram 作为旁支。配套稀疏容量、MLA 缓存、V3 协同与 RL 偏差四个实验,
|
||||||
|
并以 60 个一手论文或官方仓库节点连接 Kimi K2/K3。
|
||||||
|
|
||||||
### 12. 工具使用与长程 Agent
|
### 12. 工具使用与长程 Agent
|
||||||
|
|
||||||
WebGPT、Toolformer、ReAct、Reflexion、代码 Agent、Computer Use、环境奖励、可验证任务、沙箱、百万 Token 轨迹与 K3 Agentic RL。
|
WebGPT、Toolformer、ReAct、Reflexion、代码 Agent、Computer Use、环境奖励、可验证任务、沙箱、百万 Token 轨迹与 K3 Agentic RL。
|
||||||
|
|||||||
+2
-1
@@ -23,7 +23,8 @@
|
|||||||
"check:multimodal-browser": "node scripts/check-multimodal-browser.mjs",
|
"check:multimodal-browser": "node scripts/check-multimodal-browser.mjs",
|
||||||
"check:inference-serving-browser": "node scripts/check-inference-serving-browser.mjs",
|
"check:inference-serving-browser": "node scripts/check-inference-serving-browser.mjs",
|
||||||
"check:evaluation-browser": "node scripts/check-evaluation-browser.mjs",
|
"check:evaluation-browser": "node scripts/check-evaluation-browser.mjs",
|
||||||
"check:representation-browser": "node scripts/check-representation-browser.mjs"
|
"check:representation-browser": "node scripts/check-representation-browser.mjs",
|
||||||
|
"check:deepseek-browser": "node scripts/check-deepseek-browser.mjs"
|
||||||
},
|
},
|
||||||
"dependencies": {
|
"dependencies": {
|
||||||
"@astrojs/sitemap": "3.7.3",
|
"@astrojs/sitemap": "3.7.3",
|
||||||
|
|||||||
@@ -0,0 +1,189 @@
|
|||||||
|
# DeepSeek 技术谱系二轮:Grok 候选线索
|
||||||
|
|
||||||
|
> 状态:**全部未核验**
|
||||||
|
> 生成方式:本机 Grok CLI Headless,2026-07-29
|
||||||
|
> 用途:查漏、形成问题、发现可能的教学断点。
|
||||||
|
> 禁止用途:不得直接承载正文事实、公式、年份、模型数字、benchmark 结论或因果归因。
|
||||||
|
|
||||||
|
## 1. Grok 对现有页面的审计候选
|
||||||
|
|
||||||
|
现有 `src/pages/deepseek/index.astro` 已具备从 DeepSeek LLM 到 V4 的静态骨架,但可能存在以下缺口:
|
||||||
|
|
||||||
|
1. Dense 基线缺少“为什么可归因需要基线”的实验;
|
||||||
|
2. DeepSeekMoE 缺少组合空间、激活预算和通信税三本账;
|
||||||
|
3. MLA 缺少矩阵吸收、decoupled RoPE 与缓存元素的完整推导;
|
||||||
|
4. V3 只概览 FP8、DualPipe、MTP、loss-free balance,没有让读者操作它们;
|
||||||
|
5. GRPO 只展示组内归一,缺少 clip、KL、全同奖励零信号和 rollout 成本;
|
||||||
|
6. R1 缺少训练轨迹证据边界;
|
||||||
|
7. DAPO / Dr.GRPO 只在推理专题出现,DeepSeek 谱系本身断链;
|
||||||
|
8. V3.2 缺少 indexer warm-up、稀疏训练和 Agent 数据合成;
|
||||||
|
9. V4 的 CSA/HCA、mHC、Muon、混合精度和 effort 只写了摘要;
|
||||||
|
10. DeepSeek → K3 映射缺少“直接祖先 / 同题新解 / 同期不同路线”的反例。
|
||||||
|
|
||||||
|
以上只是候选审计,正式取舍见 `DEEPSEEK_RESEARCH.md`。
|
||||||
|
|
||||||
|
## 2. 二十四张候选问题账
|
||||||
|
|
||||||
|
| 编号 | 候选问题 | 禁止偷换 |
|
||||||
|
|---|---|---|
|
||||||
|
| Q01 | 为什么先建立 dense 坐标系? | 不把后续 MoE 的所有收益归因于稀疏性 |
|
||||||
|
| Q02 | 总参数、激活参数、FLOPs 和显存怎样分账? | 不把 active params 等同同规模 dense 成本 |
|
||||||
|
| Q03 | 细粒度专家怎样改变组合空间? | 不把组合数直接当能力 |
|
||||||
|
| Q04 | shared expert 隔离了什么? | 不写成免费公共知识库 |
|
||||||
|
| Q05 | 稀疏计算为什么带来 all-to-all? | 不用理论 FLOPs 代替墙钟 |
|
||||||
|
| Q06 | aux loss 与语言建模目标为何可能冲突? | aux-loss-free 不等于无均衡 |
|
||||||
|
| Q07 | KV Cache 为什么是服务重复税? | 不把权重显存与 KV 状态混在一起 |
|
||||||
|
| Q08 | MHA、GQA、MLA 分别压缩什么? | MLA 不是 MQA 改名 |
|
||||||
|
| Q09 | joint latent 保留了什么? | 不声称低秩必然无损 |
|
||||||
|
| Q10 | RoPE 为什么妨碍权重吸收? | 不省略位置分支缓存 |
|
||||||
|
| Q11 | FP8 是哪些张量的什么角色? | 不写“全模型八位” |
|
||||||
|
| Q12 | DualPipe 隐藏哪些气泡? | 不写“消灭所有等待” |
|
||||||
|
| Q13 | MTP 的训练与推理角色怎样分开? | 不写“取代自回归” |
|
||||||
|
| Q14 | GRPO 去掉 critic 后成本去了哪里? | 不写“几乎零额外成本” |
|
||||||
|
| Q15 | 可验证奖励能塑造什么? | 不外推到无法验证的开放任务 |
|
||||||
|
| Q16 | R1-Zero 真正隔离了哪个变量? | 不写“无数据从零推理” |
|
||||||
|
| Q17 | R1 的四阶段分别修什么? | 不把 R1 写成纯 RL |
|
||||||
|
| Q18 | 蒸馏为何不是重演探索? | 不把 SFT 学生称为小型 R1-Zero |
|
||||||
|
| Q19 | GRPO 的长度、难度与 clip 偏差是什么? | DAPO / Dr.GRPO 不是官方 R1 配方 |
|
||||||
|
| Q20 | DSA indexer 与固定稀疏模式差在哪? | top-k 不保证召回 |
|
||||||
|
| Q21 | Agentic task synthesis 怎样形成环境闭环? | 不把静态题提升归因成 Agent 能力 |
|
||||||
|
| Q22 | CSA 和 HCA 分别压缩哪本账? | 不把二者合写成一个缩写 |
|
||||||
|
| Q23 | mHC 与 Muon 针对的稳定性对象有何不同? | 不写 Muon 替代全部 AdamW |
|
||||||
|
| Q24 | DeepSeek 与 K3 如何逐对象对照? | 不按总参数或榜单做单轴排名 |
|
||||||
|
|
||||||
|
## 3. 候选历史链
|
||||||
|
|
||||||
|
```text
|
||||||
|
Dense / Scaling
|
||||||
|
→ 容量随激活计算一起涨
|
||||||
|
DeepSeekMoE
|
||||||
|
→ 容量与激活计算分开,但通信与路由变贵
|
||||||
|
V2 / MLA
|
||||||
|
→ 缓存从完整多头 K/V 改为 joint latent,但位置支路与恢复计算仍在
|
||||||
|
V3
|
||||||
|
→ loss-free balance + MTP + FP8 + DualPipe,把模型与集群一起优化
|
||||||
|
DeepSeekMath / GRPO
|
||||||
|
→ 用组内相对奖励省掉 critic,但 rollout、奖励和目标偏差仍在
|
||||||
|
R1-Zero / R1
|
||||||
|
→ 先隔离纯规则奖励实验,再用 cold start、SFT mix 与通用 RL 修复
|
||||||
|
DAPO / Dr.GRPO
|
||||||
|
→ 从复现断点反查 clip、采样、聚合、截断、长度和难度偏差
|
||||||
|
V3.2
|
||||||
|
→ DSA 降长上下文主注意力成本;Agent 合成把推理放入环境
|
||||||
|
V4
|
||||||
|
→ CSA/HCA、mHC、Muon、低精度和异构缓存围绕 1M 联合设计
|
||||||
|
K3
|
||||||
|
→ MLA/MoE 有明确祖先,其余多为同题新解或同期不同路线
|
||||||
|
```
|
||||||
|
|
||||||
|
## 4. 候选一手节点池
|
||||||
|
|
||||||
|
### DeepSeek 主线
|
||||||
|
|
||||||
|
- DeepSeek LLM — `2401.02954`
|
||||||
|
- DeepSeek-Coder — `2401.14196`
|
||||||
|
- DeepSeekMoE — `2401.06066`
|
||||||
|
- DeepSeekMath — `2402.03300`
|
||||||
|
- DeepSeek-V2 — `2405.04434`
|
||||||
|
- DeepSeek-Coder-V2 — `2406.11931`
|
||||||
|
- ESFT — `2407.01906`
|
||||||
|
- DeepSeek-Prover-V1.5 — `2408.08152`
|
||||||
|
- DeepSeek-V3 — `2412.19437`
|
||||||
|
- DeepSeek-R1 — `2501.12948`
|
||||||
|
- DAPO — `2503.14476`
|
||||||
|
- Understanding R1-Zero-Like Training / Dr.GRPO — `2503.20783`
|
||||||
|
- DeepSeek-Prover-V2 — `2504.21801`
|
||||||
|
- DeepEP — `github.com/deepseek-ai/DeepEP`
|
||||||
|
- DualPipe — `github.com/deepseek-ai/DualPipe`
|
||||||
|
- DeepGEMM — `github.com/deepseek-ai/DeepGEMM`
|
||||||
|
- DeepSeek-V3.2 — `2512.02556`
|
||||||
|
- Engram — `2601.07372`
|
||||||
|
- mHC — `2512.24880`
|
||||||
|
- DeepSeek-V4 — `2606.19348`
|
||||||
|
|
||||||
|
### 机制祖先候选
|
||||||
|
|
||||||
|
- Conditional Computation — `cs/0008102`
|
||||||
|
- Sparsely-Gated MoE — `1701.06538`
|
||||||
|
- GShard — `2006.16668`
|
||||||
|
- Switch Transformer — `2101.03961`
|
||||||
|
- ST-MoE — `2202.08906`
|
||||||
|
- MQA — `1911.02150`
|
||||||
|
- GQA — `2305.13245`
|
||||||
|
- RoPE — `2104.09864`
|
||||||
|
- FlashAttention — `2205.14135`
|
||||||
|
- ZeRO — `1910.02054`
|
||||||
|
- GPipe — `1811.06965`
|
||||||
|
- PipeDream — `1806.03377`
|
||||||
|
- PPO — `1707.06347`
|
||||||
|
- InstructGPT — `2203.02155`
|
||||||
|
- Multi-Token Prediction — `2404.19737`
|
||||||
|
- Muon is Scalable — `2502.16982`
|
||||||
|
|
||||||
|
### Kimi 对照候选
|
||||||
|
|
||||||
|
- Kimi k1.5 — `2501.12599`
|
||||||
|
- Kimi K2 — `2507.20534`
|
||||||
|
- Kimi Linear — `2510.26692`
|
||||||
|
- LatentMoE — `2601.18089`
|
||||||
|
- Attention Residuals — `2603.15031`
|
||||||
|
- Kimi K3 — `2607.24653`
|
||||||
|
|
||||||
|
候选池并不等于最终阅读链;正式链必须剔除综述、二手页面和无法核验的年份。
|
||||||
|
|
||||||
|
## 5. 四个候选交互实验
|
||||||
|
|
||||||
|
### A. Sparse Capacity Ledger
|
||||||
|
|
||||||
|
- 输入:专家总数、每 Token 激活数、shared 数、专家宽度、EP 节点数;
|
||||||
|
- 输出:总/激活参数、组合数、教学通信压力;
|
||||||
|
- 误解:总参数等于每 Token 成本、细粒度必然更快、shared 免费。
|
||||||
|
|
||||||
|
### B. MLA Cache Workbench
|
||||||
|
|
||||||
|
- 输入:层数、heads、head dim、GQA groups、latent dim、RoPE dim、上下文、精度;
|
||||||
|
- 输出:MHA/GQA/MLA 每 Token 元素与总缓存、压缩基线;
|
||||||
|
- 误解:V2 的 93.3% 是普适常数、MLA 等于 GQA、RoPE 没有缓存。
|
||||||
|
|
||||||
|
### C. Algorithm–System Co-design Board
|
||||||
|
|
||||||
|
- 输入:micro-batches、pipeline stages、通信/计算比、精度角色、MTP 开关;
|
||||||
|
- 输出:toy bubble、暴露通信、训练目标密度与角色合同;
|
||||||
|
- 误解:DualPipe 清零气泡、FP8 全路径、MTP 默认参与生成。
|
||||||
|
|
||||||
|
### D. GRPO Bias Microscope
|
||||||
|
|
||||||
|
- 输入:一组奖励、序列长度、是否 std norm、response/token aggregation、clip 上下界;
|
||||||
|
- 输出:优势、轨迹权重、零信号、长度倾向;
|
||||||
|
- 误解:去 critic 等于无成本、长 CoT 必然更优、DAPO/Dr.GRPO 是 R1 官方配方。
|
||||||
|
|
||||||
|
## 6. 二十条候选红线
|
||||||
|
|
||||||
|
1. 细粒度专家在所有任务、所有硬件上都优于粗粒度专家;
|
||||||
|
2. shared expert 是不增加计算的公共知识库;
|
||||||
|
3. 激活参数等价于同规模 dense 的完整成本;
|
||||||
|
4. MLA 只是 MQA/GQA 换名;
|
||||||
|
5. V2 的 KV/吞吐数字是 MLA 通用常数;
|
||||||
|
6. decoupled RoPE 分支无需缓存;
|
||||||
|
7. aux-loss-free 等于没有负载均衡;
|
||||||
|
8. FP8 等于全部张量和累加都使用 FP8;
|
||||||
|
9. DualPipe 消灭所有 pipeline bubble;
|
||||||
|
10. V3 的 2.788M GPU hours 可直接与不同硬件模型横比;
|
||||||
|
11. MTP 在推理时取代 next-token generation;
|
||||||
|
12. GRPO 是无偏且严格优于 PPO;
|
||||||
|
13. R1-Zero 不依赖预训练知识;
|
||||||
|
14. “aha” token 是 RL 创造推理的因果证据;
|
||||||
|
15. R1 完整模型没有 SFT;
|
||||||
|
16. 蒸馏学生重演了教师的 RL 探索;
|
||||||
|
17. DAPO/Dr.GRPO 是 DeepSeek 官方 R1 配方;
|
||||||
|
18. DSA top-k 保证不漏关键历史;
|
||||||
|
19. V4 与 K3 使用同一种百万上下文状态;
|
||||||
|
20. “直接祖先”意味着实现和超参数相同。
|
||||||
|
|
||||||
|
## 7. 正式核验结果去向
|
||||||
|
|
||||||
|
- 一手来源结论:`research/DEEPSEEK_RESEARCH.md`
|
||||||
|
- 页面实现:`src/pages/deepseek/index.astro`
|
||||||
|
- 交互实现:`src/components/DeepSeekLab.astro`
|
||||||
|
- 论文索引:`src/data/papers.ts`
|
||||||
|
|
||||||
@@ -0,0 +1,570 @@
|
|||||||
|
# DeepSeek 技术谱系二轮正式研究账本
|
||||||
|
|
||||||
|
> 研究截止:2026-07-29
|
||||||
|
> 课程角色:DeepSeek 聚光专题二轮;与 MoE、长上下文、训练系统、数值、推理、Agent、评测专题互相链接,但不替代各专题完整推导。
|
||||||
|
> 证据规则:正文事实只来自一手论文、作者官方仓库和 Kimi K3 官方报告;Grok 产物仅见 `DEEPSEEK_GROK_LEADS.md`,不承担证据。
|
||||||
|
> 简化规则:所有二维图、成本滑条和训练曲线若非论文复跑,必须标“教学模型”。
|
||||||
|
|
||||||
|
## 0. 本轮要修复什么
|
||||||
|
|
||||||
|
现有 DeepSeek 页面建立了正确的代际骨架,但还不足以让读者回答四类问题:
|
||||||
|
|
||||||
|
1. **可归因性**:一代同时改模型、数据、精度和系统,怎样知道是哪一项在起作用?
|
||||||
|
2. **对象边界**:总参数、激活参数、KV 状态、训练显存、墙钟时间和 benchmark 分数不能混算。
|
||||||
|
3. **训练轨迹**:R1-Zero 的长度曲线、DAPO 的熵崩、Dr.GRPO 的长度偏差分别说明什么?
|
||||||
|
4. **跨模型对照**:V4 与 K3 都支持 1M,并不意味着状态表示、注意力和服务系统相同。
|
||||||
|
|
||||||
|
二轮页面应成为“论文主线的总装图”,而不是再写一遍七个专题。
|
||||||
|
|
||||||
|
## 1. 二十四张正式问题账
|
||||||
|
|
||||||
|
| 编号 | 对象 | 读者问题 | 正式回答边界 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Q01 | Dense 坐标系 | 为什么 DeepSeek LLM 不是可跳过的序章? | 它固定 tokenizer、数据、架构和 scaling 试验的起点;并不单独证明后续所有设计 |
|
||||||
|
| Q02 | 参数角色 | 671B / 37B 各表示什么? | total 是装下的容量,activated 是每 Token 经过的专家参数子集;都不等于端到端 FLOPs |
|
||||||
|
| Q03 | 专家粒度 | 为什么切小专家还要多选? | DeepSeekMoE 把每个专家缩成 `1/m`,总数和激活数同乘 `m`,近似保持专家计算 |
|
||||||
|
| Q04 | Shared expert | 为什么把公共知识单独隔离? | 始终激活的 shared experts 减少 routed experts 重复;仍付激活计算 |
|
||||||
|
| Q05 | 通信税 | 为什么稀疏 FLOPs 不等于便宜? | 路由会产生 dispatch/combine、跨节点 all-to-all、负载长尾和权重访问 |
|
||||||
|
| Q06 | 均衡 | aux-loss-free 到底去掉了什么? | V3 的 expert bias 影响选择、不进入最终 gate weight;仍有 sequence-wise auxiliary loss 防极端失衡 |
|
||||||
|
| Q07 | KV 状态 | 为什么 V2 把服务状态当架构问题? | 权重只装一次,KV 随请求、层、Token 增长,直接限制并发和长上下文 |
|
||||||
|
| Q08 | Attention 压缩 | MQA/GQA/MLA 的差别是什么? | MQA/GQA 共享 K/V 头;MLA 联合低秩压缩 K/V 内容并在计算中恢复 |
|
||||||
|
| Q09 | 矩阵吸收 | MLA 为什么不必显式恢复完整 content key/value? | 无位置项时可利用矩阵乘结合律把上投影吸收到 query/output 投影 |
|
||||||
|
| Q10 | 位置分叉 | 为什么要 decoupled RoPE? | RoPE 位于 key/query 路径中会阻断上述吸收,因此 V2 另设小的 RoPE key/query 分支并缓存 key |
|
||||||
|
| Q11 | FP8 合同 | “FP8 训练”包含哪些角色? | V3 主要 GEMM 用 FP8,配 tile/block scaling、较高精度累加和高精度敏感算子;不是全路径 FP8 |
|
||||||
|
| Q12 | Pipeline | DualPipe 隐藏了什么? | 成对前后向 chunk 的计算—通信重叠并从两端注入 micro-batch;减少而非清零 bubble |
|
||||||
|
| Q13 | MTP | 训练和推理各怎样使用 MTP? | 顺序模块增加未来 Token 监督;推理可丢弃,也可复用于 speculative draft |
|
||||||
|
| Q14 | GRPO | 去掉 critic 后还剩什么? | policy/reference、同题多 rollout、reward/verifier、clip 和 KL;主要省 value model |
|
||||||
|
| Q15 | 可验证奖励 | R1-Zero 的奖励能覆盖哪些任务? | 论文使用数学、代码、逻辑等可规则验证域和格式奖励;复杂开放任务仍是公开限制 |
|
||||||
|
| Q16 | 纯 RL 实验 | R1-Zero 证明了什么? | 强 V3 Base 在无 reasoning SFT 的设置下可被规则奖励继续塑造;不证明无预训练先验 |
|
||||||
|
| Q17 | R1 pipeline | 正式 R1 为什么不是纯 RL? | cold start → reasoning RL → rejection/SFT mix → general RL,各自修可读性、广度和对齐 |
|
||||||
|
| Q18 | 蒸馏 | 为什么学生不是“小号 R1-Zero”? | 报告中的 1.5B–70B 学生主要对约 800K 教师样本做 SFT,没有重演同一 RL |
|
||||||
|
| Q19 | 复现反查 | DAPO/Dr.GRPO 修的是 R1 的什么? | 它们是公开后续研究:分别处理 clip/采样/聚合/截断和长度/难度归一偏差;不是已披露 R1 内部配方 |
|
||||||
|
| Q20 | DSA | 可学习 indexer 为什么不是固定稀疏? | indexer 对历史内容评分,主 attention 只读 top-k;有 warm-up 和 sparse training,仍可能漏检 |
|
||||||
|
| Q21 | Agent 数据 | V3.2 怎样把 reasoning 放进环境? | specialist distillation + mixed RL;真实/合成工具环境、任务、解法和 verifier 构成数据闭环 |
|
||||||
|
| Q22 | V4 Attention | CSA 与 HCA 各压什么? | CSA 先压缩再稀疏 top-k;HCA 用更大压缩率保留所有压缩 entries,不做同类 top-k |
|
||||||
|
| Q23 | V4 稳定化 | mHC、Muon、QK/RMSNorm、clamp 各管什么? | 分别管残差混合、矩阵更新、attention 尺度和 FFN 极值,不能归成一个“稳定性技巧” |
|
||||||
|
| Q24 | K3 对照 | 哪些是祖先,哪些不是? | DeepSeekMoE/MLA 有明确结构继承;QB/KDA/AttnRes/SiTU/MOPD 多为同题新解或同期路线 |
|
||||||
|
|
||||||
|
## 2. 十次历史转向:不要画成产品发布日期
|
||||||
|
|
||||||
|
### W1 / Dense:先制造可比较坐标系
|
||||||
|
|
||||||
|
DeepSeek LLM(`2401.02954`)的历史作用不是“第一代也很强”,而是:
|
||||||
|
|
||||||
|
- 在 7B / 67B dense 模型上固定中英数据、BBPE、训练配方;
|
||||||
|
- 用小模型研究 scaling behavior,再选择大模型超参数;
|
||||||
|
- 把 dedup/filter/remix 和 91-dump 全局去重写成可检查步骤;
|
||||||
|
- 给后续 MoE、MLA 和训练系统提供 dense 对照。
|
||||||
|
|
||||||
|
证据缓存:`research/sources/scaling-laws/2401.02954.txt`。
|
||||||
|
|
||||||
|
### W2 / DeepSeekMoE:容量与激活计算第一次显式分开
|
||||||
|
|
||||||
|
DeepSeekMoE(`2401.06066`)提出:
|
||||||
|
|
||||||
|
1. **Fine-grained expert segmentation**:把 `N` 个专家各切成 `m` 份,总数 `mN`,激活数由 `K` 增至 `mK`,保持专家激活宽度近似不变;
|
||||||
|
2. **Shared expert isolation**:固定激活 `K_s` 个 shared experts,routed 激活数相应减少,使公共变换不必在多个 routed experts 中重复学习。
|
||||||
|
|
||||||
|
论文给的是特定规模和 benchmark 下的消融证据,不是所有 MoE/硬件上的普遍优越性。
|
||||||
|
|
||||||
|
证据缓存:`research/sources/moe/2401.06066.txt:249-330,606-666`。
|
||||||
|
|
||||||
|
### W3 / V2:把推理状态纳入模型结构
|
||||||
|
|
||||||
|
标准 MHA 每层每 Token 的缓存元素与 `2 n_h d_h` 成正比。V2 MLA 的 content 缓存核心变为:
|
||||||
|
|
||||||
|
```text
|
||||||
|
c_t^KV = W_DKV h_t
|
||||||
|
cache_content = d_c
|
||||||
|
cache_total = d_c + d_h^R
|
||||||
|
```
|
||||||
|
|
||||||
|
其中 `d_h^R` 是 decoupled RoPE key 分支。V2 配置使用:
|
||||||
|
|
||||||
|
- `d_c = 512`(论文以 `4 d_h` 表达);
|
||||||
|
- decoupled RoPE per-head dim `d_h^R = 64`;
|
||||||
|
- 每层每 Token 缓存 `d_c + d_h^R` 个元素,不是只有 `d_c`。
|
||||||
|
|
||||||
|
#### 权重吸收为何重要
|
||||||
|
|
||||||
|
对 content path:
|
||||||
|
|
||||||
|
```text
|
||||||
|
qᵀ(W_UK c) = (W_UKᵀ q)ᵀc
|
||||||
|
W_O(W_UV c) = (W_O W_UV)c
|
||||||
|
```
|
||||||
|
|
||||||
|
因此推理实现可以在投影权重中吸收上投影,而非先物化所有 heads 的完整 K/V。若直接把 RoPE 施加到 content key,上式之间会插入位置相关旋转矩阵,无法做同样的固定权重吸收。V2 因而将 RoPE 分支解耦。
|
||||||
|
|
||||||
|
证据缓存:`research/sources/long-context/2405.04434.txt:330-419`。
|
||||||
|
|
||||||
|
V2 的 `−42.5% / −93.3% / 5.76×` 必须始终写成“报告相对 DeepSeek 67B 的特定设置”,不能成为 MLA 常数。
|
||||||
|
|
||||||
|
### W4 / V3:算法—数值—系统协同,不是四个孤立卖点
|
||||||
|
|
||||||
|
V3(`2412.19437`)保留 MLA + DeepSeekMoE,并新增四条互锁机制。
|
||||||
|
|
||||||
|
#### 2.4.1 Auxiliary-loss-free balance
|
||||||
|
|
||||||
|
- 每个 routed expert 有动态 bias `b_i`;
|
||||||
|
- bias 参与 top-k 选择;
|
||||||
|
- 真正乘到专家输出上的 gate value 不包含 bias;
|
||||||
|
- 过载 expert 的 bias 下调,低载 expert 上调;
|
||||||
|
- 报告仍保留 sequence-wise auxiliary loss 防止单序列极端失衡。
|
||||||
|
|
||||||
|
所以“aux-loss-free”只描述主要全局 balance 策略。
|
||||||
|
|
||||||
|
#### 2.4.2 Sequential MTP
|
||||||
|
|
||||||
|
V3 的第 `k` 个 MTP module 接收上一深度 state 与未来 Token embedding,预测额外未来 Token:
|
||||||
|
|
||||||
|
```text
|
||||||
|
L = L_NTP + λ · mean_k(L_MTP^k)
|
||||||
|
```
|
||||||
|
|
||||||
|
MTP embedding/output head 与主模型共享。报告明确:
|
||||||
|
|
||||||
|
- 训练目标主要为 densify signals / pre-plan representations;
|
||||||
|
- 推理时可以直接丢弃 MTP module;
|
||||||
|
- 也可以将其改作 speculative decoding。
|
||||||
|
|
||||||
|
这与“并行一次输出多 Token”不是同一概念。
|
||||||
|
|
||||||
|
#### 2.4.3 FP8 mixed precision
|
||||||
|
|
||||||
|
必须用角色合同描述:
|
||||||
|
|
||||||
|
| 角色 | V3 报告处理 |
|
||||||
|
|---|---|
|
||||||
|
| 高密度 GEMM inputs | 细粒度量化后 FP8 |
|
||||||
|
| GEMM accumulation | Tensor Core 路径外补 FP32 精确累加策略 |
|
||||||
|
| master weights / optimizer | 高精度保存与更新 |
|
||||||
|
| 敏感算子 | BF16/FP32 |
|
||||||
|
| activation cache | 部分低精度保存 |
|
||||||
|
| communication | 结合低精度减少带宽 |
|
||||||
|
|
||||||
|
“V3 用 FP8”不能压缩成单一 dtype 标签。
|
||||||
|
|
||||||
|
#### 2.4.4 DualPipe + DeepEP
|
||||||
|
|
||||||
|
DualPipe:
|
||||||
|
|
||||||
|
- 从 pipeline 两端注入 micro-batches;
|
||||||
|
- 在成对前/后向 chunk 中重叠计算与通信;
|
||||||
|
- 相对经典 schedule 减少 bubble;
|
||||||
|
- 仍有 divisibility、activation memory 和 stage balance 条件。
|
||||||
|
|
||||||
|
DeepEP 官方仓库承载高吞吐/低延迟 expert dispatch/combine kernel;它是系统实现节点,不是 V3 模型算法的新 loss。
|
||||||
|
|
||||||
|
证据:
|
||||||
|
|
||||||
|
- `research/sources/moe/2412.19437.txt:437-448,532-602,644-713,772-999`
|
||||||
|
- `https://github.com/deepseek-ai/DualPipe`
|
||||||
|
- `https://github.com/deepseek-ai/DeepEP`
|
||||||
|
|
||||||
|
### W5 / DeepSeekMath:GRPO 首先是一笔 critic 账
|
||||||
|
|
||||||
|
DeepSeekMath(`2402.03300`)对同一 prompt 采样 `G` 个输出,以组内 reward 构造 outcome advantage:
|
||||||
|
|
||||||
|
```text
|
||||||
|
A_i = (r_i - mean(r_1…r_G)) / (std(r_1…r_G) + ε)
|
||||||
|
```
|
||||||
|
|
||||||
|
完整目标仍含:
|
||||||
|
|
||||||
|
- importance ratio;
|
||||||
|
- clipping;
|
||||||
|
- reference-policy KL;
|
||||||
|
- 每题多 rollout;
|
||||||
|
- reward/verifier 执行。
|
||||||
|
|
||||||
|
所以 GRPO 去掉 value model,不是去掉 RL 系统。
|
||||||
|
|
||||||
|
该论文在特定 7B 数学设置中观察到 Maj@K 改善而 Pass@K 未同样改善,应解释为输出分布重排证据,而非基础覆盖能力的普遍增长。
|
||||||
|
|
||||||
|
证据缓存:`research/sources/reasoning/2402.03300.txt`。
|
||||||
|
|
||||||
|
### W6 / R1-Zero → R1:先做隔离实验,再做可用模型
|
||||||
|
|
||||||
|
#### R1-Zero
|
||||||
|
|
||||||
|
- base:DeepSeek-V3 Base;
|
||||||
|
- 无 reasoning SFT;
|
||||||
|
- GRPO;
|
||||||
|
- accuracy reward + format reward;
|
||||||
|
- reasoning 域使用规则验证,避免大规模 neural RM reward hacking;
|
||||||
|
- 训练中报告 AIME accuracy 与平均 response length 轨迹。
|
||||||
|
|
||||||
|
这证明在该强 base 和验证域上,RL 能进一步塑造搜索/反思行为;不证明知识与算法从零产生。
|
||||||
|
|
||||||
|
#### 正式 R1
|
||||||
|
|
||||||
|
```text
|
||||||
|
V3 Base
|
||||||
|
→ cold-start reasoning data
|
||||||
|
→ reasoning-oriented RL
|
||||||
|
→ rejection sampling + reasoning/general SFT mix
|
||||||
|
→ general RL with rule + preference/safety rewards
|
||||||
|
```
|
||||||
|
|
||||||
|
R1-Zero 的可读性和语言混合问题是正式 R1 增加 cold start 与后续阶段的直接理由。
|
||||||
|
|
||||||
|
#### Distillation
|
||||||
|
|
||||||
|
六个 1.5B–70B 学生使用约 800K R1 生成/筛选样本进行 SFT。它说明强教师轨迹可迁移,不说明学生内部重演了大规模 RL 探索。
|
||||||
|
|
||||||
|
证据缓存:`research/sources/reasoning/2501.12948.txt:85-225,324-437,742-844`。
|
||||||
|
|
||||||
|
### W7 / DAPO 与 Dr.GRPO:复现不是尾注,而是算法显微镜
|
||||||
|
|
||||||
|
二者都不是 DeepSeek 官方 R1 配方;它们是公开后续研究。
|
||||||
|
|
||||||
|
#### DAPO(`2503.14476`)
|
||||||
|
|
||||||
|
论文公开四项技术:
|
||||||
|
|
||||||
|
1. **Clip-Higher**:上下 clip 解耦,提高上界,缓解熵崩;
|
||||||
|
2. **Dynamic Sampling**:过滤 reward 全同、优势为零的组并补采;
|
||||||
|
3. **Token-Level Policy Gradient Loss**:跨 batch Token 聚合,改变长短 response 的权重;
|
||||||
|
4. **Overlong Reward Shaping**:过滤或平滑惩罚被硬截断的长回答,降低 reward noise。
|
||||||
|
|
||||||
|
DAPO 报告的 AIME 分数绑定 Qwen2.5-32B Base、数据、系统和协议,不可写成“算法无条件超过 R1”。
|
||||||
|
|
||||||
|
证据缓存:`research/sources/reasoning/2503.14476.txt:79-91,234-459`。
|
||||||
|
|
||||||
|
#### Dr.GRPO(`2503.20783`)
|
||||||
|
|
||||||
|
论文指出两类偏差:
|
||||||
|
|
||||||
|
- response-level length bias:response loss 除以自身长度,使每个 Token 的总权重依赖长度;
|
||||||
|
- question-level difficulty bias:优势除以组内 reward std,使低 std 问题获得更大尺度。
|
||||||
|
|
||||||
|
其实现用固定全局最大 Token 数作分母,并移除组 std normalization。论文还观察 DeepSeek-V3 Base 在 RL 前即可生成 “aha/wait” 表达,因此单个措辞不能作为 RL 创造反思的因果证据。
|
||||||
|
|
||||||
|
这是一组明确假设和实验,不等于“推翻所有 GRPO”。
|
||||||
|
|
||||||
|
证据缓存:`research/sources/reasoning/2503.20783.txt:319-343,521-611`。
|
||||||
|
|
||||||
|
### W8 / V3.2:Attention 与 Agent 数据同时转向
|
||||||
|
|
||||||
|
#### DSA
|
||||||
|
|
||||||
|
DeepSeek Sparse Attention:
|
||||||
|
|
||||||
|
1. lightning indexer 对 query–history 计算 index score;
|
||||||
|
2. 选择 top-k KV entries;
|
||||||
|
3. 主 attention 只在选中 entries 上计算;
|
||||||
|
4. 在 MLA 架构下实例化;
|
||||||
|
5. 先 dense warm-up 初始化 indexer,再 sparse continued pre-training 对齐。
|
||||||
|
|
||||||
|
主 attention 从 `O(L²)` 降到 `O(Lk)`;indexer 本身仍扫描历史并有额外成本。top-k 是容量约束,不是 recall 保证。
|
||||||
|
|
||||||
|
#### Specialist distillation + mixed RL
|
||||||
|
|
||||||
|
V3.2 将 reasoning、general agent、agentic coding/search 与 human alignment specialists 蒸馏到同一模型,再做 mixed RL。
|
||||||
|
|
||||||
|
Agent 数据必须区分:
|
||||||
|
|
||||||
|
| 类型 | 环境 | Prompt |
|
||||||
|
|---|---|---|
|
||||||
|
| Code agent | 真实 | 抽取 |
|
||||||
|
| Search agent | 真实 API | 合成 |
|
||||||
|
| General agent | 合成工具环境 | 合成 |
|
||||||
|
|
||||||
|
报告给出 1,827 个 general-agent environments,并通过 `<environment, tools, task, verifier>` 闭环生成。不能将其简化为“更多工具调用文本”。
|
||||||
|
|
||||||
|
证据缓存:`research/sources/long-context/2512.02556.txt:122-234,296-371,557-636`。
|
||||||
|
|
||||||
|
### W9 / V4:百万上下文是一组异构状态
|
||||||
|
|
||||||
|
V4(`2606.19348`)不只是把 V3.2 context 拉长。
|
||||||
|
|
||||||
|
#### CSA
|
||||||
|
|
||||||
|
- 先按相邻 hidden states 形成压缩 KV entries;
|
||||||
|
- indexer 在压缩 entries 上做 DSA-style top-k;
|
||||||
|
- 主 attention 只读选中的压缩 entries;
|
||||||
|
- 兼有序列压缩与稀疏选择。
|
||||||
|
|
||||||
|
#### HCA
|
||||||
|
|
||||||
|
- 使用显著更大的 compression rate;
|
||||||
|
- 不做与 CSA 相同的 overlapped compression / top-k sparse selection;
|
||||||
|
- 保留所有更少的压缩 entries;
|
||||||
|
- 追求更激进的固定状态压缩。
|
||||||
|
|
||||||
|
#### 共同细节
|
||||||
|
|
||||||
|
- head-wise query 与 compressed KV RMSNorm;
|
||||||
|
- 最后 64 维 partial RoPE;
|
||||||
|
- shared-KV MQA;
|
||||||
|
- grouped output projection;
|
||||||
|
- 混合层还包含短窗状态,服务端因此维护异构 KV/state cache。
|
||||||
|
|
||||||
|
#### mHC
|
||||||
|
|
||||||
|
将 residual mapping `B_l` 投影到 doubly stochastic matrices 的 Birkhoff polytope:
|
||||||
|
|
||||||
|
```text
|
||||||
|
B_l ≥ 0
|
||||||
|
rowsum(B_l) = 1
|
||||||
|
colsum(B_l) = 1
|
||||||
|
||B_l||₂ ≤ 1
|
||||||
|
```
|
||||||
|
|
||||||
|
V4 expansion factor 为 4。它改变相邻层 residual streams 的混合,不等于 AttnRes 沿历史层检索。
|
||||||
|
|
||||||
|
#### Muon 与稳定化
|
||||||
|
|
||||||
|
- 主要二维矩阵采用 Muon;
|
||||||
|
- embedding、prediction head、RMSNorm weights 等保留 AdamW;
|
||||||
|
- attention Q/K 路径做额外 Norm;
|
||||||
|
- SwiGLU linear branch clamp 到 `[-10,10]`,gate upper cap `10`;
|
||||||
|
- Muon、mHC、Norm、clamp 解决不同对象。
|
||||||
|
|
||||||
|
#### Reasoning effort
|
||||||
|
|
||||||
|
V4 支持多个 effort mode。任何 benchmark 必须带 model variant、effort、context、tool budget 和 harness;“V4 分数”不是单一协议。
|
||||||
|
|
||||||
|
证据缓存:`research/sources/long-context/2606.19348.txt:318-741,783-853,1203-1312,1369-1488,1587-1742`。
|
||||||
|
|
||||||
|
### W10 / K3:继承图必须允许“没有箭头”
|
||||||
|
|
||||||
|
| DeepSeek 节点 | K3 落点 | 关系类型 | 禁止结论 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| DeepSeekMoE fine-grained + shared/routed | Stable LatentMoE shared/routed | 明确结构祖先 | 不代表 expert 数和 router 相同 |
|
||||||
|
| V2 MLA | 周期性 Gated MLA | 明确采用并改造 | K3 不是全 MLA |
|
||||||
|
| V3 loss-free balance | Quantile Balancing | 同问题新方案 | QB 不是 expert bias 改名 |
|
||||||
|
| V3 FP8 | MXFP4 weights + MXFP8 activations QAT | 低精度方向延伸 | 精度角色合同不同 |
|
||||||
|
| GRPO / R1 | multi-domain/multi-effort RL | 共享范式 | K3 未公开等同 R1 的训练轨迹 |
|
||||||
|
| V3.2/V4 long context | 3 KDA + 1 NoPE Gated MLA | 同目标不同状态 | 不把 DSA/CSA/HCA 套到 KDA |
|
||||||
|
| V4 mHC | Block Attention Residuals | 同期不同深度拓扑 | mHC 不是 AttnRes |
|
||||||
|
| V4 Muon | Per-Head Muon | 同优化器族不同参数分组 | 不写相同 optimizer recipe |
|
||||||
|
| R1 distillation | MOPD | 都有 teacher/student | MOPD student rollout 是 on-policy,不能等同离线 SFT |
|
||||||
|
|
||||||
|
K3 直接证据:`research/sources/kimi-k3/k3_tech_report.txt`。
|
||||||
|
|
||||||
|
特别边界:
|
||||||
|
|
||||||
|
- K3 主模型的 93 层以 12 层为 AttnRes block,得到 8 个 layer blocks,加 embedding 共 9 个来源;
|
||||||
|
- K3 报告的推理芯片 nano-model 原型另使用 block size 2;那是芯片概念验证配置,不能回填主模型架构。
|
||||||
|
|
||||||
|
## 3. 四个交互实验合同
|
||||||
|
|
||||||
|
### Lab 01 / Sparse Capacity Ledger
|
||||||
|
|
||||||
|
**输入**
|
||||||
|
|
||||||
|
- architecture:Dense / coarse MoE / DeepSeekMoE / V3;
|
||||||
|
- experts `E`;
|
||||||
|
- routed top-k `k`;
|
||||||
|
- shared experts `s`;
|
||||||
|
- expert width ratio;
|
||||||
|
- EP nodes。
|
||||||
|
|
||||||
|
**输出**
|
||||||
|
|
||||||
|
- total expert units;
|
||||||
|
- active expert units;
|
||||||
|
- theoretical combinations `C(E,k)`(只作组合空间,不作能力);
|
||||||
|
- toy compute ratio;
|
||||||
|
- toy communication pressure;
|
||||||
|
- shared/routed 角色说明。
|
||||||
|
|
||||||
|
**强制边界**
|
||||||
|
|
||||||
|
- 组合数使用对数或科学计数,避免溢出;
|
||||||
|
- 通信为教学指标,不写 GB/s;
|
||||||
|
- total params 与 active params 分列。
|
||||||
|
|
||||||
|
### Lab 02 / MLA Cache Workbench
|
||||||
|
|
||||||
|
**输入**
|
||||||
|
|
||||||
|
- layers、context、batch;
|
||||||
|
- MHA heads、head dim;
|
||||||
|
- GQA KV groups;
|
||||||
|
- MLA latent dim、RoPE dim;
|
||||||
|
- bytes/element。
|
||||||
|
|
||||||
|
**公式**
|
||||||
|
|
||||||
|
```text
|
||||||
|
MHA elements/token/layer = 2 · n_h · d_h
|
||||||
|
GQA elements/token/layer = 2 · n_kv · d_h
|
||||||
|
MLA elements/token/layer = d_c + d_h^R
|
||||||
|
total bytes = per-token-layer · L · T · B · bytes
|
||||||
|
```
|
||||||
|
|
||||||
|
**输出**
|
||||||
|
|
||||||
|
- 元素、GiB、相对当前 MHA 基线的 reduction;
|
||||||
|
- 显式显示“作者报告值 ≠ 当前教学配置”;
|
||||||
|
- weight absorption / RoPE 分叉图。
|
||||||
|
|
||||||
|
### Lab 03 / V3 Co-design Board
|
||||||
|
|
||||||
|
**输入**
|
||||||
|
|
||||||
|
- pipeline stages;
|
||||||
|
- micro-batches;
|
||||||
|
- compute/communication ratio;
|
||||||
|
- schedule:1F1B / dual-ended toy;
|
||||||
|
- precision contract:BF16 / naive FP8 / mixed FP8;
|
||||||
|
- MTP:off / train / speculative。
|
||||||
|
|
||||||
|
**输出**
|
||||||
|
|
||||||
|
- toy bubble fraction;
|
||||||
|
- exposed communication;
|
||||||
|
- activation/master/accumulator dtype 角色;
|
||||||
|
- NTP/MTP supervision count;
|
||||||
|
- MTP inference role。
|
||||||
|
|
||||||
|
**边界**
|
||||||
|
|
||||||
|
- 不声称复现 DualPipe schedule;
|
||||||
|
- bubble/communication 是方向模型;
|
||||||
|
- naive FP8 必须显示风险,不让它看起来更先进。
|
||||||
|
|
||||||
|
### Lab 04 / GRPO Bias Microscope
|
||||||
|
|
||||||
|
**输入**
|
||||||
|
|
||||||
|
- 4–8 条 rollout rewards;
|
||||||
|
- 每条长度;
|
||||||
|
- algorithm:GRPO / DAPO-style / Dr.GRPO;
|
||||||
|
- clip low/high;
|
||||||
|
- std norm;
|
||||||
|
- response-level / token-level aggregation;
|
||||||
|
- overlong threshold。
|
||||||
|
|
||||||
|
**输出**
|
||||||
|
|
||||||
|
- normalized advantage;
|
||||||
|
- 每个 response / token 的 toy gradient weight;
|
||||||
|
- reward 全同零信号;
|
||||||
|
- length/difficulty bias 提示;
|
||||||
|
- DAPO/Dr.GRPO 与 R1 的 provenance 标签。
|
||||||
|
|
||||||
|
**边界**
|
||||||
|
|
||||||
|
- 不模拟完整 optimizer 或真实 policy ratio;
|
||||||
|
- 不把 toy gradient 当训练曲线;
|
||||||
|
- DAPO/Dr.GRPO 明确标“后续公开研究,不是 R1 已披露配方”。
|
||||||
|
|
||||||
|
## 4. 六十节点正式阅读链
|
||||||
|
|
||||||
|
| # | 年份 | 节点 | 一手链接 | 在本页承担的角色 |
|
||||||
|
|---:|---:|---|---|---|
|
||||||
|
| 01 | 1991 | Adaptive Mixtures of Local Experts | https://proceedings.neurips.cc/paper/1991/hash/59b90e1005a220e2ebc542eb9d950b1e-Abstract.html | 专家门控前史 |
|
||||||
|
| 02 | 2000 | Learning to Reason with Neural Networks / Conditional Computation | https://arxiv.org/abs/cs/0008102 | 条件计算 |
|
||||||
|
| 03 | 2003 | A Neural Probabilistic Language Model | https://www.jmlr.org/papers/v3/bengio03a.html | Dense LM 坐标 |
|
||||||
|
| 04 | 2017 | Attention Is All You Need | https://arxiv.org/abs/1706.03762 | Transformer 主干 |
|
||||||
|
| 05 | 2017 | Outrageously Large Neural Networks | https://arxiv.org/abs/1701.06538 | 稀疏 MoE |
|
||||||
|
| 06 | 2017 | Proximal Policy Optimization Algorithms | https://arxiv.org/abs/1707.06347 | GRPO 对照 |
|
||||||
|
| 07 | 2018 | GPipe | https://arxiv.org/abs/1811.06965 | Pipeline 前史 |
|
||||||
|
| 08 | 2018 | PipeDream | https://arxiv.org/abs/1806.03377 | Pipeline schedule |
|
||||||
|
| 09 | 2019 | Fast Transformer Decoding / MQA | https://arxiv.org/abs/1911.02150 | KV 共享 |
|
||||||
|
| 10 | 2019 | Megatron-LM | https://arxiv.org/abs/1909.08053 | 模型并行 |
|
||||||
|
| 11 | 2019 | ZeRO | https://arxiv.org/abs/1910.02054 | 状态分片 |
|
||||||
|
| 12 | 2019 | RMSNorm | https://arxiv.org/abs/1910.07467 | 尺度控制 |
|
||||||
|
| 13 | 2020 | GShard | https://arxiv.org/abs/2006.16668 | 大规模 MoE |
|
||||||
|
| 14 | 2020 | QK-Normalization | https://arxiv.org/abs/2010.04245 | attention logit 稳定 |
|
||||||
|
| 15 | 2021 | Switch Transformers | https://arxiv.org/abs/2101.03961 | coarse top-1 MoE |
|
||||||
|
| 16 | 2021 | RoFormer / RoPE | https://arxiv.org/abs/2104.09864 | MLA 位置分叉 |
|
||||||
|
| 17 | 2022 | ST-MoE | https://arxiv.org/abs/2202.08906 | MoE 稳定性 |
|
||||||
|
| 18 | 2022 | DeepNet | https://arxiv.org/abs/2203.00555 | 深层残差 |
|
||||||
|
| 19 | 2022 | InstructGPT | https://arxiv.org/abs/2203.02155 | SFT/RM/PPO 合同 |
|
||||||
|
| 20 | 2022 | FlashAttention | https://arxiv.org/abs/2205.14135 | IO-aware exact attention |
|
||||||
|
| 21 | 2022 | Process and Outcome Feedback | https://arxiv.org/abs/2211.14275 | reasoning reward 前史 |
|
||||||
|
| 22 | 2022 | Self-Consistency | https://arxiv.org/abs/2203.11171 | 多采样聚合 |
|
||||||
|
| 23 | 2023 | GQA | https://arxiv.org/abs/2305.13245 | KV 分组 |
|
||||||
|
| 24 | 2023 | Let's Verify Step by Step | https://arxiv.org/abs/2305.20050 | verifier / PRM |
|
||||||
|
| 25 | 2023 | Direct Preference Optimization | https://arxiv.org/abs/2305.18290 | RL 外偏好路线 |
|
||||||
|
| 26 | 2023 | FlashAttention-2 | https://arxiv.org/abs/2307.08691 | attention kernel |
|
||||||
|
| 27 | 2023 | PagedAttention / vLLM | https://arxiv.org/abs/2309.06180 | KV 服务状态 |
|
||||||
|
| 28 | 2024 | DeepSeek LLM | https://arxiv.org/abs/2401.02954 | Dense/scaling 基线 |
|
||||||
|
| 29 | 2024 | DeepSeek-Coder | https://arxiv.org/abs/2401.14196 | 代码数据旁支 |
|
||||||
|
| 30 | 2024 | DeepSeekMoE | https://arxiv.org/abs/2401.06066 | 细粒度 + shared |
|
||||||
|
| 31 | 2024 | DeepSeekMath | https://arxiv.org/abs/2402.03300 | 数学数据 + GRPO |
|
||||||
|
| 32 | 2024 | RLOO | https://arxiv.org/abs/2402.14740 | critic-free 对照 |
|
||||||
|
| 33 | 2024 | DeepSeek-V2 | https://arxiv.org/abs/2405.04434 | MLA + MoE |
|
||||||
|
| 34 | 2024 | Better & Faster LLMs via MTP | https://arxiv.org/abs/2404.19737 | MTP 祖先 |
|
||||||
|
| 35 | 2024 | DeepSeek-Coder-V2 | https://arxiv.org/abs/2406.11931 | V2 continued pretrain 旁支 |
|
||||||
|
| 36 | 2024 | ESFT | https://arxiv.org/abs/2407.01906 | 专家特化微调 |
|
||||||
|
| 37 | 2024 | DeepSeek-Prover-V1.5 | https://arxiv.org/abs/2408.08152 | proof feedback RL |
|
||||||
|
| 38 | 2024 | Hyper-Connections | https://arxiv.org/abs/2409.19606 | mHC 前身 |
|
||||||
|
| 39 | 2024 | DeepSeek-V3 | https://arxiv.org/abs/2412.19437 | FP8/DualPipe/MTP |
|
||||||
|
| 40 | 2025 | DeepSeek-R1 | https://arxiv.org/abs/2501.12948 | R1-Zero/R1/蒸馏 |
|
||||||
|
| 41 | 2025 | Muon is Scalable for LLM Training | https://arxiv.org/abs/2502.16982 | V4 optimizer 前史 |
|
||||||
|
| 42 | 2025 | DAPO | https://arxiv.org/abs/2503.14476 | GRPO 工程修正 |
|
||||||
|
| 43 | 2025 | Understanding R1-Zero-Like Training | https://arxiv.org/abs/2503.20783 | Dr.GRPO / 偏差 |
|
||||||
|
| 44 | 2025 | DeepSeek-Prover-V2 | https://arxiv.org/abs/2504.21801 | subgoal + RL |
|
||||||
|
| 45 | 2025 | DeepEP | https://github.com/deepseek-ai/DeepEP | Expert Parallel kernel |
|
||||||
|
| 46 | 2025 | DualPipe | https://github.com/deepseek-ai/DualPipe | V3/R1 pipeline 实现 |
|
||||||
|
| 47 | 2025 | DeepGEMM | https://github.com/deepseek-ai/DeepGEMM | FP8 GEMM 实现 |
|
||||||
|
| 48 | 2025 | DeepSeek-VL2 | https://arxiv.org/abs/2412.10302 | 多模态理解旁支 |
|
||||||
|
| 49 | 2025 | Janus-Pro | https://arxiv.org/abs/2501.17811 | 统一理解/生成旁支 |
|
||||||
|
| 50 | 2025 | Kimi k1.5 | https://arxiv.org/abs/2501.12599 | 同期 reasoning RL |
|
||||||
|
| 51 | 2025 | Kimi K2 | https://arxiv.org/abs/2507.20534 | MLA/MoE/Muon 对照 |
|
||||||
|
| 52 | 2025 | Kimi Linear | https://arxiv.org/abs/2510.26692 | KDA 前身 |
|
||||||
|
| 53 | 2025 | DeepSeek-V3.2 | https://arxiv.org/abs/2512.02556 | DSA + Agent |
|
||||||
|
| 54 | 2025 | mHC | https://arxiv.org/abs/2512.24880 | 受约束 residual |
|
||||||
|
| 55 | 2026 | Engram | https://arxiv.org/abs/2601.07372 | 条件记忆新稀疏轴 |
|
||||||
|
| 56 | 2026 | LatentMoE | https://arxiv.org/abs/2601.18089 | K3 routed latent 前身 |
|
||||||
|
| 57 | 2026 | Attention Residuals | https://arxiv.org/abs/2603.15031 | K3 深度路由 |
|
||||||
|
| 58 | 2026 | DeepSeek-V4 | https://arxiv.org/abs/2606.19348 | CSA/HCA/mHC/Muon |
|
||||||
|
| 59 | 2026 | Kimi K3 | https://arxiv.org/abs/2607.24653 | 对照锚点 |
|
||||||
|
| 60 | 2026 | Kimi K3 official code/model repository | https://github.com/MoonshotAI/Kimi-K3 | 开放实现边界 |
|
||||||
|
|
||||||
|
## 5. 允许进入正文的报告数字
|
||||||
|
|
||||||
|
所有数字必须带比较对象:
|
||||||
|
|
||||||
|
| 数字 | 允许写法 | 禁止写法 |
|
||||||
|
|---|---|---|
|
||||||
|
| V2 `42.5% / 93.3% / 5.76×` | V2 报告相对 DeepSeek 67B 特定设置 | MLA 固有加速 |
|
||||||
|
| V3 `671B / 37B` | total / activated params | 等价 37B dense 端到端成本 |
|
||||||
|
| V3 `14.8T` | 报告预训练 token 总量 | 数据质量证明 |
|
||||||
|
| V3 `2.788M H800 hours` | 报告完整训练口径;硬件限定 | 跨模型统一成本 |
|
||||||
|
| R1 `~800K` | 教师生成/筛选的 distill SFT samples | 小模型自主 RL 数据 |
|
||||||
|
| V3.2 `1,827 environments` | general-agent 合成环境数 | 全部 Agent 数据规模 |
|
||||||
|
| V4 `1.6T/49B`、`284B/13B` | Pro/Flash total/active | 两模型性能排序 |
|
||||||
|
| V4 `27%/10%` 等 | 报告相对 V3.2、1M context 的 FLOPs/KV | 所有服务栈固定比例 |
|
||||||
|
| K3 `2.8T/104B` | total/active | 与 V4 单轴优劣 |
|
||||||
|
|
||||||
|
## 6. 事实审计红线
|
||||||
|
|
||||||
|
- [x] DAPO 与 Dr.GRPO 不写成 DeepSeek 官方 R1 recipe。
|
||||||
|
- [x] “aha/wait” 不写成 RL 从零创造推理的因果证据。
|
||||||
|
- [x] R1 与 R1-Zero 分开。
|
||||||
|
- [x] Distill students 不写成重跑 RL。
|
||||||
|
- [x] aux-loss-free 不写成没有任何 auxiliary balance。
|
||||||
|
- [x] MLA content cache 与 decoupled RoPE cache 都进入公式。
|
||||||
|
- [x] FP8 用完整角色合同。
|
||||||
|
- [x] MTP 训练、可丢弃推理与 speculative role 分开。
|
||||||
|
- [x] DSA indexer 成本与漏检风险保留。
|
||||||
|
- [x] CSA 与 HCA 分开。
|
||||||
|
- [x] mHC 与 AttnRes 分开。
|
||||||
|
- [x] Muon 与 AdamW 参数分组保留。
|
||||||
|
- [x] V4 与 K3 按状态对象对照,不按 1M 标签归并。
|
||||||
|
- [x] 主模型 AttnRes block size 12 与 MiniTriton benchmark block size 2 分开。
|
||||||
|
- [x] 所有 benchmark/成本数字带报告、配置和比较对象。
|
||||||
|
|
||||||
|
## 7. 页面验收合同
|
||||||
|
|
||||||
|
- 至少 24 张问题账;
|
||||||
|
- 至少 20 个正文目录;
|
||||||
|
- 60 个一手/官方阅读节点;
|
||||||
|
- 四个独立可操作实验;
|
||||||
|
- DeepSeekMath 必须在主时间线中;
|
||||||
|
- DAPO / Dr.GRPO 必须标后续公开研究;
|
||||||
|
- MLA 实验必须把 RoPE cache 算进去;
|
||||||
|
- V3 实验必须显示 FP8 角色而不是单一开关;
|
||||||
|
- R1 pipeline 必须同时可见 Zero 与正式 R1;
|
||||||
|
- V4/K3 表必须包含“直接祖先 / 同题新解 / 同期不同路线”;
|
||||||
|
- 桌面与 390px 移动端无文档级横向溢出;
|
||||||
|
- tabs 支持键盘方向键;
|
||||||
|
- toy model、作者报告和公式推导使用不同标签;
|
||||||
|
- 专属 Chrome 回归并纳入全站回归。
|
||||||
@@ -228,8 +228,8 @@ if (numeric(reliability.initial.passAt) <= numeric(reliability.k2.passAt) || num
|
|||||||
if (reliability.nonIdempotent.sideRisk === "LOW") failures.push("非幂等写操作风险没有提升");
|
if (reliability.nonIdempotent.sideRisk === "LOW") failures.push("非幂等写操作风险没有提升");
|
||||||
if (numeric(rl.wait.utilization) >= numeric(rl.full.utilization) || numeric(rl.wait.lostWork) <= numeric(rl.full.lostWork)) failures.push("wait-all 长尾/重算方向异常");
|
if (numeric(rl.wait.utilization) >= numeric(rl.full.utilization) || numeric(rl.wait.lostWork) <= numeric(rl.full.lostWork)) failures.push("wait-all 长尾/重算方向异常");
|
||||||
if (!rl.wait.takeaway.includes("wait-all") || rl.keyboardSelected !== "rl" || rl.keyboardVisible !== "rl") failures.push("长程 RL 解释或键盘导航异常");
|
if (!rl.wait.takeaway.includes("wait-all") || rl.keyboardSelected !== "rl" || rl.keyboardVisible !== "rl") failures.push("长程 RL 解释或键盘导航异常");
|
||||||
if (home.releaseCards !== 15 || !home.firstRelease.includes("hidden state") || home.firstHref !== "/architecture/representation/") failures.push("首页评测首发入口异常");
|
if (home.releaseCards !== 16 || !home.firstRelease.includes("从 Dense 到百万上下文") || home.firstHref !== "/deepseek/") failures.push("首页评测首发入口异常");
|
||||||
if (home.paperCount !== "480" || papers.total !== 480 || !papers.hasAgentFilter || papers.agentVisible < 52) failures.push("论文库 Agent 标签或论文总数异常");
|
if (home.paperCount !== "486" || papers.total !== 486 || !papers.hasAgentFilter || papers.agentVisible < 52) failures.push("论文库 Agent 标签或论文总数异常");
|
||||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
||||||
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||||
|
|
||||||
|
|||||||
@@ -226,8 +226,8 @@ if (!update.steps[0].includes("Fixed preference")) failures.push("DPO 更新流
|
|||||||
if (!recipe.family.includes("Multi-effort") || !recipe.regime.includes("9 RL experts") || !recipe.constraints.includes("verbosity")) failures.push("K3 配方合同异常");
|
if (!recipe.family.includes("Multi-effort") || !recipe.regime.includes("9 RL experts") || !recipe.constraints.includes("verbosity")) failures.push("K3 配方合同异常");
|
||||||
if (!recipe.path.some((step) => step.includes("3 domains × 3 efforts")) || !recipe.path.some((step) => step.includes("MOPD"))) failures.push("K3 配方路径异常");
|
if (!recipe.path.some((step) => step.includes("3 domains × 3 efforts")) || !recipe.path.some((step) => step.includes("MOPD"))) failures.push("K3 配方路径异常");
|
||||||
if (recipe.keyboardSelected !== "recipe" || recipe.keyboardVisible !== "recipe") failures.push("实验 tab 键盘导航异常");
|
if (recipe.keyboardSelected !== "recipe" || recipe.keyboardVisible !== "recipe") failures.push("实验 tab 键盘导航异常");
|
||||||
if (home.releaseCards !== 15 || !home.firstRelease.includes("hidden state") || home.firstHref !== "/architecture/representation/") failures.push("首页评测首发入口异常");
|
if (home.releaseCards !== 16 || !home.firstRelease.includes("从 Dense 到百万上下文") || home.firstHref !== "/deepseek/") failures.push("首页评测首发入口异常");
|
||||||
if (home.paperCount !== "480" || papers.total !== 480 || !papers.hasAlignmentFilter || papers.alignmentVisible < 35) failures.push("论文库后训练标签或论文总数异常");
|
if (home.paperCount !== "486" || papers.total !== 486 || !papers.hasAlignmentFilter || papers.alignmentVisible < 35) failures.push("论文库后训练标签或论文总数异常");
|
||||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
||||||
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||||
|
|
||||||
|
|||||||
@@ -234,11 +234,11 @@ if (layout.navLinks !== 20 || mobile.mobileLinks !== 20 || home.navLinks !== 20)
|
|||||||
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
|
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
|
||||||
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
||||||
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
|
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
|
||||||
if (home.releaseCards !== 15 || !home.firstRelease.includes("hidden state") || home.firstHref !== "/architecture/representation/") {
|
if (home.releaseCards !== 16 || !home.firstRelease.includes("从 Dense 到百万上下文") || home.firstHref !== "/deepseek/") {
|
||||||
failures.push("首页 Transformer 新章入口异常");
|
failures.push("首页 Transformer 新章入口异常");
|
||||||
}
|
}
|
||||||
if (home.paperCount !== "480") failures.push(`首页论文总数异常:${home.paperCount}`);
|
if (home.paperCount !== "486") failures.push(`首页论文总数异常:${home.paperCount}`);
|
||||||
if (!papers.hasDataFilter || papers.total !== 480 || papers.visible < 25) failures.push("论文库数据标签或论文总数异常");
|
if (!papers.hasDataFilter || papers.total !== 486 || papers.visible < 25) failures.push("论文库数据标签或论文总数异常");
|
||||||
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
|
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
|
||||||
|
|
||||||
socket.close();
|
socket.close();
|
||||||
|
|||||||
@@ -0,0 +1,303 @@
|
|||||||
|
import { writeFileSync } from "node:fs";
|
||||||
|
|
||||||
|
const cdpPort = process.env.CDP_PORT ?? "9227";
|
||||||
|
const baseUrl = process.env.SITE_URL ?? "http://127.0.0.1:4327";
|
||||||
|
const pages = await fetch(`http://127.0.0.1:${cdpPort}/json/list`).then((response) => response.json());
|
||||||
|
const page = pages.find((entry) => entry.type === "page");
|
||||||
|
if (!page) throw new Error(`CDP ${cdpPort} 没有可用页面`);
|
||||||
|
|
||||||
|
const socket = new WebSocket(page.webSocketDebuggerUrl);
|
||||||
|
await new Promise((resolve, reject) => {
|
||||||
|
socket.addEventListener("open", resolve, { once: true });
|
||||||
|
socket.addEventListener("error", reject, { once: true });
|
||||||
|
});
|
||||||
|
|
||||||
|
let nextId = 0;
|
||||||
|
const pending = new Map();
|
||||||
|
const exceptions = [];
|
||||||
|
socket.addEventListener("message", (event) => {
|
||||||
|
const message = JSON.parse(event.data);
|
||||||
|
if (message.id && pending.has(message.id)) {
|
||||||
|
const { resolve, reject } = pending.get(message.id);
|
||||||
|
pending.delete(message.id);
|
||||||
|
if (message.error) reject(new Error(message.error.message));
|
||||||
|
else resolve(message.result);
|
||||||
|
}
|
||||||
|
if (message.method === "Runtime.exceptionThrown") {
|
||||||
|
exceptions.push(message.params.exceptionDetails.exception?.description ?? message.params.exceptionDetails.text);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
const command = (method, params = {}) => new Promise((resolve, reject) => {
|
||||||
|
const id = ++nextId;
|
||||||
|
pending.set(id, { resolve, reject });
|
||||||
|
socket.send(JSON.stringify({ id, method, params }));
|
||||||
|
});
|
||||||
|
const pause = (milliseconds) => new Promise((resolve) => setTimeout(resolve, milliseconds));
|
||||||
|
const evaluate = async (expression) => {
|
||||||
|
const result = await command("Runtime.evaluate", { expression, returnByValue: true, awaitPromise: true });
|
||||||
|
if (result.exceptionDetails) throw new Error(result.exceptionDetails.exception?.description ?? result.exceptionDetails.text);
|
||||||
|
return result.result.value;
|
||||||
|
};
|
||||||
|
const navigate = async (path) => {
|
||||||
|
await command("Page.navigate", { url: `${baseUrl}${path}` });
|
||||||
|
for (let attempt = 0; attempt < 70; attempt += 1) {
|
||||||
|
await pause(100);
|
||||||
|
if (await evaluate("document.readyState === 'complete'")) return;
|
||||||
|
}
|
||||||
|
throw new Error(`${path} 加载超时`);
|
||||||
|
};
|
||||||
|
const screenshot = async (path) => {
|
||||||
|
const result = await command("Page.captureScreenshot", { format: "png", captureBeyondViewport: false });
|
||||||
|
writeFileSync(path, Buffer.from(result.data, "base64"));
|
||||||
|
};
|
||||||
|
|
||||||
|
await command("Page.enable");
|
||||||
|
await command("Runtime.enable");
|
||||||
|
await command("Emulation.setDeviceMetricsOverride", {
|
||||||
|
width: 1440,
|
||||||
|
height: 1100,
|
||||||
|
deviceScaleFactor: 1,
|
||||||
|
mobile: false,
|
||||||
|
});
|
||||||
|
|
||||||
|
await navigate("/deepseek/");
|
||||||
|
await screenshot("/tmp/llm-atlas-deepseek-desktop.png");
|
||||||
|
|
||||||
|
const overview = await evaluate(`(() => ({
|
||||||
|
title: document.querySelector("h1")?.textContent.trim(),
|
||||||
|
sections: document.querySelectorAll(".article-section").length,
|
||||||
|
tocLinks: document.querySelectorAll(".side-rail a").length,
|
||||||
|
ledgers: document.querySelectorAll(".ledger-card").length,
|
||||||
|
waves: document.querySelectorAll(".wave-grid > article").length,
|
||||||
|
paperLinks: document.querySelectorAll("[data-deepseek-paper-chain] a").length,
|
||||||
|
labTabs: document.querySelectorAll("[data-ds-tab]").length,
|
||||||
|
labPanels: document.querySelectorAll("[data-ds-panel]").length,
|
||||||
|
branches: document.querySelectorAll(".branch-grid > a").length,
|
||||||
|
followups: document.querySelectorAll(".lineage-row.followup").length,
|
||||||
|
navLinks: document.querySelectorAll(".top-nav a").length,
|
||||||
|
activeNav: document.querySelector('.top-nav a[aria-current="page"]')?.textContent.trim(),
|
||||||
|
documentOverflow: document.documentElement.scrollWidth - document.documentElement.clientWidth,
|
||||||
|
}))()`);
|
||||||
|
|
||||||
|
const capacity = await evaluate(`(() => {
|
||||||
|
const root = document.querySelector("[data-deepseek-lab]");
|
||||||
|
const read = () => ({
|
||||||
|
panel: root.querySelector("[data-ds-panel]:not([hidden])").dataset.dsPanel,
|
||||||
|
total: root.querySelector("[data-total-capacity]").textContent.trim(),
|
||||||
|
active: root.querySelector("[data-active-compute]").textContent.trim(),
|
||||||
|
combinations: root.querySelector("[data-combinations]").textContent.trim(),
|
||||||
|
communication: root.querySelector("[data-communication]").textContent.trim(),
|
||||||
|
name: root.querySelector("[data-capacity-name]").textContent.trim(),
|
||||||
|
explain: root.querySelector("[data-capacity-explain]").textContent.trim(),
|
||||||
|
});
|
||||||
|
const initial = read();
|
||||||
|
root.querySelector('[data-capacity-preset="dense"]').click();
|
||||||
|
const dense = read();
|
||||||
|
root.querySelector('[data-capacity-preset="deepseekmoe"]').click();
|
||||||
|
const fine = read();
|
||||||
|
root.querySelector('[data-capacity-preset="v3"]').click();
|
||||||
|
const v3 = read();
|
||||||
|
return { initial, dense, fine, v3 };
|
||||||
|
})()`);
|
||||||
|
|
||||||
|
const cache = await evaluate(`(() => {
|
||||||
|
const root = document.querySelector("[data-deepseek-lab]");
|
||||||
|
root.querySelector('[data-ds-tab="cache"]').click();
|
||||||
|
const read = () => ({
|
||||||
|
panel: root.querySelector("[data-ds-panel]:not([hidden])").dataset.dsPanel,
|
||||||
|
mha: root.querySelector("[data-mha-elements]").textContent.trim(),
|
||||||
|
gqa: root.querySelector("[data-gqa-elements]").textContent.trim(),
|
||||||
|
mla: root.querySelector("[data-mla-elements]").textContent.trim(),
|
||||||
|
rope: root.querySelector("[data-rope-cache]").textContent.trim(),
|
||||||
|
selected: root.querySelector("[data-selected-cache]").textContent.trim(),
|
||||||
|
baseline: root.querySelector("[data-mha-cache]").textContent.trim(),
|
||||||
|
reduction: root.querySelector("[data-cache-reduction]").textContent.trim(),
|
||||||
|
boundary: root.querySelector("[data-cache-boundary]").textContent.trim(),
|
||||||
|
});
|
||||||
|
const initial = read();
|
||||||
|
const rope = root.querySelector("[data-rope-dim]");
|
||||||
|
rope.value = "0";
|
||||||
|
rope.dispatchEvent(new Event("input", { bubbles: true }));
|
||||||
|
const noRope = read();
|
||||||
|
const context = root.querySelector("[data-context-length]");
|
||||||
|
context.value = "1048576";
|
||||||
|
context.dispatchEvent(new Event("input", { bubbles: true }));
|
||||||
|
const million = read();
|
||||||
|
return { initial, noRope, million };
|
||||||
|
})()`);
|
||||||
|
|
||||||
|
const codesign = await evaluate(`(() => {
|
||||||
|
const root = document.querySelector("[data-deepseek-lab]");
|
||||||
|
root.querySelector('[data-ds-tab="codesign"]').click();
|
||||||
|
const readSchedule = () => ({
|
||||||
|
bubble: root.querySelector("[data-bubble]").textContent.trim(),
|
||||||
|
exposed: root.querySelector("[data-exposed-comm]").textContent.trim(),
|
||||||
|
});
|
||||||
|
const oneWay = readSchedule();
|
||||||
|
root.querySelector('[data-schedule="dual"]').click();
|
||||||
|
const dual = readSchedule();
|
||||||
|
root.querySelector('[data-precision="naive"]').click();
|
||||||
|
const naive = {
|
||||||
|
risk: root.querySelector("[data-risk-label]").textContent.trim(),
|
||||||
|
accum: root.querySelector("[data-accum-dtype]").textContent.trim(),
|
||||||
|
explain: root.querySelector("[data-precision-explain]").textContent.trim(),
|
||||||
|
};
|
||||||
|
root.querySelector('[data-precision="mixed"]').click();
|
||||||
|
const mixed = {
|
||||||
|
risk: root.querySelector("[data-risk-label]").textContent.trim(),
|
||||||
|
accum: root.querySelector("[data-accum-dtype]").textContent.trim(),
|
||||||
|
sensitive: root.querySelector("[data-sensitive-dtype]").textContent.trim(),
|
||||||
|
};
|
||||||
|
root.querySelector('[data-mtp-role="off"]').click();
|
||||||
|
const off = root.querySelector("[data-mtp-supervision]").textContent.trim();
|
||||||
|
root.querySelector('[data-mtp-role="draft"]').click();
|
||||||
|
const draft = {
|
||||||
|
supervision: root.querySelector("[data-mtp-supervision]").textContent.trim(),
|
||||||
|
cost: root.querySelector("[data-mtp-main-cost]").textContent.trim(),
|
||||||
|
explain: root.querySelector("[data-mtp-explain]").textContent.trim(),
|
||||||
|
};
|
||||||
|
return { panel: root.querySelector("[data-ds-panel]:not([hidden])").dataset.dsPanel, oneWay, dual, naive, mixed, off, draft };
|
||||||
|
})()`);
|
||||||
|
|
||||||
|
const rl = await evaluate(`(() => {
|
||||||
|
const root = document.querySelector("[data-deepseek-lab]");
|
||||||
|
root.querySelector('[data-ds-tab="rl"]').click();
|
||||||
|
const read = () => ({
|
||||||
|
signal: root.querySelector("[data-signal-state]").textContent.trim(),
|
||||||
|
mean: root.querySelector("[data-reward-mean]").textContent.trim(),
|
||||||
|
std: root.querySelector("[data-reward-std]").textContent.trim(),
|
||||||
|
effective: root.querySelector("[data-effective]").textContent.trim(),
|
||||||
|
provenance: root.querySelector("[data-provenance]").textContent.trim(),
|
||||||
|
algorithm: root.querySelector("[data-algorithm-name]").textContent.trim(),
|
||||||
|
boundary: root.querySelector("[data-rl-boundary]").textContent.trim(),
|
||||||
|
weights: [...root.querySelectorAll("[data-advantage-rows] > div span:last-child")].map((node) => node.textContent.trim()),
|
||||||
|
});
|
||||||
|
const initial = read();
|
||||||
|
root.querySelector('[data-reward-preset="same"]').click();
|
||||||
|
const same = read();
|
||||||
|
root.querySelector('[data-reward-preset="longwrong"]').click();
|
||||||
|
root.querySelector('[data-rl-algorithm="dapo"]').click();
|
||||||
|
const dapo = read();
|
||||||
|
root.querySelector('[data-rl-algorithm="dr"]').click();
|
||||||
|
const dr = read();
|
||||||
|
root.querySelector('[data-r1-mode="r1"]').click();
|
||||||
|
const r1 = root.querySelector("[data-r1-mode-explain]").textContent.trim();
|
||||||
|
root.querySelector('[data-r1-mode="distill"]').click();
|
||||||
|
const distill = root.querySelector("[data-r1-mode-explain]").textContent.trim();
|
||||||
|
const first = root.querySelector('[data-ds-tab="capacity"]');
|
||||||
|
first.focus();
|
||||||
|
first.dispatchEvent(new KeyboardEvent("keydown", { key: "ArrowRight", bubbles: true }));
|
||||||
|
return {
|
||||||
|
initial, same, dapo, dr, r1, distill,
|
||||||
|
keyboardSelected: root.querySelector('[data-ds-tab][aria-selected="true"]').dataset.dsTab,
|
||||||
|
keyboardVisible: root.querySelector("[data-ds-panel]:not([hidden])").dataset.dsPanel,
|
||||||
|
};
|
||||||
|
})()`);
|
||||||
|
|
||||||
|
await evaluate(`(() => {
|
||||||
|
document.querySelector("[data-deepseek-lab]").scrollIntoView({ block: "start", behavior: "instant" });
|
||||||
|
window.scrollBy(0, -82);
|
||||||
|
})()`);
|
||||||
|
await pause(180);
|
||||||
|
await screenshot("/tmp/llm-atlas-deepseek-lab-desktop.png");
|
||||||
|
|
||||||
|
await navigate("/");
|
||||||
|
const home = await evaluate(`(() => ({
|
||||||
|
releaseCards: document.querySelectorAll(".release-card").length,
|
||||||
|
firstRelease: document.querySelector(".release-card h2").textContent.trim(),
|
||||||
|
firstHref: document.querySelector(".release-card").getAttribute("href"),
|
||||||
|
paperCount: document.querySelector(".hero-stats div:nth-child(3) b").textContent.trim(),
|
||||||
|
navLinks: document.querySelectorAll(".top-nav a").length,
|
||||||
|
}))()`);
|
||||||
|
|
||||||
|
await navigate("/papers/");
|
||||||
|
const papers = await evaluate(`(() => {
|
||||||
|
const button = [...document.querySelectorAll("[data-filter]")].find((node) => node.textContent.trim() === "DeepSeek");
|
||||||
|
button?.click();
|
||||||
|
return {
|
||||||
|
total: document.querySelectorAll("[data-paper]").length,
|
||||||
|
visible: document.querySelectorAll("[data-paper]:not([hidden])").length,
|
||||||
|
hasFilter: Boolean(button),
|
||||||
|
hasCoder: document.body.textContent.includes("DeepSeek-Coder-V2"),
|
||||||
|
hasEngram: document.body.textContent.includes("Conditional Memory via Scalable Lookup"),
|
||||||
|
};
|
||||||
|
})()`);
|
||||||
|
|
||||||
|
await command("Emulation.setDeviceMetricsOverride", {
|
||||||
|
width: 390,
|
||||||
|
height: 844,
|
||||||
|
deviceScaleFactor: 1,
|
||||||
|
mobile: true,
|
||||||
|
});
|
||||||
|
await navigate("/deepseek/");
|
||||||
|
const mobile = await evaluate(`(() => {
|
||||||
|
const root = document.querySelector("[data-deepseek-lab]");
|
||||||
|
root.scrollIntoView({ block: "start", behavior: "instant" });
|
||||||
|
const toggle = document.querySelector("#menu-toggle");
|
||||||
|
toggle?.click();
|
||||||
|
return {
|
||||||
|
documentOverflow: document.documentElement.scrollWidth - document.documentElement.clientWidth,
|
||||||
|
menuVisible: getComputedStyle(toggle).display !== "none",
|
||||||
|
menuOpen: toggle.getAttribute("aria-expanded"),
|
||||||
|
mobileLinks: document.querySelectorAll("#mobile-nav a").length,
|
||||||
|
tabs: root.querySelectorAll("[data-ds-tab]").length,
|
||||||
|
offenders: [...document.querySelectorAll("body *")]
|
||||||
|
.filter((node) => !node.closest(".paper-chain, .advantage-table, .precision-table, .mapping-table, [data-deepseek-lab]"))
|
||||||
|
.filter((node) => node.getBoundingClientRect().right > document.documentElement.clientWidth + 1)
|
||||||
|
.slice(0, 12)
|
||||||
|
.map((node) => ({
|
||||||
|
tag: node.tagName,
|
||||||
|
className: typeof node.className === "string" ? node.className : "",
|
||||||
|
right: Math.round(node.getBoundingClientRect().right),
|
||||||
|
width: Math.round(node.getBoundingClientRect().width),
|
||||||
|
})),
|
||||||
|
};
|
||||||
|
})()`);
|
||||||
|
await evaluate(`(() => {
|
||||||
|
document.querySelector("#menu-toggle")?.click();
|
||||||
|
window.scrollBy(0, -82);
|
||||||
|
})()`);
|
||||||
|
await pause(180);
|
||||||
|
await screenshot("/tmp/llm-atlas-deepseek-mobile.png");
|
||||||
|
|
||||||
|
const report = { overview, capacity, cache, codesign, rl, home, papers, mobile, exceptions };
|
||||||
|
console.log(JSON.stringify(report, null, 2));
|
||||||
|
|
||||||
|
const numeric = (text) => Number.parseFloat(text.replaceAll(",", ""));
|
||||||
|
const failures = [];
|
||||||
|
if (!overview.title.includes("为什么转向")) failures.push("专题标题异常");
|
||||||
|
if (overview.sections !== 25 || overview.tocLinks !== 25) failures.push("二十四个编号专题加阅读链的目录结构异常");
|
||||||
|
if (overview.ledgers !== 24 || overview.waves !== 10) failures.push("二十四张问题账或十次转向结构异常");
|
||||||
|
if (overview.paperLinks !== 60 || overview.branches !== 5 || overview.followups !== 1) failures.push("论文链、旁支或公开后续标记异常");
|
||||||
|
if (overview.labTabs !== 4 || overview.labPanels !== 4) failures.push("四联实验结构异常");
|
||||||
|
if (overview.navLinks !== 20 || home.navLinks !== 20 || mobile.mobileLinks !== 20 || overview.activeNav !== "DeepSeek") failures.push("全站导航未同步 DeepSeek");
|
||||||
|
if (overview.documentOverflow > 1 || mobile.documentOverflow > 1) failures.push("桌面或移动端存在文档级横向溢出");
|
||||||
|
if (capacity.initial.panel !== "capacity" || capacity.initial.total !== "32.1× FFN" || capacity.initial.active !== "1.13× FFN") failures.push("V3 稀疏容量初始账异常");
|
||||||
|
if (!capacity.dense.name.includes("DENSE") || capacity.dense.communication !== "NONE" || numeric(capacity.dense.total) !== numeric(capacity.dense.active)) failures.push("Dense 容量预设异常");
|
||||||
|
if (!capacity.fine.name.includes("FINE-GRAINED") || !capacity.fine.explain.includes("shared")) failures.push("DeepSeekMoE 预设异常");
|
||||||
|
if (!capacity.v3.combinations.includes("10^") || capacity.v3.communication !== "HIGH") failures.push("V3 路由组合或通信方向异常");
|
||||||
|
if (cache.initial.panel !== "cache" || numeric(cache.initial.mha) !== 32768 || numeric(cache.initial.gqa) !== 2048 || numeric(cache.initial.mla) !== 576 || numeric(cache.initial.rope) !== 64) failures.push("MLA 精确元素账异常");
|
||||||
|
if (numeric(cache.initial.reduction) !== 98.2 || numeric(cache.noRope.mla) !== 512 || numeric(cache.noRope.reduction) <= numeric(cache.initial.reduction)) failures.push("RoPE cache 或 MLA reduction 异常");
|
||||||
|
if (!cache.million.selected.includes("GiB") || !cache.million.boundary.includes("1,048,576")) failures.push("百万 Token 缓存账异常");
|
||||||
|
if (numeric(codesign.dual.bubble) >= numeric(codesign.oneWay.bubble) || numeric(codesign.dual.exposed) >= numeric(codesign.oneWay.exposed)) failures.push("Dual-ended toy 没有减少空泡或暴露通信");
|
||||||
|
if (codesign.naive.risk !== "CRITICAL" || codesign.naive.accum !== "FP8" || codesign.mixed.risk !== "MANAGED" || !codesign.mixed.accum.includes("FP32")) failures.push("FP8 角色合同异常");
|
||||||
|
if (!codesign.off.includes("1 token") || !codesign.draft.supervision.includes("draft") || !codesign.draft.explain.includes("验收率")) failures.push("MTP 生命周期异常");
|
||||||
|
if (rl.initial.signal !== "GROUP-RELATIVE SIGNAL" || rl.same.signal !== "ZERO GROUP SIGNAL" || !rl.same.boundary.includes("优势为零")) failures.push("GRPO 零方差信号异常");
|
||||||
|
if (!rl.dapo.provenance.includes("2503.14476") || !rl.dapo.algorithm.includes("FOLLOW-UP") || !rl.dr.provenance.includes("2503.20783")) failures.push("DAPO / Dr.GRPO 来源边界异常");
|
||||||
|
if (!rl.r1.includes("cold start") || !rl.distill.includes("没有重演")) failures.push("R1 / distill 身份切换异常");
|
||||||
|
if (rl.keyboardSelected !== "cache" || rl.keyboardVisible !== "cache") failures.push("实验键盘 tab 导航异常");
|
||||||
|
if (home.releaseCards !== 16 || !home.firstRelease.includes("从 Dense 到百万上下文") || home.firstHref !== "/deepseek/" || home.paperCount !== "486") failures.push("首页 DeepSeek 首发入口或论文数异常");
|
||||||
|
if (papers.total !== 486 || !papers.hasFilter || papers.visible < 20 || !papers.hasCoder || !papers.hasEngram) failures.push("论文库 DeepSeek 聚光异常");
|
||||||
|
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
||||||
|
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
||||||
|
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||||
|
|
||||||
|
if (failures.length) {
|
||||||
|
console.error(`\nFAIL\n- ${failures.join("\n- ")}`);
|
||||||
|
process.exitCode = 1;
|
||||||
|
} else {
|
||||||
|
console.log("\nPASS DeepSeek browser regression");
|
||||||
|
}
|
||||||
|
|
||||||
|
socket.close();
|
||||||
@@ -277,8 +277,8 @@ if (numeric(system.initial.success) <= numeric(system.initial.model) || numeric(
|
|||||||
if (numeric(system.cheap.success) >= numeric(system.initial.success) || numeric(system.cheap.cost) !== 4) failures.push("低预算没有降低成功率 / 成本");
|
if (numeric(system.cheap.success) >= numeric(system.initial.success) || numeric(system.cheap.cost) !== 4) failures.push("低预算没有降低成功率 / 成本");
|
||||||
if (numeric(system.locked.unsafe) !== 0 || numeric(system.locked.overrefusal) <= numeric(system.initial.overrefusal)) failures.push("安全壳没有展现危险服从 / 过拒权衡");
|
if (numeric(system.locked.unsafe) !== 0 || numeric(system.locked.overrefusal) <= numeric(system.initial.overrefusal)) failures.push("安全壳没有展现危险服从 / 过拒权衡");
|
||||||
if (system.keyboardSelected !== "judge" || system.keyboardVisible !== "judge") failures.push("实验键盘 tab 导航异常");
|
if (system.keyboardSelected !== "judge" || system.keyboardVisible !== "judge") failures.push("实验键盘 tab 导航异常");
|
||||||
if (home.releaseCards !== 15 || !home.firstRelease.includes("hidden state") || home.firstHref !== "/architecture/representation/") failures.push("首页评测首发入口异常");
|
if (home.releaseCards !== 16 || !home.firstRelease.includes("从 Dense 到百万上下文") || home.firstHref !== "/deepseek/") failures.push("首页评测首发入口异常");
|
||||||
if (home.paperCount !== "480" || home.topicCount !== "17" || papers.total !== 480 || !papers.hasFilter || papers.visible < 80) failures.push("首页 / 论文库评测索引异常");
|
if (home.paperCount !== "486" || home.topicCount !== "17" || papers.total !== 486 || !papers.hasFilter || papers.visible < 80) failures.push("首页 / 论文库评测索引异常");
|
||||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
||||||
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
||||||
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||||
|
|||||||
@@ -264,8 +264,8 @@ if (!fleet.k3.avoided.includes("320K") || fleet.k3.shortSlo !== "PROTECTED") fai
|
|||||||
if (!fleet.failed.state.includes("SECONDARY RE-PREFILL") || !fleet.failed.recompute.includes("FAILED PRIMARY")) failures.push("缓存故障没有触发原子失效后的重算");
|
if (!fleet.failed.state.includes("SECONDARY RE-PREFILL") || !fleet.failed.recompute.includes("FAILED PRIMARY")) failures.push("缓存故障没有触发原子失效后的重算");
|
||||||
if (fleet.bursty.shortSlo !== "VIOLATED") failures.push("平均并发阈值没有暴露长请求突发");
|
if (fleet.bursty.shortSlo !== "VIOLATED") failures.push("平均并发阈值没有暴露长请求突发");
|
||||||
if (fleet.keyboardSelected !== "phase" || fleet.keyboardVisible !== "phase") failures.push("实验键盘 tab 导航异常");
|
if (fleet.keyboardSelected !== "phase" || fleet.keyboardVisible !== "phase") failures.push("实验键盘 tab 导航异常");
|
||||||
if (home.releaseCards !== 15 || !home.firstRelease.includes("hidden state") || home.firstHref !== "/architecture/representation/") failures.push("首页评测首发入口异常");
|
if (home.releaseCards !== 16 || !home.firstRelease.includes("从 Dense 到百万上下文") || home.firstHref !== "/deepseek/") failures.push("首页评测首发入口异常");
|
||||||
if (home.paperCount !== "480" || papers.total !== 480 || !papers.hasFilter || papers.visible !== 45) failures.push("论文库推理服务标签或总数异常");
|
if (home.paperCount !== "486" || papers.total !== 486 || !papers.hasFilter || papers.visible !== 46) failures.push("论文库推理服务标签或总数异常");
|
||||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
||||||
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
||||||
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||||
|
|||||||
@@ -177,7 +177,7 @@ if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentO
|
|||||||
}
|
}
|
||||||
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
||||||
if (!mobile.menuVisible) failures.push("移动端菜单按钮未显示");
|
if (!mobile.menuVisible) failures.push("移动端菜单按钮未显示");
|
||||||
if (home.releaseCards !== 15) failures.push(`首页新章卡数量异常:${home.releaseCards}`);
|
if (home.releaseCards !== 16) failures.push(`首页新章卡数量异常:${home.releaseCards}`);
|
||||||
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
|
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
|
||||||
|
|
||||||
socket.close();
|
socket.close();
|
||||||
|
|||||||
@@ -246,8 +246,8 @@ if (ocr.unreported.status !== "OUT OF EVIDENCE" || ocr.unreported.accuracy !== "
|
|||||||
if (loop.toolsStart.state !== "OPEN" || loop.toolsEnd.state !== "VERIFIED" || loop.toolsEnd.evidence !== "97%" || loop.toolsEnd.tools !== "3") failures.push("vision-in-the-loop 终局异常");
|
if (loop.toolsStart.state !== "OPEN" || loop.toolsEnd.state !== "VERIFIED" || loop.toolsEnd.evidence !== "97%" || loop.toolsEnd.tools !== "3") failures.push("vision-in-the-loop 终局异常");
|
||||||
if (loop.cotEnd.state !== "FAILED" || !loop.cotEnd.takeaway.includes("不能凭空增加")) failures.push("文字 CoT 与新观察没有分开");
|
if (loop.cotEnd.state !== "FAILED" || !loop.cotEnd.takeaway.includes("不能凭空增加")) failures.push("文字 CoT 与新观察没有分开");
|
||||||
if (loop.keyboardSelected !== "connector" || loop.keyboardVisible !== "connector") failures.push("实验键盘 tab 导航异常");
|
if (loop.keyboardSelected !== "connector" || loop.keyboardVisible !== "connector") failures.push("实验键盘 tab 导航异常");
|
||||||
if (home.releaseCards !== 15 || !home.firstRelease.includes("hidden state") || home.firstHref !== "/architecture/representation/") failures.push("首页评测首发入口异常");
|
if (home.releaseCards !== 16 || !home.firstRelease.includes("从 Dense 到百万上下文") || home.firstHref !== "/deepseek/") failures.push("首页评测首发入口异常");
|
||||||
if (home.paperCount !== "480" || papers.total !== 480 || !papers.hasFilter || papers.multimodalVisible < 59) failures.push("论文库多模态标签或总数异常");
|
if (home.paperCount !== "486" || papers.total !== 486 || !papers.hasFilter || papers.multimodalVisible < 59) failures.push("论文库多模态标签或总数异常");
|
||||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
||||||
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
||||||
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||||
|
|||||||
@@ -277,11 +277,11 @@ if (layout.navLinks !== 20 || mobile.mobileLinks !== 20 || home.navLinks !== 20)
|
|||||||
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
|
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
|
||||||
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
||||||
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
|
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
|
||||||
if (home.releaseCards !== 15 || !home.firstRelease.includes("hidden state") || home.firstHref !== "/architecture/representation/") {
|
if (home.releaseCards !== 16 || !home.firstRelease.includes("从 Dense 到百万上下文") || home.firstHref !== "/deepseek/") {
|
||||||
failures.push("首页 Transformer 新章入口异常");
|
failures.push("首页 Transformer 新章入口异常");
|
||||||
}
|
}
|
||||||
if (home.paperCount !== "480") failures.push(`首页论文总数异常:${home.paperCount}`);
|
if (home.paperCount !== "486") failures.push(`首页论文总数异常:${home.paperCount}`);
|
||||||
if (!papers.hasOptimizerFilter || papers.total !== 480 || papers.visible < 8) failures.push("论文库优化器标签或论文总数异常");
|
if (!papers.hasOptimizerFilter || papers.total !== 486 || papers.visible < 8) failures.push("论文库优化器标签或论文总数异常");
|
||||||
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
|
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
|
||||||
|
|
||||||
socket.close();
|
socket.close();
|
||||||
|
|||||||
@@ -289,7 +289,7 @@ if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentO
|
|||||||
}
|
}
|
||||||
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
||||||
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
|
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
|
||||||
if (home.releaseCards !== 15 || !home.firstRelease.includes("hidden state")) failures.push("首页评测新章入口异常");
|
if (home.releaseCards !== 16 || !home.firstRelease.includes("从 Dense 到百万上下文")) failures.push("首页评测新章入口异常");
|
||||||
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
|
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
|
||||||
|
|
||||||
socket.close();
|
socket.close();
|
||||||
|
|||||||
@@ -288,8 +288,8 @@ if (numeric(residual.attnres.states) !== 9 || !residual.attnres.routeExplain.inc
|
|||||||
if (!residual.clamp.activation.includes("V4") || !residual.clamp.bound.includes("100")) failures.push("DeepSeek-V4 clamp 展示异常");
|
if (!residual.clamp.activation.includes("V4") || !residual.clamp.bound.includes("100")) failures.push("DeepSeek-V4 clamp 展示异常");
|
||||||
if (!residual.situ.activation.includes("KIMI") || !residual.situ.bound.includes("100")) failures.push("K3 SiTU 上界展示异常");
|
if (!residual.situ.activation.includes("KIMI") || !residual.situ.bound.includes("100")) failures.push("K3 SiTU 上界展示异常");
|
||||||
if (residual.keyboardSelected !== "position" || residual.keyboardVisible !== "position") failures.push("实验键盘 tab 导航异常");
|
if (residual.keyboardSelected !== "position" || residual.keyboardVisible !== "position") failures.push("实验键盘 tab 导航异常");
|
||||||
if (home.releaseCards !== 15 || !home.firstRelease.includes("hidden state") || home.firstHref !== "/architecture/representation/") failures.push("首页表示新章入口异常");
|
if (home.releaseCards !== 16 || !home.firstRelease.includes("从 Dense 到百万上下文") || home.firstHref !== "/deepseek/") failures.push("首页表示新章入口异常");
|
||||||
if (home.paperCount !== "480" || home.topicCount !== "17" || papers.total !== 480 || !papers.hasFilter || papers.visible < 30) failures.push("首页 / 论文库表示索引异常");
|
if (home.paperCount !== "486" || home.topicCount !== "17" || papers.total !== 486 || !papers.hasFilter || papers.visible < 30) failures.push("首页 / 论文库表示索引异常");
|
||||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
||||||
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
||||||
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||||
|
|||||||
@@ -273,10 +273,10 @@ if (layout.navLinks !== 20 || mobile.mobileLinks !== 20 || home.navLinks !== 20)
|
|||||||
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
|
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
|
||||||
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
||||||
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
|
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
|
||||||
if (home.releaseCards !== 15 || !home.firstRelease.includes("hidden state") || home.firstHref !== "/architecture/representation/") {
|
if (home.releaseCards !== 16 || !home.firstRelease.includes("从 Dense 到百万上下文") || home.firstHref !== "/deepseek/") {
|
||||||
failures.push("首页 Transformer 新章入口异常");
|
failures.push("首页 Transformer 新章入口异常");
|
||||||
}
|
}
|
||||||
if (home.paperCount !== "480") failures.push(`首页论文总数异常:${home.paperCount}`);
|
if (home.paperCount !== "486") failures.push(`首页论文总数异常:${home.paperCount}`);
|
||||||
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
|
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
|
||||||
|
|
||||||
socket.close();
|
socket.close();
|
||||||
|
|||||||
@@ -233,7 +233,7 @@ if (layout.articleSections !== 16 || layout.paperLinks !== 37 || layout.labTabs
|
|||||||
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
|
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
|
||||||
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
||||||
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
|
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
|
||||||
if (home.releaseCards !== 15 || !home.firstRelease.includes("hidden state")) failures.push("首页评测新章入口异常");
|
if (home.releaseCards !== 16 || !home.firstRelease.includes("从 Dense 到百万上下文")) failures.push("首页评测新章入口异常");
|
||||||
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
|
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
|
||||||
|
|
||||||
socket.close();
|
socket.close();
|
||||||
|
|||||||
@@ -172,7 +172,7 @@ const home = await evaluate(`(() => ({
|
|||||||
releaseCards: document.querySelectorAll(".release-card").length,
|
releaseCards: document.querySelectorAll(".release-card").length,
|
||||||
firstRelease: document.querySelector(".release-card h2").textContent,
|
firstRelease: document.querySelector(".release-card h2").textContent,
|
||||||
firstHref: document.querySelector(".release-card").getAttribute("href"),
|
firstHref: document.querySelector(".release-card").getAttribute("href"),
|
||||||
paperCount: [...document.querySelectorAll(".hero-stats b")].map((node) => node.textContent.trim()).find((value) => value === "480"),
|
paperCount: [...document.querySelectorAll(".hero-stats b")].map((node) => node.textContent.trim()).find((value) => value === "486"),
|
||||||
}))()`);
|
}))()`);
|
||||||
|
|
||||||
await navigate("/papers/");
|
await navigate("/papers/");
|
||||||
@@ -236,8 +236,8 @@ if (block.family.trim() !== "Hybrid MoE" || !block.kv.includes("3 KDA : 1 Gated
|
|||||||
if (!block.path.some((step) => step.includes("KDA × 3")) || !block.note.includes("AttnRes")) failures.push("K3 Block 路径异常");
|
if (!block.path.some((step) => step.includes("KDA × 3")) || !block.note.includes("AttnRes")) failures.push("K3 Block 路径异常");
|
||||||
if (block.context.trim() !== "128K" || numeric(block.mha) !== 400 || numeric(block.kda) !== 1) failures.push("KV 成本缩放异常");
|
if (block.context.trim() !== "128K" || numeric(block.mha) !== 400 || numeric(block.kda) !== 1) failures.push("KV 成本缩放异常");
|
||||||
if (block.keyboardSelected !== "block" || block.keyboardVisible !== "block") failures.push("实验 tab 键盘导航异常");
|
if (block.keyboardSelected !== "block" || block.keyboardVisible !== "block") failures.push("实验 tab 键盘导航异常");
|
||||||
if (home.releaseCards !== 15 || !home.firstRelease.includes("hidden state") || home.firstHref !== "/architecture/representation/") failures.push("首页评测首发入口异常");
|
if (home.releaseCards !== 16 || !home.firstRelease.includes("从 Dense 到百万上下文") || home.firstHref !== "/deepseek/") failures.push("首页评测首发入口异常");
|
||||||
if (home.paperCount !== "480" || papers.total !== 480 || papers.transformerVisible < 30) failures.push("论文库或首页论文数量异常");
|
if (home.paperCount !== "486" || papers.total !== 486 || papers.transformerVisible < 30) failures.push("论文库或首页论文数量异常");
|
||||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
||||||
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,623 @@
|
|||||||
|
<figure class="ds-lab" data-deepseek-lab>
|
||||||
|
<div class="ds-lab-head">
|
||||||
|
<div>
|
||||||
|
<p>INTERACTIVE / DEEPSEEK SYSTEM ATLAS</p>
|
||||||
|
<h3>四本账,把“模型创新”拆回可计算对象</h3>
|
||||||
|
</div>
|
||||||
|
<p>
|
||||||
|
这里混合精确计数与显式 toy model:参数组合和 KV 元素按公式计算;通信、bubble 与梯度权重只展示方向。
|
||||||
|
每个面板都标出证据边界,不能拿来替代真实 checkpoint 或集群复跑。
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="ds-tabs" role="tablist" aria-label="选择 DeepSeek 技术实验">
|
||||||
|
<button type="button" role="tab" data-ds-tab="capacity" aria-selected="true">
|
||||||
|
<span>01</span><b>稀疏容量账</b><small>MoE · shared · communication</small>
|
||||||
|
</button>
|
||||||
|
<button type="button" role="tab" data-ds-tab="cache" aria-selected="false" tabindex="-1">
|
||||||
|
<span>02</span><b>MLA 缓存账</b><small>MHA · GQA · latent · RoPE</small>
|
||||||
|
</button>
|
||||||
|
<button type="button" role="tab" data-ds-tab="codesign" aria-selected="false" tabindex="-1">
|
||||||
|
<span>03</span><b>V3 协同账</b><small>FP8 · DualPipe · MTP</small>
|
||||||
|
</button>
|
||||||
|
<button type="button" role="tab" data-ds-tab="rl" aria-selected="false" tabindex="-1">
|
||||||
|
<span>04</span><b>RL 偏差镜</b><small>GRPO · DAPO · Dr.GRPO</small>
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<section class="ds-panel" data-ds-panel="capacity">
|
||||||
|
<div class="panel-intro">
|
||||||
|
<div><span>WORKBENCH 01 / SPARSE CAPACITY</span><h4>总参数很大,不代表每个 Token 都经过全部专家</h4></div>
|
||||||
|
<p>切换架构或自定义专家配置,分开观察总容量、激活计算和跨设备通信;组合数只是可选路径,不是能力分数。</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="preset-row" role="group" aria-label="选择容量预设">
|
||||||
|
<button type="button" data-capacity-preset="dense"><b>DENSE</b><small>一条完整 FFN</small></button>
|
||||||
|
<button type="button" data-capacity-preset="coarse"><b>COARSE MOE</b><small>8 选 2</small></button>
|
||||||
|
<button type="button" data-capacity-preset="deepseekmoe"><b>DEEPSEEKMOE</b><small>细粒度 + shared</small></button>
|
||||||
|
<button type="button" data-capacity-preset="v3" class="active"><b>V3</b><small>256 routed · 8 active</small></button>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="control-grid five">
|
||||||
|
<label><span>routed experts E <output data-experts-label>256</output></span><input data-experts type="range" min="1" max="384" value="256" /></label>
|
||||||
|
<label><span>routed top-k <output data-topk-label>8</output></span><input data-topk type="range" min="0" max="16" value="8" /></label>
|
||||||
|
<label><span>shared experts <output data-shared-label>1</output></span><input data-shared type="range" min="0" max="4" value="1" /></label>
|
||||||
|
<label><span>expert width <output data-expert-width-label>0.125×</output></span><input data-expert-width type="range" min="1" max="100" value="12.5" step=".5" /></label>
|
||||||
|
<label><span>EP nodes <output data-ep-label>4</output></span><input data-ep-nodes type="range" min="1" max="16" value="4" /></label>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="capacity-stage">
|
||||||
|
<div class="router-box">
|
||||||
|
<span>TOKEN</span><b>h<sub>t</sub></b><small>router scores</small>
|
||||||
|
</div>
|
||||||
|
<i>→</i>
|
||||||
|
<div class="expert-map" data-expert-map aria-label="教学专家池"></div>
|
||||||
|
<i>→</i>
|
||||||
|
<div class="router-box output"><span>COMBINE</span><b>Σ gᵢEᵢ(h)</b><small>shared + routed</small></div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="metric-grid four">
|
||||||
|
<article><span>TOTAL EXPERT CAPACITY</span><b data-total-capacity>33.0× FFN</b><p>按 expert width 折算,不含 attention</p></article>
|
||||||
|
<article><span>ACTIVE EXPERT COMPUTE</span><b data-active-compute>2.0× FFN</b><p>每 Token 的教学 FFN 单位</p></article>
|
||||||
|
<article><span>ROUTE COMBINATIONS</span><b data-combinations>≈ 10¹⁵</b><p>只表示可组合路径,不代表专长质量</p></article>
|
||||||
|
<article class="dark"><span>COMMUNICATION PRESSURE</span><b data-communication>HIGH</b><p>教学指标;不是 GB/s 实测</p></article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="evidence-boundary">
|
||||||
|
<b data-capacity-name>DEEPSEEK-V3 / CAPACITY CONTRACT</b>
|
||||||
|
<p data-capacity-explain>V3 每层含 1 个 shared 与 256 个 routed experts,每 Token 激活 8 个 routed;理论稀疏计算仍需要 expert dispatch/combine 与负载均衡。</p>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="ds-panel" data-ds-panel="cache" hidden>
|
||||||
|
<div class="panel-intro">
|
||||||
|
<div><span>WORKBENCH 02 / INFERENCE STATE</span><h4>MLA 省的不是模型权重,而是每个请求不断增长的历史状态</h4></div>
|
||||||
|
<p>所有方案都与当前 MHA 基线比较;MLA 的缓存必须包含 joint latent 和 decoupled RoPE key,不能只报 latent。</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="control-grid four">
|
||||||
|
<label><span>layers L <output data-cache-layers-label>60</output></span><input data-cache-layers type="range" min="12" max="96" step="12" value="60" /></label>
|
||||||
|
<label><span>context T</span><select data-context-length><option value="4096">4K</option><option value="32768">32K</option><option value="131072" selected>128K</option><option value="1048576">1M</option></select></label>
|
||||||
|
<label><span>batch B <output data-batch-label>4</output></span><input data-batch type="range" min="1" max="32" value="4" /></label>
|
||||||
|
<label><span>storage</span><select data-cache-bytes><option value="1">FP8 / 1 byte</option><option value="2" selected>BF16 / 2 bytes</option><option value="4">FP32 / 4 bytes</option></select></label>
|
||||||
|
</div>
|
||||||
|
<div class="control-grid five">
|
||||||
|
<label><span>MHA heads <output data-heads-label>128</output></span><input data-heads type="range" min="8" max="160" step="8" value="128" /></label>
|
||||||
|
<label><span>head dim dₕ <output data-head-dim-label>128</output></span><input data-head-dim type="range" min="32" max="256" step="32" value="128" /></label>
|
||||||
|
<label><span>GQA KV groups <output data-kv-groups-label>8</output></span><input data-kv-groups type="range" min="1" max="32" value="8" /></label>
|
||||||
|
<label><span>MLA latent d꜀ <output data-latent-label>512</output></span><input data-latent type="range" min="128" max="2048" step="128" value="512" /></label>
|
||||||
|
<label><span>RoPE key dᴿ <output data-rope-dim-label>64</output></span><input data-rope-dim type="range" min="0" max="128" step="16" value="64" /></label>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="cache-comparison">
|
||||||
|
<article data-cache-card="mha">
|
||||||
|
<span>MHA</span><b data-mha-elements>32,768</b><small>elements / token / layer</small>
|
||||||
|
<div><i data-mha-bar></i></div><p>完整 K 与 V,所有 heads 分开缓存。</p>
|
||||||
|
</article>
|
||||||
|
<article data-cache-card="gqa">
|
||||||
|
<span>GQA</span><b data-gqa-elements>2,048</b><small>elements / token / layer</small>
|
||||||
|
<div><i data-gqa-bar></i></div><p>Query heads 分组共享 K/V。</p>
|
||||||
|
</article>
|
||||||
|
<article class="selected" data-cache-card="mla">
|
||||||
|
<span>MLA / V2-LIKE</span><b data-mla-elements>576</b><small>d꜀ + dᴿ elements</small>
|
||||||
|
<div><i data-mla-bar></i></div><p>joint latent + decoupled RoPE key。</p>
|
||||||
|
</article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="absorption-grid">
|
||||||
|
<div class="absorb-flow">
|
||||||
|
<span>CONTENT PATH / ASSOCIATIVITY</span>
|
||||||
|
<code>qᵀ(W<sub>UK</sub>c) = (W<sub>UK</sub>ᵀq)ᵀc</code>
|
||||||
|
<div><b>q</b><i>× absorbed W</i><b>c<sup>KV</sup></b><i>→</i><b>score</b></div>
|
||||||
|
<p>固定线性上投影可吸收到 query/output 权重,不必先物化完整多头 content K/V。</p>
|
||||||
|
</div>
|
||||||
|
<div class="rope-flow">
|
||||||
|
<span>POSITION PATH / DECOUPLED</span>
|
||||||
|
<code>q<sup>R</sup>ᵀ R(j−i) k<sup>R</sup></code>
|
||||||
|
<div><b>qᴿ</b><i>rotate</i><b>kᴿ</b><i>cache</i><b data-rope-cache>64</b></div>
|
||||||
|
<p>RoPE 是位置相关运算,无法作为固定矩阵吸收;因此 key 的小位置分支仍要缓存。</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="metric-grid four">
|
||||||
|
<article><span>SELECTED MLA CACHE</span><b data-selected-cache>32.96 GiB</b><p>当前 L × T × B × bytes</p></article>
|
||||||
|
<article><span>MHA BASELINE</span><b data-mha-cache>1.88 TiB</b><p>同一配置,不是报告基线</p></article>
|
||||||
|
<article><span>REDUCTION VS MHA</span><b data-cache-reduction>98.2%</b><p>由当前滑条计算</p></article>
|
||||||
|
<article class="dark"><span>REPORT ≠ TOY</span><b>93.3%</b><p>V2 报告相对 DeepSeek 67B</p></article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="evidence-boundary">
|
||||||
|
<b>EXACT ELEMENT COUNT / CONFIGURATION-DEPENDENT BYTES</b>
|
||||||
|
<p data-cache-boundary>元素公式是精确账;GiB 只按所选 dtype 计算,没有加入 allocator、page、quantization metadata 或 kernel workspace。</p>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="ds-panel" data-ds-panel="codesign" hidden>
|
||||||
|
<div class="panel-intro">
|
||||||
|
<div><span>WORKBENCH 03 / V3 CO-DESIGN</span><h4>FP8、DualPipe 与 MTP 改的是三种不同成本</h4></div>
|
||||||
|
<p>调度台只显示方向:bubble 与通信是 toy schedule;精度角色和 MTP 生命周期来自 V3 报告。</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="codesign-grid">
|
||||||
|
<div class="schedule-workbench">
|
||||||
|
<div class="subhead"><span>PIPELINE</span><h5>依赖等待怎样暴露设备空闲?</h5></div>
|
||||||
|
<div class="choice-row" role="group" aria-label="选择教学 pipeline schedule">
|
||||||
|
<button type="button" data-schedule="onef1b" class="active"><b>1F1B-LIKE</b><small>单向填充</small></button>
|
||||||
|
<button type="button" data-schedule="dual"><b>DUAL-ENDED TOY</b><small>两端注入 + overlap</small></button>
|
||||||
|
</div>
|
||||||
|
<div class="control-grid three compact">
|
||||||
|
<label><span>PP stages <output data-stages-label>8</output></span><input data-stages type="range" min="2" max="16" step="2" value="8" /></label>
|
||||||
|
<label><span>micro-batches <output data-micro-label>20</output></span><input data-micro type="range" min="4" max="64" step="2" value="20" /></label>
|
||||||
|
<label><span>comm / compute <output data-comm-label>0.60×</output></span><input data-comm-ratio type="range" min="0" max="150" value="60" /></label>
|
||||||
|
</div>
|
||||||
|
<div class="pipeline-strip" data-pipeline-strip aria-label="教学流水线时隙"></div>
|
||||||
|
<div class="metric-grid two">
|
||||||
|
<article><span>TOY BUBBLE</span><b data-bubble>25.9%</b><p>只由 stages / micro-batches 构造</p></article>
|
||||||
|
<article class="dark"><span>EXPOSED COMM</span><b data-exposed-comm>45.0%</b><p>重叠假设后的教学比例</p></article>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="precision-workbench">
|
||||||
|
<div class="subhead"><span>PRECISION CONTRACT</span><h5>一个“FP8”标签不够描述训练</h5></div>
|
||||||
|
<div class="choice-row" role="group" aria-label="选择精度配方">
|
||||||
|
<button type="button" data-precision="bf16"><b>BF16</b><small>统一高精度基线</small></button>
|
||||||
|
<button type="button" data-precision="naive"><b>NAIVE FP8</b><small>错误的一键强转</small></button>
|
||||||
|
<button type="button" data-precision="mixed" class="active"><b>V3 MIXED</b><small>角色分治</small></button>
|
||||||
|
</div>
|
||||||
|
<div class="dtype-contract">
|
||||||
|
<article><span>GEMM INPUTS</span><b data-gemm-dtype>FP8 / tiled scale</b></article>
|
||||||
|
<article><span>ACCUMULATION</span><b data-accum-dtype>FP32-assisted</b></article>
|
||||||
|
<article><span>MASTER / OPTIMIZER</span><b data-master-dtype>FP32 / BF16 roles</b></article>
|
||||||
|
<article><span>SENSITIVE OPS</span><b data-sensitive-dtype>BF16 / FP32</b></article>
|
||||||
|
</div>
|
||||||
|
<div class="risk-meter"><span>NUMERICAL RISK</span><div><i data-risk-bar></i></div><b data-risk-label>MANAGED</b></div>
|
||||||
|
<p data-precision-explain>V3 把主要 GEMM、缩放、累加、master state、敏感算子和通信分别设计;“混合精度”才是完整对象。</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="mtp-workbench">
|
||||||
|
<div class="subhead"><span>MULTI-TOKEN PREDICTION</span><h5>同一模块在生命周期中有三种角色</h5></div>
|
||||||
|
<div class="choice-row mtp" role="group" aria-label="选择 MTP 角色">
|
||||||
|
<button type="button" data-mtp-role="off"><b>OFF</b><small>只做 NTP</small></button>
|
||||||
|
<button type="button" data-mtp-role="train" class="active"><b>TRAIN</b><small>额外未来监督</small></button>
|
||||||
|
<button type="button" data-mtp-role="draft"><b>SPECULATIVE</b><small>候选草稿</small></button>
|
||||||
|
</div>
|
||||||
|
<div class="mtp-flow" data-mtp-flow></div>
|
||||||
|
<div class="metric-grid three">
|
||||||
|
<article><span>SUPERVISION DEPTH</span><b data-mtp-supervision>2 token targets</b><p>主 NTP + 一层 V3-style MTP</p></article>
|
||||||
|
<article><span>MAIN MODEL COST</span><b data-mtp-main-cost>TRAIN + MODULE</b><p>推理可丢弃额外 module</p></article>
|
||||||
|
<article class="dark"><span>AUTOREGRESSIVE?</span><b data-mtp-ar>YES</b><p>草稿仍需验证,不取代主模型</p></article>
|
||||||
|
</div>
|
||||||
|
<p class="boundary-copy" data-mtp-explain>V3 的顺序 MTP 首先增加训练信号;报告明确推理可直接丢弃,也可复用于 speculative decoding。</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="evidence-boundary">
|
||||||
|
<b>TOY SCHEDULE / SOURCE-GROUNDED ROLE CONTRACT</b>
|
||||||
|
<p data-codesign-boundary>本台没有复跑 DualPipe 或 FP8 kernel;它只阻止把三个独立机制压成“V3 训练便宜”一句话。</p>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="ds-panel" data-ds-panel="rl" hidden>
|
||||||
|
<div class="panel-intro">
|
||||||
|
<div><span>WORKBENCH 04 / POLICY GRADIENT</span><h4>奖励相同、回答长度和归一方式都会改变“谁被学得更多”</h4></div>
|
||||||
|
<p>DAPO 与 Dr.GRPO 是 R1 之后的公开研究,不是 DeepSeek 已披露的 R1 内部配方;这里用 toy weight 暴露目标函数偏差。</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="preset-row rl-presets" role="group" aria-label="选择 rollout 奖励预设">
|
||||||
|
<button type="button" data-reward-preset="mixed" class="active"><b>MIXED</b><small>0 · 1 · .7 · .2</small></button>
|
||||||
|
<button type="button" data-reward-preset="same"><b>ALL SAME</b><small>零组内信号</small></button>
|
||||||
|
<button type="button" data-reward-preset="longwrong"><b>LONG WRONG</b><small>长度偏差探针</small></button>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="rollout-controls">
|
||||||
|
{[
|
||||||
|
["1", "0", "512"],
|
||||||
|
["2", "1", "768"],
|
||||||
|
["3", ".7", "1536"],
|
||||||
|
["4", ".2", "3072"],
|
||||||
|
].map(([index, reward, length]) => (
|
||||||
|
<label>
|
||||||
|
<span>y{index}</span>
|
||||||
|
<small>reward</small><input data-rollout-reward type="number" min="-1" max="1" step=".1" value={reward} />
|
||||||
|
<small>tokens</small><input data-rollout-length type="number" min="64" max="8192" step="64" value={length} />
|
||||||
|
</label>
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="choice-row algorithm-row" role="group" aria-label="选择策略梯度教学算法">
|
||||||
|
<button type="button" data-rl-algorithm="grpo" class="active"><b>GRPO</b><small>std norm · response avg</small></button>
|
||||||
|
<button type="button" data-rl-algorithm="dapo"><b>DAPO-STYLE</b><small>token loss · dynamic filter</small></button>
|
||||||
|
<button type="button" data-rl-algorithm="dr"><b>Dr.GRPO</b><small>no std · fixed denominator</small></button>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="control-grid three compact">
|
||||||
|
<label><span>clip low <output data-clip-low-label>0.20</output></span><input data-clip-low type="range" min="5" max="40" value="20" /></label>
|
||||||
|
<label><span>clip high <output data-clip-high-label>0.20</output></span><input data-clip-high type="range" min="5" max="80" value="20" /></label>
|
||||||
|
<label><span>overlong threshold <output data-overlong-label>4096</output></span><input data-overlong type="range" min="512" max="8192" step="512" value="4096" /></label>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="rl-stage">
|
||||||
|
<div class="advantage-table">
|
||||||
|
<div class="head"><span>ROLLOUT</span><span>R</span><span>LENGTH</span><span>Â</span><span>TOY WEIGHT</span></div>
|
||||||
|
<div data-advantage-rows></div>
|
||||||
|
</div>
|
||||||
|
<div class="rl-diagnosis">
|
||||||
|
<span data-algorithm-name>GRPO / ORIGINAL FAMILY</span>
|
||||||
|
<h5 data-signal-state>GROUP-RELATIVE SIGNAL</h5>
|
||||||
|
<p data-algorithm-explain>奖励按组均值和标准差归一;response-level loss 再按各自长度平均,可能改变长短回答的 Token 权重。</p>
|
||||||
|
<dl>
|
||||||
|
<div><dt>reward mean</dt><dd data-reward-mean>0.475</dd></div>
|
||||||
|
<div><dt>reward std</dt><dd data-reward-std>0.396</dd></div>
|
||||||
|
<div><dt>effective samples</dt><dd data-effective>4 / 4</dd></div>
|
||||||
|
<div><dt>provenance</dt><dd data-provenance>DeepSeekMath / R1</dd></div>
|
||||||
|
</dl>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="pipeline-switch">
|
||||||
|
<div class="subhead"><span>TRAINING IDENTITY</span><h5>R1-Zero、R1 与 Distill 不是三个名字相近的同一步</h5></div>
|
||||||
|
<div class="choice-row" role="group" aria-label="选择 R1 训练身份">
|
||||||
|
<button type="button" data-r1-mode="zero" class="active"><b>R1-ZERO</b><small>base → rule RL</small></button>
|
||||||
|
<button type="button" data-r1-mode="r1"><b>R1</b><small>cold start → multi-stage</small></button>
|
||||||
|
<button type="button" data-r1-mode="distill"><b>DISTILL</b><small>teacher traces → SFT</small></button>
|
||||||
|
</div>
|
||||||
|
<div class="r1-mode-flow" data-r1-mode-flow></div>
|
||||||
|
<p data-r1-mode-explain>R1-Zero 从 V3 Base 直接进行规则奖励 GRPO;“无 reasoning SFT”不等于“无预训练知识”。</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="evidence-boundary">
|
||||||
|
<b data-rl-boundary-title>TOY GRADIENT WEIGHT / NOT A TRAINING REPLAY</b>
|
||||||
|
<p data-rl-boundary>优势公式与论文定义对齐;权重只用来显示归一和长度方向,没有 policy ratio、真实 token probability、KL 或 optimizer。</p>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<figcaption>
|
||||||
|
<span>证据分层</span>
|
||||||
|
KV 元素与组合参数按公式;FP8/MTP/R1 角色来自官方报告;bubble、通信、数值风险和梯度 weight 为本站教学模型。
|
||||||
|
</figcaption>
|
||||||
|
</figure>
|
||||||
|
|
||||||
|
<script>
|
||||||
|
const roots = document.querySelectorAll<HTMLElement>("[data-deepseek-lab]");
|
||||||
|
roots.forEach((root) => {
|
||||||
|
const one = <T extends Element>(selector: string) => root.querySelector<T>(selector);
|
||||||
|
const all = <T extends Element>(selector: string) => [...root.querySelectorAll<T>(selector)];
|
||||||
|
const value = (selector: string) => Number(one<HTMLInputElement | HTMLSelectElement>(selector)?.value ?? 0);
|
||||||
|
const set = (selector: string, text: string) => {
|
||||||
|
const node = one<HTMLElement>(selector);
|
||||||
|
if (node) node.textContent = text;
|
||||||
|
};
|
||||||
|
const compact = (number: number) => {
|
||||||
|
if (number >= 1024 ** 4) return `${(number / 1024 ** 4).toFixed(2)} TiB`;
|
||||||
|
if (number >= 1024 ** 3) return `${(number / 1024 ** 3).toFixed(2)} GiB`;
|
||||||
|
if (number >= 1024 ** 2) return `${(number / 1024 ** 2).toFixed(2)} MiB`;
|
||||||
|
return `${number.toLocaleString()} B`;
|
||||||
|
};
|
||||||
|
|
||||||
|
const tabs = all<HTMLButtonElement>("[data-ds-tab]");
|
||||||
|
const panels = all<HTMLElement>("[data-ds-panel]");
|
||||||
|
const selectTab = (tab: HTMLButtonElement) => {
|
||||||
|
tabs.forEach((candidate) => {
|
||||||
|
const selected = candidate === tab;
|
||||||
|
candidate.setAttribute("aria-selected", String(selected));
|
||||||
|
candidate.tabIndex = selected ? 0 : -1;
|
||||||
|
});
|
||||||
|
panels.forEach((panel) => panel.hidden = panel.dataset.dsPanel !== tab.dataset.dsTab);
|
||||||
|
};
|
||||||
|
tabs.forEach((tab, index) => {
|
||||||
|
tab.addEventListener("click", () => selectTab(tab));
|
||||||
|
tab.addEventListener("keydown", (event) => {
|
||||||
|
if (!["ArrowLeft", "ArrowRight", "Home", "End"].includes(event.key)) return;
|
||||||
|
event.preventDefault();
|
||||||
|
let target = index;
|
||||||
|
if (event.key === "ArrowRight") target = (index + 1) % tabs.length;
|
||||||
|
if (event.key === "ArrowLeft") target = (index - 1 + tabs.length) % tabs.length;
|
||||||
|
if (event.key === "Home") target = 0;
|
||||||
|
if (event.key === "End") target = tabs.length - 1;
|
||||||
|
tabs[target].focus();
|
||||||
|
selectTab(tabs[target]);
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
const capacityPresets: Record<string, [number, number, number, number, number, string, string]> = {
|
||||||
|
dense: [1, 1, 0, 100, 1, "DENSE / ONE ACTIVE PATH", "Dense FFN 让总容量与每 Token 激活计算绑定;没有 expert dispatch,但容量增长会直接增加计算。"],
|
||||||
|
coarse: [8, 2, 0, 100, 2, "COARSE MOE / 8 CHOOSE 2", "少数大专家提供条件计算;每个专家仍覆盖较宽知识,跨节点路由开始产生通信。"],
|
||||||
|
deepseekmoe: [63, 7, 1, 25, 4, "DEEPSEEKMOE / FINE-GRAINED", "论文代表配置把标准专家切为 0.25×,使用 1 shared + 63 routed、激活 7 routed;具体规模随实验变化。"],
|
||||||
|
v3: [256, 8, 1, 12.5, 4, "DEEPSEEK-V3 / CAPACITY CONTRACT", "V3 每层含 1 个 shared 与 256 个 routed experts,每 Token 激活 8 个 routed;理论稀疏计算仍需要 expert dispatch/combine 与负载均衡。"],
|
||||||
|
};
|
||||||
|
const logChoose = (n: number, k: number) => {
|
||||||
|
const safeK = Math.min(k, n - k);
|
||||||
|
if (safeK <= 0) return 0;
|
||||||
|
let result = 0;
|
||||||
|
for (let index = 1; index <= safeK; index += 1) result += Math.log10(n - safeK + index) - Math.log10(index);
|
||||||
|
return result;
|
||||||
|
};
|
||||||
|
const renderCapacity = () => {
|
||||||
|
const experts = Math.max(1, Math.round(value("[data-experts]")));
|
||||||
|
const topk = Math.min(experts, Math.round(value("[data-topk]")));
|
||||||
|
const shared = Math.round(value("[data-shared]"));
|
||||||
|
const width = value("[data-expert-width]") / 100;
|
||||||
|
const nodes = Math.round(value("[data-ep-nodes]"));
|
||||||
|
const total = (experts + shared) * width;
|
||||||
|
const active = (topk + shared) * width;
|
||||||
|
const logComb = logChoose(experts, topk);
|
||||||
|
const remoteShare = nodes <= 1 ? 0 : (nodes - 1) / nodes;
|
||||||
|
const communication = topk * remoteShare;
|
||||||
|
const level = communication === 0 ? "NONE" : communication < .8 ? "LOW" : communication < 2 ? "MEDIUM" : "HIGH";
|
||||||
|
set("[data-experts-label]", String(experts));
|
||||||
|
set("[data-topk-label]", String(topk));
|
||||||
|
set("[data-shared-label]", String(shared));
|
||||||
|
set("[data-expert-width-label]", `${width.toFixed(3).replace(/0+$/, "").replace(/\.$/, "")}×`);
|
||||||
|
set("[data-ep-label]", String(nodes));
|
||||||
|
set("[data-total-capacity]", `${total.toFixed(1)}× FFN`);
|
||||||
|
set("[data-active-compute]", `${active.toFixed(2)}× FFN`);
|
||||||
|
set("[data-combinations]", topk === 0 ? "1 route" : logComb < 6 ? Math.round(10 ** logComb).toLocaleString() : `≈ 10^${Math.floor(logComb)}`);
|
||||||
|
set("[data-communication]", level);
|
||||||
|
const map = one<HTMLElement>("[data-expert-map]");
|
||||||
|
if (map) {
|
||||||
|
const visible = Math.min(24, experts);
|
||||||
|
map.innerHTML = [
|
||||||
|
...Array.from({ length: shared }, (_, index) => `<i class="shared active"><b>S${index + 1}</b><small>shared</small></i>`),
|
||||||
|
...Array.from({ length: visible }, (_, index) => `<i class="${index < topk ? "active" : ""}"><b>E${index + 1}</b><small>${index < topk ? "route" : "idle"}</small></i>`),
|
||||||
|
].join("");
|
||||||
|
}
|
||||||
|
};
|
||||||
|
all<HTMLButtonElement>("[data-capacity-preset]").forEach((button) => button.addEventListener("click", () => {
|
||||||
|
const preset = capacityPresets[button.dataset.capacityPreset ?? "v3"];
|
||||||
|
const selectors = ["[data-experts]", "[data-topk]", "[data-shared]", "[data-expert-width]", "[data-ep-nodes]"];
|
||||||
|
selectors.forEach((selector, index) => {
|
||||||
|
const input = one<HTMLInputElement>(selector);
|
||||||
|
if (input) input.value = String(preset[index]);
|
||||||
|
});
|
||||||
|
set("[data-capacity-name]", preset[5]);
|
||||||
|
set("[data-capacity-explain]", preset[6]);
|
||||||
|
all("[data-capacity-preset]").forEach((candidate) => candidate.classList.toggle("active", candidate === button));
|
||||||
|
renderCapacity();
|
||||||
|
}));
|
||||||
|
all<HTMLInputElement>("[data-experts],[data-topk],[data-shared],[data-expert-width],[data-ep-nodes]").forEach((input) => input.addEventListener("input", renderCapacity));
|
||||||
|
|
||||||
|
const renderCache = () => {
|
||||||
|
const layers = value("[data-cache-layers]");
|
||||||
|
const context = value("[data-context-length]");
|
||||||
|
const batch = value("[data-batch]");
|
||||||
|
const bytes = value("[data-cache-bytes]");
|
||||||
|
const heads = value("[data-heads]");
|
||||||
|
const headDim = value("[data-head-dim]");
|
||||||
|
const groups = Math.min(value("[data-kv-groups]"), heads);
|
||||||
|
const latent = value("[data-latent]");
|
||||||
|
const rope = value("[data-rope-dim]");
|
||||||
|
const mha = 2 * heads * headDim;
|
||||||
|
const gqa = 2 * groups * headDim;
|
||||||
|
const mla = latent + rope;
|
||||||
|
const total = (elements: number) => elements * layers * context * batch * bytes;
|
||||||
|
const max = Math.max(mha, gqa, mla);
|
||||||
|
set("[data-cache-layers-label]", String(layers));
|
||||||
|
set("[data-batch-label]", String(batch));
|
||||||
|
set("[data-heads-label]", String(heads));
|
||||||
|
set("[data-head-dim-label]", String(headDim));
|
||||||
|
set("[data-kv-groups-label]", String(groups));
|
||||||
|
set("[data-latent-label]", String(latent));
|
||||||
|
set("[data-rope-dim-label]", String(rope));
|
||||||
|
set("[data-mha-elements]", mha.toLocaleString());
|
||||||
|
set("[data-gqa-elements]", gqa.toLocaleString());
|
||||||
|
set("[data-mla-elements]", mla.toLocaleString());
|
||||||
|
set("[data-rope-cache]", String(rope));
|
||||||
|
set("[data-selected-cache]", compact(total(mla)));
|
||||||
|
set("[data-mha-cache]", compact(total(mha)));
|
||||||
|
set("[data-cache-reduction]", `${((1 - mla / mha) * 100).toFixed(1)}%`);
|
||||||
|
const bars: [string, number][] = [["[data-mha-bar]", mha], ["[data-gqa-bar]", gqa], ["[data-mla-bar]", mla]];
|
||||||
|
bars.forEach(([selector, number]) => {
|
||||||
|
const node = one<HTMLElement>(selector);
|
||||||
|
if (node) node.style.width = `${Math.max(2, number / max * 100)}%`;
|
||||||
|
});
|
||||||
|
set("[data-cache-boundary]", `当前是 ${layers} 层 × ${context.toLocaleString()} Token × batch ${batch} × ${bytes} byte;GiB 未加入 allocator、page、quantization metadata 或 kernel workspace。`);
|
||||||
|
};
|
||||||
|
all<HTMLInputElement | HTMLSelectElement>("[data-cache-layers],[data-context-length],[data-batch],[data-cache-bytes],[data-heads],[data-head-dim],[data-kv-groups],[data-latent],[data-rope-dim]").forEach((control) => control.addEventListener("input", renderCache));
|
||||||
|
|
||||||
|
let schedule = "onef1b";
|
||||||
|
let precision = "mixed";
|
||||||
|
let mtpRole = "train";
|
||||||
|
const renderSchedule = () => {
|
||||||
|
const stages = value("[data-stages]");
|
||||||
|
const micro = value("[data-micro]");
|
||||||
|
const comm = value("[data-comm-ratio]") / 100;
|
||||||
|
const effectiveStages = schedule === "dual" ? Math.max(1, stages / 2) : stages;
|
||||||
|
const bubble = (effectiveStages - 1) / (micro + effectiveStages - 1);
|
||||||
|
const overlap = schedule === "dual" ? .78 : .25;
|
||||||
|
const exposed = Math.max(0, comm * (1 - overlap));
|
||||||
|
set("[data-stages-label]", String(stages));
|
||||||
|
set("[data-micro-label]", String(micro));
|
||||||
|
set("[data-comm-label]", `${comm.toFixed(2)}×`);
|
||||||
|
set("[data-bubble]", `${(bubble * 100).toFixed(1)}%`);
|
||||||
|
set("[data-exposed-comm]", `${(exposed * 100).toFixed(1)}%`);
|
||||||
|
const strip = one<HTMLElement>("[data-pipeline-strip]");
|
||||||
|
if (strip) strip.innerHTML = Array.from({ length: 30 }, (_, index) => {
|
||||||
|
const warm = index < effectiveStages - 1 || index >= 30 - (effectiveStages - 1);
|
||||||
|
const commSlot = !warm && index % (schedule === "dual" ? 7 : 4) === 0;
|
||||||
|
return `<i class="${warm ? "bubble" : commSlot ? "comm" : "compute"}"><small>${warm ? "idle" : commSlot ? "a2a" : index % 2 ? "B" : "F"}</small></i>`;
|
||||||
|
}).join("");
|
||||||
|
};
|
||||||
|
const precisionDetails: Record<string, [string, string, string, string, number, string, string]> = {
|
||||||
|
bf16: ["BF16", "BF16/FP32", "FP32", "BF16/FP32", 28, "LOWER", "BF16 基线保留更宽动态范围,但增加存储、带宽与高密度算术成本。"],
|
||||||
|
naive: ["FP8 / one scale", "FP8", "FP8", "FP8", 96, "CRITICAL", "一键把输入、累加、master state 和敏感算子全转 FP8 会暴露溢出、舍入与更新失真;这不是 V3 配方。"],
|
||||||
|
mixed: ["FP8 / tiled scale", "FP32-assisted", "FP32 / BF16 roles", "BF16 / FP32", 46, "MANAGED", "V3 把主要 GEMM、缩放、累加、master state、敏感算子和通信分别设计;“混合精度”才是完整对象。"],
|
||||||
|
};
|
||||||
|
const renderPrecision = () => {
|
||||||
|
const detail = precisionDetails[precision];
|
||||||
|
set("[data-gemm-dtype]", detail[0]);
|
||||||
|
set("[data-accum-dtype]", detail[1]);
|
||||||
|
set("[data-master-dtype]", detail[2]);
|
||||||
|
set("[data-sensitive-dtype]", detail[3]);
|
||||||
|
set("[data-risk-label]", detail[5]);
|
||||||
|
set("[data-precision-explain]", detail[6]);
|
||||||
|
const bar = one<HTMLElement>("[data-risk-bar]");
|
||||||
|
if (bar) bar.style.width = `${detail[4]}%`;
|
||||||
|
};
|
||||||
|
const renderMtp = () => {
|
||||||
|
const flow = one<HTMLElement>("[data-mtp-flow]");
|
||||||
|
const details: Record<string, [string, string, string, string, string]> = {
|
||||||
|
off: ["1 token target", "MAIN ONLY", "YES", "hₜ → next token", "关闭 MTP 后只有主 next-token loss;这不是 V3 报告采用的预训练目标。"],
|
||||||
|
train: ["2 token targets", "TRAIN + MODULE", "YES", "hₜ → t+1 · MTP₁(hₜ,t+1) → t+2", "V3 的顺序 MTP 首先增加训练信号;报告明确推理可直接丢弃,也可复用于 speculative decoding。"],
|
||||||
|
draft: ["draft + verify", "EXTRA DRAFT COST", "YES", "MTP draft → main model verify → accept/reject", "把 MTP module 当草稿器仍需主模型验证;验收率与 kernel 决定是否真实加速。"],
|
||||||
|
};
|
||||||
|
const detail = details[mtpRole];
|
||||||
|
set("[data-mtp-supervision]", detail[0]);
|
||||||
|
set("[data-mtp-main-cost]", detail[1]);
|
||||||
|
set("[data-mtp-ar]", detail[2]);
|
||||||
|
set("[data-mtp-explain]", detail[4]);
|
||||||
|
if (flow) flow.innerHTML = detail[3].split("→").map((item, index, array) => `<b>${item.trim()}</b>${index < array.length - 1 ? "<i>→</i>" : ""}`).join("");
|
||||||
|
};
|
||||||
|
all<HTMLButtonElement>("[data-schedule]").forEach((button) => button.addEventListener("click", () => {
|
||||||
|
schedule = button.dataset.schedule ?? "onef1b";
|
||||||
|
all("[data-schedule]").forEach((candidate) => candidate.classList.toggle("active", candidate === button));
|
||||||
|
renderSchedule();
|
||||||
|
}));
|
||||||
|
all<HTMLInputElement>("[data-stages],[data-micro],[data-comm-ratio]").forEach((control) => control.addEventListener("input", renderSchedule));
|
||||||
|
all<HTMLButtonElement>("[data-precision]").forEach((button) => button.addEventListener("click", () => {
|
||||||
|
precision = button.dataset.precision ?? "mixed";
|
||||||
|
all("[data-precision]").forEach((candidate) => candidate.classList.toggle("active", candidate === button));
|
||||||
|
renderPrecision();
|
||||||
|
}));
|
||||||
|
all<HTMLButtonElement>("[data-mtp-role]").forEach((button) => button.addEventListener("click", () => {
|
||||||
|
mtpRole = button.dataset.mtpRole ?? "train";
|
||||||
|
all("[data-mtp-role]").forEach((candidate) => candidate.classList.toggle("active", candidate === button));
|
||||||
|
renderMtp();
|
||||||
|
}));
|
||||||
|
|
||||||
|
let algorithm = "grpo";
|
||||||
|
let r1Mode = "zero";
|
||||||
|
const rewardPresets: Record<string, [number[], number[]]> = {
|
||||||
|
mixed: [[0, 1, .7, .2], [512, 768, 1536, 3072]],
|
||||||
|
same: [[1, 1, 1, 1], [512, 768, 1536, 3072]],
|
||||||
|
longwrong: [[1, .7, .2, 0], [384, 768, 2048, 6144]],
|
||||||
|
};
|
||||||
|
const renderRl = () => {
|
||||||
|
const rewards = all<HTMLInputElement>("[data-rollout-reward]").map((input) => Number(input.value));
|
||||||
|
const lengths = all<HTMLInputElement>("[data-rollout-length]").map((input) => Math.max(1, Number(input.value)));
|
||||||
|
const mean = rewards.reduce((sum, number) => sum + number, 0) / rewards.length;
|
||||||
|
const variance = rewards.reduce((sum, number) => sum + (number - mean) ** 2, 0) / rewards.length;
|
||||||
|
const std = Math.sqrt(variance);
|
||||||
|
const threshold = value("[data-overlong]");
|
||||||
|
const centered = rewards.map((reward) => reward - mean);
|
||||||
|
const advantages = centered.map((number) => algorithm === "dr" ? number : std > 1e-8 ? number / std : 0);
|
||||||
|
const weights = advantages.map((advantage, index) => {
|
||||||
|
if (algorithm === "dr") return advantage / threshold * 1024;
|
||||||
|
if (algorithm === "dapo") return lengths[index] > threshold ? 0 : advantage * lengths[index] / Math.max(...lengths);
|
||||||
|
return advantage / lengths[index] * 1024;
|
||||||
|
});
|
||||||
|
const effective = algorithm === "dapo" ? rewards.filter((reward, index) => Math.abs(reward - mean) > 1e-8 && lengths[index] <= threshold).length : rewards.length;
|
||||||
|
const rows = one<HTMLElement>("[data-advantage-rows]");
|
||||||
|
if (rows) rows.innerHTML = rewards.map((reward, index) => {
|
||||||
|
const width = Math.min(100, Math.abs(weights[index]) / Math.max(...weights.map(Math.abs), .001) * 100);
|
||||||
|
return `<div><b>y${index + 1}</b><span>${reward.toFixed(2)}</span><span>${lengths[index].toLocaleString()}</span><span>${advantages[index].toFixed(2)}</span><span class="${weights[index] >= 0 ? "positive" : "negative"}"><i style="width:${width}%"></i>${weights[index].toFixed(2)}</span></div>`;
|
||||||
|
}).join("");
|
||||||
|
const details: Record<string, [string, string, string]> = {
|
||||||
|
grpo: ["GRPO / ORIGINAL FAMILY", "奖励按组均值和标准差归一;response-level loss 再按各自长度平均,可能改变长短回答的 Token 权重。", "DeepSeekMath / R1"],
|
||||||
|
dapo: ["DAPO-STYLE / FOLLOW-UP", "这里用 token-level 聚合方向和 overlong filter 展示 DAPO 的两个修正;真实 DAPO 还包含 Clip-Higher 与 Dynamic Sampling。", "DAPO · arXiv:2503.14476"],
|
||||||
|
dr: ["Dr.GRPO / FOLLOW-UP", "去掉组 std normalization,并用固定全局长度分母,暴露原 GRPO 的 response-length 与 question-difficulty 尺度问题。", "Dr.GRPO · arXiv:2503.20783"],
|
||||||
|
};
|
||||||
|
const detail = details[algorithm];
|
||||||
|
set("[data-algorithm-name]", detail[0]);
|
||||||
|
set("[data-algorithm-explain]", detail[1]);
|
||||||
|
set("[data-provenance]", detail[2]);
|
||||||
|
set("[data-reward-mean]", mean.toFixed(3));
|
||||||
|
set("[data-reward-std]", std.toFixed(3));
|
||||||
|
set("[data-effective]", `${effective} / ${rewards.length}`);
|
||||||
|
set("[data-signal-state]", std < 1e-8 ? "ZERO GROUP SIGNAL" : algorithm === "dapo" && effective < rewards.length ? "FILTER / RESAMPLE" : "GROUP-RELATIVE SIGNAL");
|
||||||
|
set("[data-clip-low-label]", (value("[data-clip-low]") / 100).toFixed(2));
|
||||||
|
set("[data-clip-high-label]", (value("[data-clip-high]") / 100).toFixed(2));
|
||||||
|
set("[data-overlong-label]", String(threshold));
|
||||||
|
set("[data-rl-boundary]", std < 1e-8
|
||||||
|
? "四条奖励完全相同:组内中心化后优势为零。DAPO Dynamic Sampling 会过滤此类无梯度组并补采,但 rollout 成本不会消失。"
|
||||||
|
: "优势公式与论文定义对齐;weight 只显示归一和长度方向,没有 policy ratio、真实 token probability、KL 或 optimizer。");
|
||||||
|
};
|
||||||
|
all<HTMLButtonElement>("[data-reward-preset]").forEach((button) => button.addEventListener("click", () => {
|
||||||
|
const preset = rewardPresets[button.dataset.rewardPreset ?? "mixed"];
|
||||||
|
all<HTMLInputElement>("[data-rollout-reward]").forEach((input, index) => { input.value = String(preset[0][index]); });
|
||||||
|
all<HTMLInputElement>("[data-rollout-length]").forEach((input, index) => { input.value = String(preset[1][index]); });
|
||||||
|
all("[data-reward-preset]").forEach((candidate) => candidate.classList.toggle("active", candidate === button));
|
||||||
|
renderRl();
|
||||||
|
}));
|
||||||
|
all<HTMLInputElement>("[data-rollout-reward],[data-rollout-length],[data-clip-low],[data-clip-high],[data-overlong]").forEach((control) => control.addEventListener("input", renderRl));
|
||||||
|
all<HTMLButtonElement>("[data-rl-algorithm]").forEach((button) => button.addEventListener("click", () => {
|
||||||
|
algorithm = button.dataset.rlAlgorithm ?? "grpo";
|
||||||
|
all("[data-rl-algorithm]").forEach((candidate) => candidate.classList.toggle("active", candidate === button));
|
||||||
|
renderRl();
|
||||||
|
}));
|
||||||
|
const r1Details: Record<string, [string[], string]> = {
|
||||||
|
zero: [["V3 BASE", "RULE REWARD", "GRPO", "R1-ZERO"], "R1-Zero 从 V3 Base 直接进行规则奖励 GRPO;“无 reasoning SFT”不等于“无预训练知识”。"],
|
||||||
|
r1: [["V3 BASE", "COLD START", "REASONING RL", "SFT MIX", "GENERAL RL"], "正式 R1 用 cold start 修可读性与语言,再经 reasoning RL、rejection/SFT mix 和通用 RL;它不是纯 RL 单阶段。"],
|
||||||
|
distill: [["R1 TEACHER", "≈800K FILTERED TRACES", "SFT", "1.5B–70B STUDENTS"], "报告中的蒸馏学生主要学习 R1 生成/筛选轨迹,没有重演同一大规模 RL 探索过程。"],
|
||||||
|
};
|
||||||
|
const renderR1 = () => {
|
||||||
|
const detail = r1Details[r1Mode];
|
||||||
|
const flow = one<HTMLElement>("[data-r1-mode-flow]");
|
||||||
|
if (flow) flow.innerHTML = detail[0].map((item, index) => `<b>${item}</b>${index < detail[0].length - 1 ? "<i>→</i>" : ""}`).join("");
|
||||||
|
set("[data-r1-mode-explain]", detail[1]);
|
||||||
|
};
|
||||||
|
all<HTMLButtonElement>("[data-r1-mode]").forEach((button) => button.addEventListener("click", () => {
|
||||||
|
r1Mode = button.dataset.r1Mode ?? "zero";
|
||||||
|
all("[data-r1-mode]").forEach((candidate) => candidate.classList.toggle("active", candidate === button));
|
||||||
|
renderR1();
|
||||||
|
}));
|
||||||
|
|
||||||
|
renderCapacity();
|
||||||
|
renderCache();
|
||||||
|
renderSchedule();
|
||||||
|
renderPrecision();
|
||||||
|
renderMtp();
|
||||||
|
renderRl();
|
||||||
|
renderR1();
|
||||||
|
});
|
||||||
|
</script>
|
||||||
|
|
||||||
|
<style>
|
||||||
|
.ds-lab { max-width:100%; margin:36px 0; overflow:hidden; box-sizing:border-box; border:1px solid var(--line-strong); background:var(--paper); box-shadow:0 24px 70px rgba(8,18,30,.16); }
|
||||||
|
.ds-lab-head { display:grid; grid-template-columns:1fr 1fr; gap:38px; align-items:end; padding:30px; color:#f4efe7; background:linear-gradient(135deg,#172437,#263b52); }
|
||||||
|
.ds-lab-head p:first-child,.panel-intro span,.subhead span { margin:0 0 9px; color:#d49a68; font:600 .55rem var(--mono); letter-spacing:.13em; }
|
||||||
|
.ds-lab-head h3 { margin:0; color:#fff; font:650 clamp(1.3rem,2.6vw,2.05rem) var(--serif); }
|
||||||
|
.ds-lab-head > p { margin:0; color:#b7c1cc; font-size:.65rem; line-height:1.8; }
|
||||||
|
.ds-tabs { display:grid; grid-template-columns:repeat(4,1fr); border-bottom:1px solid var(--line-strong); }
|
||||||
|
.ds-tabs button { min-height:94px; display:grid; grid-template-columns:32px 1fr; grid-template-rows:auto auto; gap:4px 8px; padding:18px; border:0; border-right:1px solid var(--line); color:var(--ink); background:var(--paper-raised); text-align:left; cursor:pointer; }
|
||||||
|
.ds-tabs button:last-child { border-right:0; }.ds-tabs button[aria-selected="true"] { color:#fff; background:var(--navy); }
|
||||||
|
.ds-tabs span { grid-row:1/3; color:var(--copper); font:600 .6rem var(--mono); }.ds-tabs b { font:650 .84rem var(--serif); }.ds-tabs small { color:var(--muted); font:.49rem var(--mono); }
|
||||||
|
.ds-tabs button[aria-selected="true"] small { color:#aab7c5; }
|
||||||
|
.ds-panel { padding:30px; }.panel-intro { display:grid; grid-template-columns:1.1fr .9fr; gap:34px; align-items:end; margin-bottom:24px; }
|
||||||
|
.panel-intro h4 { margin:0; font:650 clamp(1.2rem,2.5vw,1.85rem) var(--serif); }.panel-intro > p { margin:0; color:var(--muted); font-size:.65rem; line-height:1.8; }
|
||||||
|
.preset-row,.choice-row { display:grid; grid-template-columns:repeat(4,1fr); gap:8px; margin:18px 0; }
|
||||||
|
.preset-row button,.choice-row button { min-height:63px; padding:12px; border:1px solid var(--line); color:var(--ink); background:var(--paper-raised); text-align:left; cursor:pointer; }
|
||||||
|
.preset-row button.active,.choice-row button.active { color:#fff; border-color:var(--navy); background:var(--navy); }
|
||||||
|
.preset-row b,.choice-row b { display:block; font:.59rem var(--mono); }.preset-row small,.choice-row small { display:block; margin-top:7px; color:var(--muted); font:.48rem var(--mono); }
|
||||||
|
.preset-row button.active small,.choice-row button.active small { color:#b6c2ce; }
|
||||||
|
.control-grid { display:grid; gap:9px; margin:16px 0 22px; }.control-grid.five { grid-template-columns:repeat(5,1fr); }.control-grid.four { grid-template-columns:repeat(4,1fr); }.control-grid.three { grid-template-columns:repeat(3,1fr); }
|
||||||
|
.control-grid label { display:flex; flex-direction:column; justify-content:space-between; gap:9px; min-height:72px; padding:12px; border:1px solid var(--line); background:var(--paper-raised); color:var(--muted); font:.52rem var(--mono); }
|
||||||
|
.control-grid label > span { display:flex; justify-content:space-between; gap:6px; }.control-grid input[type="range"] { width:100%; accent-color:var(--copper); }.control-grid select { min-height:31px; border:1px solid var(--line); color:var(--ink); background:var(--paper); font:.55rem var(--mono); }
|
||||||
|
.capacity-stage { display:grid; grid-template-columns:130px 25px 1fr 25px 145px; gap:12px; align-items:center; margin:24px 0; padding:22px; background:#172437; color:#fff; }
|
||||||
|
.capacity-stage > i { color:#d49a68; text-align:center; font-style:normal; }.router-box { padding:18px; border:1px solid #405269; }.router-box span { color:#d49a68; font:.5rem var(--mono); }.router-box b { display:block; margin:10px 0; font:650 1.2rem var(--serif); }.router-box small { color:#98a9ba; font:.48rem var(--mono); }
|
||||||
|
.expert-map { display:grid; grid-template-columns:repeat(8,1fr); gap:5px; }.expert-map i { min-height:43px; display:flex; flex-direction:column; justify-content:center; align-items:center; border:1px solid #405269; color:#7d8fa3; font-style:normal; }.expert-map i.active { color:#fff; border-color:#d49a68; background:rgba(212,154,104,.18); }.expert-map i.shared { background:#8c5538; }
|
||||||
|
.expert-map b { font:.5rem var(--mono); }.expert-map small { margin-top:4px; font:.4rem var(--mono); }
|
||||||
|
.metric-grid { display:grid; border-top:1px solid var(--line); border-left:1px solid var(--line); margin:20px 0; }.metric-grid.four { grid-template-columns:repeat(4,1fr); }.metric-grid.three { grid-template-columns:repeat(3,1fr); }.metric-grid.two { grid-template-columns:repeat(2,1fr); }
|
||||||
|
.metric-grid article { min-height:125px; padding:18px; border-right:1px solid var(--line); border-bottom:1px solid var(--line); }.metric-grid span { color:var(--copper); font:.49rem var(--mono); letter-spacing:.08em; }.metric-grid b { display:block; margin:14px 0 9px; font:650 1.12rem var(--serif); }.metric-grid p { margin:0; color:var(--muted); font-size:.53rem; line-height:1.55; }.metric-grid .dark { color:#fff; background:var(--navy); }.metric-grid .dark p { color:#9fb0c2; }
|
||||||
|
.evidence-boundary { margin-top:18px; padding:18px 20px; border-left:3px solid var(--copper); background:rgba(193,124,68,.08); }.evidence-boundary b { font:.57rem var(--mono); letter-spacing:.08em; }.evidence-boundary p { margin:8px 0 0; color:var(--muted); font-size:.61rem; line-height:1.7; }
|
||||||
|
.cache-comparison { display:grid; grid-template-columns:repeat(3,1fr); border:1px solid var(--line); }.cache-comparison article { padding:20px; border-right:1px solid var(--line); }.cache-comparison article:last-child { border-right:0; }.cache-comparison article.selected { background:rgba(193,124,68,.08); }.cache-comparison span { color:var(--copper); font:.54rem var(--mono); }.cache-comparison b { display:block; margin:12px 0 4px; font:650 1.2rem var(--serif); }.cache-comparison small { color:var(--muted-light); font:.48rem var(--mono); }.cache-comparison article > div { height:9px; margin:16px 0; background:var(--line); }.cache-comparison i { display:block; height:100%; background:var(--copper); }.cache-comparison p { color:var(--muted); font-size:.57rem; line-height:1.6; }
|
||||||
|
.absorption-grid { display:grid; grid-template-columns:1fr 1fr; margin-top:18px; border:1px solid var(--line); }.absorption-grid > div { padding:22px; }.absorption-grid > div:first-child { border-right:1px solid var(--line); }.absorption-grid span { color:var(--copper); font:.5rem var(--mono); }.absorption-grid code { display:block; margin:14px 0; color:var(--ink); font:.73rem var(--mono); }.absorption-grid div > div { display:flex; align-items:center; gap:8px; }.absorption-grid div > div b { padding:9px; border:1px solid var(--line); font:.58rem var(--mono); }.absorption-grid div > div i { color:var(--copper); font:.48rem var(--mono); font-style:normal; }.absorption-grid p { color:var(--muted); font-size:.57rem; line-height:1.65; }
|
||||||
|
.codesign-grid { display:grid; grid-template-columns:1.15fr .85fr; gap:16px; }.schedule-workbench,.precision-workbench,.mtp-workbench { padding:22px; border:1px solid var(--line); background:var(--paper-raised); }.subhead h5 { margin:0; font:650 1.05rem var(--serif); }.choice-row { grid-template-columns:repeat(2,1fr); }.choice-row.mtp { grid-template-columns:repeat(3,1fr); }
|
||||||
|
.control-grid.compact label { min-height:62px; padding:10px; }.pipeline-strip { display:grid; grid-template-columns:repeat(15,1fr); gap:3px; min-height:92px; margin:18px 0; }.pipeline-strip i { display:flex; justify-content:center; align-items:center; min-height:36px; font-style:normal; }.pipeline-strip i.compute { color:#fff; background:#314a62; }.pipeline-strip i.comm { color:#fff; background:#b36c45; }.pipeline-strip i.bubble { color:#8a96a3; border:1px dashed var(--line-strong); }.pipeline-strip small { font:.4rem var(--mono); }
|
||||||
|
.dtype-contract { display:grid; grid-template-columns:1fr 1fr; border-top:1px solid var(--line); border-left:1px solid var(--line); }.dtype-contract article { min-height:85px; padding:13px; border-right:1px solid var(--line); border-bottom:1px solid var(--line); }.dtype-contract span { color:var(--copper); font:.45rem var(--mono); }.dtype-contract b { display:block; margin-top:13px; font:.61rem var(--mono); }.risk-meter { margin:18px 0; }.risk-meter > span { color:var(--muted); font:.48rem var(--mono); }.risk-meter > div { height:9px; margin:8px 0; background:var(--line); }.risk-meter i { display:block; height:100%; background:var(--copper); }.risk-meter > b { font:.56rem var(--mono); }.precision-workbench > p,.boundary-copy { color:var(--muted); font-size:.57rem; line-height:1.65; }
|
||||||
|
.mtp-workbench { margin-top:16px; }.mtp-flow,.r1-mode-flow { display:flex; align-items:center; flex-wrap:wrap; gap:8px; padding:18px; background:#172437; color:#fff; }.mtp-flow b,.r1-mode-flow b { padding:10px; border:1px solid #405269; font:.53rem var(--mono); }.mtp-flow i,.r1-mode-flow i { color:#d49a68; font-style:normal; }
|
||||||
|
.rollout-controls { display:grid; grid-template-columns:repeat(4,1fr); gap:8px; }.rollout-controls label { display:grid; grid-template-columns:38px 1fr; gap:7px; padding:12px; border:1px solid var(--line); background:var(--paper-raised); }.rollout-controls span { grid-row:1/5; color:var(--copper); font:650 .75rem var(--serif); }.rollout-controls small { color:var(--muted); font:.45rem var(--mono); }.rollout-controls input { min-width:0; width:100%; border:1px solid var(--line); background:var(--paper); color:var(--ink); font:.54rem var(--mono); }.algorithm-row { grid-template-columns:repeat(3,1fr); }
|
||||||
|
.rl-stage { display:grid; grid-template-columns:1.15fr .85fr; border:1px solid var(--line); }.advantage-table { overflow:auto; }.advantage-table .head,.advantage-table [data-advantage-rows] > div { min-width:590px; display:grid; grid-template-columns:70px 60px 80px 60px 1fr; gap:8px; align-items:center; padding:12px 15px; border-bottom:1px solid var(--line); }.advantage-table .head { color:#fff; background:#172437; font:.45rem var(--mono); }.advantage-table [data-advantage-rows] > div { font:.55rem var(--mono); }.advantage-table [data-advantage-rows] > div:last-child { border-bottom:0; }.advantage-table [data-advantage-rows] span:last-child { position:relative; min-height:25px; display:flex; align-items:center; padding-left:7px; overflow:hidden; }.advantage-table [data-advantage-rows] span:last-child i { position:absolute; inset:0 auto 0 0; opacity:.18; }.advantage-table .positive i { background:#2e876a; }.advantage-table .negative i { background:#b1504c; }
|
||||||
|
.rl-diagnosis { padding:22px; color:#fff; background:#172437; }.rl-diagnosis > span { color:#d49a68; font:.5rem var(--mono); }.rl-diagnosis h5 { margin:14px 0; font:650 1.15rem var(--serif); }.rl-diagnosis > p { color:#b7c1cc; font-size:.58rem; line-height:1.7; }.rl-diagnosis dl { margin:18px 0 0; }.rl-diagnosis dl div { display:flex; justify-content:space-between; gap:12px; padding:8px 0; border-top:1px solid #405269; }.rl-diagnosis dt { color:#8fa2b5; font:.45rem var(--mono); }.rl-diagnosis dd { margin:0; font:.52rem var(--mono); text-align:right; }
|
||||||
|
.pipeline-switch { margin-top:18px; padding:22px; border:1px solid var(--line); }.pipeline-switch > p { color:var(--muted); font-size:.6rem; line-height:1.7; }
|
||||||
|
.ds-lab figcaption { display:flex; gap:18px; padding:16px 22px; border-top:1px solid var(--line); color:var(--muted); font-size:.54rem; line-height:1.6; }.ds-lab figcaption span { color:var(--copper); font:.49rem var(--mono); }
|
||||||
|
@media (max-width:1000px) {
|
||||||
|
.control-grid.five { grid-template-columns:repeat(3,1fr); }.codesign-grid { grid-template-columns:1fr; }.capacity-stage { grid-template-columns:110px 20px 1fr; }.capacity-stage > i:nth-of-type(2),.capacity-stage .output { display:none; }
|
||||||
|
}
|
||||||
|
@media (max-width:720px) {
|
||||||
|
.ds-lab-head,.panel-intro { grid-template-columns:1fr; }.ds-tabs { grid-template-columns:1fr; }.ds-tabs button { min-height:74px; border-right:0; border-bottom:1px solid var(--line); }
|
||||||
|
.ds-panel { padding:16px; }.preset-row,.control-grid.five,.control-grid.four,.control-grid.three,.cache-comparison,.absorption-grid,.metric-grid.four,.metric-grid.three,.metric-grid.two,.rollout-controls,.algorithm-row,.rl-stage { grid-template-columns:1fr; }
|
||||||
|
.capacity-stage { grid-template-columns:1fr; }.capacity-stage > i { transform:rotate(90deg); }.capacity-stage > i:nth-of-type(2),.capacity-stage .output { display:block; }.expert-map { grid-template-columns:repeat(6,1fr); }
|
||||||
|
.cache-comparison article,.absorption-grid > div:first-child { border-right:0; border-bottom:1px solid var(--line); }.pipeline-strip { grid-template-columns:repeat(10,1fr); }.choice-row { grid-template-columns:1fr 1fr; }
|
||||||
|
.rl-stage { border:0; gap:12px; }.advantage-table,.rl-diagnosis { border:1px solid var(--line); }.ds-lab figcaption { flex-direction:column; }
|
||||||
|
}
|
||||||
|
</style>
|
||||||
@@ -14,6 +14,13 @@ const milestones = [
|
|||||||
bridge: "把容量扩张和每 Token 计算量分开。",
|
bridge: "把容量扩张和每 Token 计算量分开。",
|
||||||
url: "https://arxiv.org/abs/2401.06066",
|
url: "https://arxiv.org/abs/2401.06066",
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
year: "2024.02",
|
||||||
|
model: "DeepSeekMath",
|
||||||
|
idea: "从大规模数学数据工程走到 GRPO:用同题多条回答的相对奖励,省去独立 critic。",
|
||||||
|
bridge: "R1 的推理 RL 不是突然出现;算法与可验证数据的预演在这里发生。",
|
||||||
|
url: "https://arxiv.org/abs/2402.03300",
|
||||||
|
},
|
||||||
{
|
{
|
||||||
year: "2024.05",
|
year: "2024.05",
|
||||||
model: "DeepSeek-V2",
|
model: "DeepSeek-V2",
|
||||||
@@ -31,10 +38,18 @@ const milestones = [
|
|||||||
{
|
{
|
||||||
year: "2025.01",
|
year: "2025.01",
|
||||||
model: "DeepSeek-R1",
|
model: "DeepSeek-R1",
|
||||||
idea: "R1-Zero 展示纯大规模 RL 可涌现推理;R1 用冷启动数据修复可读性与稳定性。",
|
idea: "R1-Zero 从强 V3 Base 直接做规则奖励 RL、没有 reasoning SFT;R1 再用冷启动与多阶段训练修复可读性和广度。",
|
||||||
bridge: "从“模仿答案”转向用可验证奖励塑造推理策略。",
|
bridge: "从“模仿答案”转向用可验证奖励塑造推理策略。",
|
||||||
url: "https://arxiv.org/abs/2501.12948",
|
url: "https://arxiv.org/abs/2501.12948",
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
year: "2025.03",
|
||||||
|
model: "DAPO / Dr.GRPO",
|
||||||
|
idea: "公开后续研究分别暴露 clipping、采样、截断、响应长度与题目难度归一偏差。",
|
||||||
|
bridge: "复现不是 R1 的内部 recipe,而是一台看清 RL 优化对象的显微镜。",
|
||||||
|
url: "https://arxiv.org/abs/2503.14476",
|
||||||
|
followup: true,
|
||||||
|
},
|
||||||
{
|
{
|
||||||
year: "2025.12",
|
year: "2025.12",
|
||||||
model: "DeepSeek-V3.2",
|
model: "DeepSeek-V3.2",
|
||||||
@@ -45,7 +60,7 @@ const milestones = [
|
|||||||
{
|
{
|
||||||
year: "2026.06",
|
year: "2026.06",
|
||||||
model: "DeepSeek-V4",
|
model: "DeepSeek-V4",
|
||||||
idea: "围绕百万 Token 上下文效率继续扩展,成为 K3 报告直接比较的开放前沿之一。",
|
idea: "CSA/HCA 构成异构长状态,mHC 约束深层残差,Muon 与数值边界共同支撑百万 Token。",
|
||||||
bridge: "长上下文不再只是位置外推,而是注意力、训练与服务的全系统问题。",
|
bridge: "长上下文不再只是位置外推,而是注意力、训练与服务的全系统问题。",
|
||||||
url: "https://arxiv.org/abs/2606.19348",
|
url: "https://arxiv.org/abs/2606.19348",
|
||||||
},
|
},
|
||||||
@@ -54,13 +69,13 @@ const milestones = [
|
|||||||
|
|
||||||
<div class="deepseek-lineage">
|
<div class="deepseek-lineage">
|
||||||
{milestones.map((item, index) => (
|
{milestones.map((item, index) => (
|
||||||
<a href={item.url} class="lineage-row" rel="noreferrer">
|
<a href={item.url} class:list={["lineage-row", { followup: item.followup }]} rel="noreferrer">
|
||||||
<div class="lineage-time">
|
<div class="lineage-time">
|
||||||
<span>{item.year}</span>
|
<span>{item.year}</span>
|
||||||
<i aria-hidden="true"></i>
|
<i aria-hidden="true"></i>
|
||||||
</div>
|
</div>
|
||||||
<div class="lineage-main">
|
<div class="lineage-main">
|
||||||
<span>DS / {String(index + 1).padStart(2, "0")}</span>
|
<span>{item.followup ? "PUBLIC FOLLOW-UP" : `DS / ${String(index + 1).padStart(2, "0")}`}</span>
|
||||||
<h3>{item.model}</h3>
|
<h3>{item.model}</h3>
|
||||||
<p>{item.idea}</p>
|
<p>{item.idea}</p>
|
||||||
</div>
|
</div>
|
||||||
@@ -93,6 +108,11 @@ const milestones = [
|
|||||||
background: var(--paper-raised);
|
background: var(--paper-raised);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
.lineage-row.followup {
|
||||||
|
border-right: 3px solid var(--copper);
|
||||||
|
background: var(--copper-pale);
|
||||||
|
}
|
||||||
|
|
||||||
.lineage-time {
|
.lineage-time {
|
||||||
display: grid;
|
display: grid;
|
||||||
grid-template-columns: 1fr 12px;
|
grid-template-columns: 1fr 12px;
|
||||||
|
|||||||
@@ -0,0 +1,110 @@
|
|||||||
|
export const deepseekLedgers = [
|
||||||
|
["Q01", "Dense 坐标系", "为什么 DeepSeek LLM 不是可跳过的序章?", "它固定 tokenizer、数据、架构和 scaling 试验的起点;并不单独证明后续所有设计。"],
|
||||||
|
["Q02", "参数角色", "671B / 37B 各表示什么?", "total 是装下的容量,activated 是每 Token 经过的专家参数子集;都不等于端到端 FLOPs。"],
|
||||||
|
["Q03", "专家粒度", "为什么切小专家还要多选?", "DeepSeekMoE 把每个专家缩成 1/m,总数和激活数同乘 m,近似保持专家计算。"],
|
||||||
|
["Q04", "Shared expert", "为什么把公共知识单独隔离?", "始终激活的 shared experts 减少 routed experts 重复;它们仍然要付激活计算。"],
|
||||||
|
["Q05", "通信税", "为什么稀疏 FLOPs 不等于便宜?", "路由会产生 dispatch/combine、跨节点 all-to-all、负载长尾和权重访问。"],
|
||||||
|
["Q06", "均衡", "aux-loss-free 到底去掉了什么?", "V3 的 expert bias 影响选择、不进入最终 gate weight;仍有 sequence-wise auxiliary loss 防极端失衡。"],
|
||||||
|
["Q07", "KV 状态", "为什么 V2 把服务状态当架构问题?", "权重只装一次,KV 随请求、层、Token 增长,直接限制并发和长上下文。"],
|
||||||
|
["Q08", "Attention 压缩", "MQA、GQA、MLA 的差别是什么?", "MQA/GQA 共享 K/V 头;MLA 联合低秩压缩 K/V 内容并在计算中恢复。"],
|
||||||
|
["Q09", "矩阵吸收", "MLA 为什么不必恢复完整 content K/V?", "无位置项时可利用矩阵乘结合律,把 K/V 上投影吸收到 query/output 投影。"],
|
||||||
|
["Q10", "位置分叉", "为什么要 decoupled RoPE?", "RoPE 会阻断固定权重吸收,所以 V2 另设小 RoPE query/key 分支,并缓存 key。"],
|
||||||
|
["Q11", "FP8 合同", "“FP8 训练”包含哪些角色?", "主要 GEMM 用 FP8,并配细粒度缩放、较高精度累加和高精度敏感算子;不是全路径 FP8。"],
|
||||||
|
["Q12", "Pipeline", "DualPipe 隐藏了什么?", "从两端注入 micro-batch,让成对前后向 chunk 与通信重叠;它减少而非清零 bubble。"],
|
||||||
|
["Q13", "MTP", "训练和推理各怎样使用 MTP?", "顺序模块增加未来 Token 监督;推理可丢弃,也可复用于 speculative draft。"],
|
||||||
|
["Q14", "GRPO", "去掉 critic 后还剩什么?", "policy/reference、同题多 rollout、reward/verifier、clip 和 KL;主要省掉 value model。"],
|
||||||
|
["Q15", "可验证奖励", "R1-Zero 的奖励能覆盖哪些任务?", "论文用数学、代码、逻辑等规则可验证域和格式奖励;开放任务仍是限制。"],
|
||||||
|
["Q16", "纯 RL 实验", "R1-Zero 究竟证明了什么?", "强 V3 Base 在无 reasoning SFT 时可被规则奖励继续塑造;不等于没有预训练先验。"],
|
||||||
|
["Q17", "R1 pipeline", "正式 R1 为什么不是纯 RL?", "cold start → reasoning RL → rejection/SFT mix → general RL,分别修可读性、广度和对齐。"],
|
||||||
|
["Q18", "蒸馏", "学生为什么不是“小号 R1-Zero”?", "1.5B–70B 学生主要对约 800K 教师样本做 SFT,没有重演同一 RL 探索。"],
|
||||||
|
["Q19", "复现反查", "DAPO / Dr.GRPO 修的是哪类问题?", "它们处理 clip、采样、聚合、截断、长度和难度偏差;是后续研究,不是已披露 R1 内部配方。"],
|
||||||
|
["Q20", "DSA", "可学习 indexer 为什么不是固定稀疏?", "indexer 对历史内容评分,主 attention 只读 top-k;它需要专门训练,也可能漏检。"],
|
||||||
|
["Q21", "Agent 数据", "V3.2 怎样把 reasoning 放进环境?", "specialist distillation + mixed RL;环境、工具、任务、解法和 verifier 构成数据闭环。"],
|
||||||
|
["Q22", "V4 Attention", "CSA 与 HCA 各压什么?", "CSA 先压缩再稀疏 top-k;HCA 更强压缩后保留全部 compressed entries。"],
|
||||||
|
["Q23", "V4 稳定化", "mHC、Muon、QK/RMSNorm、clamp 各管什么?", "它们分别管残差混合、矩阵更新、attention 尺度和 FFN 极值,不能合成一个技巧。"],
|
||||||
|
["Q24", "K3 对照", "哪些是祖先,哪些只是同题新解?", "DeepSeekMoE/MLA 有明确继承;QB、KDA、AttnRes、SiTU、MOPD 多是新解或同期路线。"],
|
||||||
|
] as const;
|
||||||
|
|
||||||
|
export const deepseekWaves = [
|
||||||
|
["W1", "2024.01", "Dense 坐标", "先固定数据、tokenizer、训练与 scaling 对照,后面的结构收益才有可比起点。"],
|
||||||
|
["W2", "2024.01", "稀疏容量", "细粒度 routed experts 加 shared experts,把总容量与单 Token 激活计算第一次清楚分开。"],
|
||||||
|
["W3", "2024.05", "服务状态", "MLA 不再只优化训练 FLOPs,而是直接改写随请求增长的 KV Cache。"],
|
||||||
|
["W4", "2024.12", "协同训练", "V3 把路由、FP8、MTP、pipeline 与通信写成同一套训练合同。"],
|
||||||
|
["W5", "2024.02 → 2025.01", "推理 RL", "DeepSeekMath 先减掉 critic;R1-Zero 再隔离规则奖励,R1 恢复可读性与通用性。"],
|
||||||
|
["W6", "2025.03", "复现显微镜", "DAPO 与 Dr.GRPO 暴露 clipping、采样、截断、长度归一和题目难度偏差。"],
|
||||||
|
["W7", "2025.12", "稀疏检索", "V3.2 用学习型 indexer 选历史,再让主 attention 读取 top-k。"],
|
||||||
|
["W8", "2025.12", "Agent 环境", "推理从静态题目进入含工具、状态转移和 verifier 的交互数据闭环。"],
|
||||||
|
["W9", "2026.06", "异构长状态", "V4 用 CSA 与 HCA 处理不同时间尺度,并联合 mHC、Muon 和数值约束。"],
|
||||||
|
["W10", "2026.07", "K3 对照", "继承图必须允许没有箭头:相同的百万上下文目标,可以有完全不同的状态机器。"],
|
||||||
|
] as const;
|
||||||
|
|
||||||
|
export const deepseekBranches = [
|
||||||
|
["代码与专家", "DeepSeek-Coder → Coder-V2 → ESFT", "代码数据配方、continued pretraining 与只微调相关专家,说明 MoE 的价值不只在通用模型参数量。", "https://arxiv.org/abs/2406.11931"],
|
||||||
|
["数学与证明", "DeepSeekMath → Prover-V1.5 → Prover-V2", "从数学语料、GRPO 走到 formal proof feedback、subgoal decomposition 与可验证证明搜索。", "https://arxiv.org/abs/2504.21801"],
|
||||||
|
["视觉与压缩", "DeepSeek-VL/VL2 → Janus → OCR", "理解、生成与光学压缩形成另一条主干;它们不应被挤进纯文本 V2→V4 时间线。", "https://arxiv.org/abs/2412.10302"],
|
||||||
|
["系统实现", "DeepEP → DualPipe → DeepGEMM / FlashMLA", "论文里的稀疏计算、流水线、FP8 与 MLA 最终必须落到可调用的通信和 kernel 实现。", "https://github.com/deepseek-ai/DeepEP"],
|
||||||
|
["条件记忆", "Engram", "把可查表的静态模式从动态网络中分离,增加一条不同于 MoE 与 attention 的稀疏轴。", "https://arxiv.org/abs/2601.07372"],
|
||||||
|
] as const;
|
||||||
|
|
||||||
|
export const deepseekPaperChain = [
|
||||||
|
["01", "1991", "Adaptive Mixtures of Local Experts", "https://proceedings.neurips.cc/paper/1991/hash/59b90e1005a220e2ebc542eb9d950b1e-Abstract.html", "专家门控前史"],
|
||||||
|
["02", "2000", "Conditional Computation", "https://arxiv.org/abs/cs/0008102", "条件计算"],
|
||||||
|
["03", "2003", "A Neural Probabilistic Language Model", "https://www.jmlr.org/papers/v3/bengio03a.html", "Dense LM 坐标"],
|
||||||
|
["04", "2017", "Attention Is All You Need", "https://arxiv.org/abs/1706.03762", "Transformer 主干"],
|
||||||
|
["05", "2017", "Outrageously Large Neural Networks", "https://arxiv.org/abs/1701.06538", "稀疏 MoE"],
|
||||||
|
["06", "2017", "Proximal Policy Optimization", "https://arxiv.org/abs/1707.06347", "GRPO 对照"],
|
||||||
|
["07", "2018", "GPipe", "https://arxiv.org/abs/1811.06965", "Pipeline 前史"],
|
||||||
|
["08", "2018", "PipeDream", "https://arxiv.org/abs/1806.03377", "Pipeline schedule"],
|
||||||
|
["09", "2019", "Fast Transformer Decoding / MQA", "https://arxiv.org/abs/1911.02150", "KV 共享"],
|
||||||
|
["10", "2019", "Megatron-LM", "https://arxiv.org/abs/1909.08053", "模型并行"],
|
||||||
|
["11", "2019", "ZeRO", "https://arxiv.org/abs/1910.02054", "状态分片"],
|
||||||
|
["12", "2019", "RMSNorm", "https://arxiv.org/abs/1910.07467", "尺度控制"],
|
||||||
|
["13", "2020", "GShard", "https://arxiv.org/abs/2006.16668", "大规模 MoE"],
|
||||||
|
["14", "2020", "QK-Normalization", "https://arxiv.org/abs/2010.04245", "attention logit 稳定"],
|
||||||
|
["15", "2021", "Switch Transformers", "https://arxiv.org/abs/2101.03961", "coarse top-1 MoE"],
|
||||||
|
["16", "2021", "RoFormer / RoPE", "https://arxiv.org/abs/2104.09864", "MLA 位置分叉"],
|
||||||
|
["17", "2022", "ST-MoE", "https://arxiv.org/abs/2202.08906", "MoE 稳定性"],
|
||||||
|
["18", "2022", "DeepNet", "https://arxiv.org/abs/2203.00555", "深层残差"],
|
||||||
|
["19", "2022", "InstructGPT", "https://arxiv.org/abs/2203.02155", "SFT/RM/PPO 合同"],
|
||||||
|
["20", "2022", "FlashAttention", "https://arxiv.org/abs/2205.14135", "IO-aware exact attention"],
|
||||||
|
["21", "2022", "Process and Outcome Feedback", "https://arxiv.org/abs/2211.14275", "reasoning reward 前史"],
|
||||||
|
["22", "2022", "Self-Consistency", "https://arxiv.org/abs/2203.11171", "多采样聚合"],
|
||||||
|
["23", "2023", "GQA", "https://arxiv.org/abs/2305.13245", "KV 分组"],
|
||||||
|
["24", "2023", "Let's Verify Step by Step", "https://arxiv.org/abs/2305.20050", "verifier / PRM"],
|
||||||
|
["25", "2023", "Direct Preference Optimization", "https://arxiv.org/abs/2305.18290", "RL 外偏好路线"],
|
||||||
|
["26", "2023", "FlashAttention-2", "https://arxiv.org/abs/2307.08691", "attention kernel"],
|
||||||
|
["27", "2023", "PagedAttention / vLLM", "https://arxiv.org/abs/2309.06180", "KV 服务状态"],
|
||||||
|
["28", "2024", "DeepSeek LLM", "https://arxiv.org/abs/2401.02954", "Dense / scaling 基线"],
|
||||||
|
["29", "2024", "DeepSeek-Coder", "https://arxiv.org/abs/2401.14196", "代码数据旁支"],
|
||||||
|
["30", "2024", "DeepSeekMoE", "https://arxiv.org/abs/2401.06066", "细粒度 + shared"],
|
||||||
|
["31", "2024", "DeepSeekMath", "https://arxiv.org/abs/2402.03300", "数学数据 + GRPO"],
|
||||||
|
["32", "2024", "RLOO", "https://arxiv.org/abs/2402.14740", "critic-free 对照"],
|
||||||
|
["33", "2024", "DeepSeek-V2", "https://arxiv.org/abs/2405.04434", "MLA + MoE"],
|
||||||
|
["34", "2024", "Better & Faster LLMs via MTP", "https://arxiv.org/abs/2404.19737", "MTP 祖先"],
|
||||||
|
["35", "2024", "DeepSeek-Coder-V2", "https://arxiv.org/abs/2406.11931", "V2 continued pretrain"],
|
||||||
|
["36", "2024", "ESFT", "https://arxiv.org/abs/2407.01906", "专家特化微调"],
|
||||||
|
["37", "2024", "DeepSeek-Prover-V1.5", "https://arxiv.org/abs/2408.08152", "proof feedback RL"],
|
||||||
|
["38", "2024", "Hyper-Connections", "https://arxiv.org/abs/2409.19606", "mHC 前身"],
|
||||||
|
["39", "2024", "DeepSeek-V3", "https://arxiv.org/abs/2412.19437", "FP8 / DualPipe / MTP"],
|
||||||
|
["40", "2025", "DeepSeek-R1", "https://arxiv.org/abs/2501.12948", "R1-Zero / R1 / 蒸馏"],
|
||||||
|
["41", "2025", "Muon is Scalable for LLM Training", "https://arxiv.org/abs/2502.16982", "V4 optimizer 前史"],
|
||||||
|
["42", "2025", "DAPO", "https://arxiv.org/abs/2503.14476", "GRPO 工程修正"],
|
||||||
|
["43", "2025", "Understanding R1-Zero-Like Training", "https://arxiv.org/abs/2503.20783", "Dr.GRPO / 偏差"],
|
||||||
|
["44", "2025", "DeepSeek-Prover-V2", "https://arxiv.org/abs/2504.21801", "subgoal + RL"],
|
||||||
|
["45", "2025", "DeepEP", "https://github.com/deepseek-ai/DeepEP", "Expert Parallel kernel"],
|
||||||
|
["46", "2025", "DualPipe", "https://github.com/deepseek-ai/DualPipe", "V3/R1 pipeline 实现"],
|
||||||
|
["47", "2025", "DeepGEMM", "https://github.com/deepseek-ai/DeepGEMM", "FP8 GEMM 实现"],
|
||||||
|
["48", "2025", "DeepSeek-VL2", "https://arxiv.org/abs/2412.10302", "多模态理解旁支"],
|
||||||
|
["49", "2025", "Janus-Pro", "https://arxiv.org/abs/2501.17811", "统一理解/生成旁支"],
|
||||||
|
["50", "2025", "Kimi k1.5", "https://arxiv.org/abs/2501.12599", "同期 reasoning RL"],
|
||||||
|
["51", "2025", "Kimi K2", "https://arxiv.org/abs/2507.20534", "MLA / MoE / Muon 对照"],
|
||||||
|
["52", "2025", "Kimi Linear", "https://arxiv.org/abs/2510.26692", "KDA 前身"],
|
||||||
|
["53", "2025", "DeepSeek-V3.2", "https://arxiv.org/abs/2512.02556", "DSA + Agent"],
|
||||||
|
["54", "2025", "mHC", "https://arxiv.org/abs/2512.24880", "受约束 residual"],
|
||||||
|
["55", "2026", "Engram", "https://arxiv.org/abs/2601.07372", "条件记忆新稀疏轴"],
|
||||||
|
["56", "2026", "LatentMoE", "https://arxiv.org/abs/2601.18089", "K3 routed latent 前身"],
|
||||||
|
["57", "2026", "Attention Residuals", "https://arxiv.org/abs/2603.15031", "K3 深度路由"],
|
||||||
|
["58", "2026", "DeepSeek-V4", "https://arxiv.org/abs/2606.19348", "CSA / HCA / mHC / Muon"],
|
||||||
|
["59", "2026", "Kimi K3", "https://arxiv.org/abs/2607.24653", "对照锚点"],
|
||||||
|
["60", "2026", "Kimi K3 official repository", "https://github.com/MoonshotAI/Kimi-K3", "开放实现边界"],
|
||||||
|
] as const;
|
||||||
@@ -3896,6 +3896,60 @@ export const papers: Paper[] = [
|
|||||||
contribution: "公开 Pre-LN decoder 训练配置与模型族,为深度、归一化和复现提供对照。",
|
contribution: "公开 Pre-LN decoder 训练配置与模型族,为深度、归一化和复现提供对照。",
|
||||||
verified: true,
|
verified: true,
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
year: 2024,
|
||||||
|
title: "DeepSeek-Coder: When the Large Language Model Meets Programming — The Rise of Code Intelligence",
|
||||||
|
url: "https://arxiv.org/abs/2401.14196",
|
||||||
|
topics: ["数据", "推理", "评测"],
|
||||||
|
contribution: "公开代码语料构造、fill-in-the-blank 目标与 1B–33B 模型族,建立 DeepSeek 代码能力旁支。",
|
||||||
|
spotlight: "DeepSeek",
|
||||||
|
verified: true,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
year: 2024,
|
||||||
|
title: "DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence",
|
||||||
|
url: "https://arxiv.org/abs/2406.11931",
|
||||||
|
topics: ["数据", "MoE", "推理", "评测"],
|
||||||
|
contribution: "从 DeepSeek-V2 checkpoint continued pretrain,扩展代码语言、上下文和数学/代码能力。",
|
||||||
|
spotlight: "DeepSeek",
|
||||||
|
verified: true,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
year: 2024,
|
||||||
|
title: "Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models",
|
||||||
|
url: "https://arxiv.org/abs/2407.01906",
|
||||||
|
topics: ["MoE", "后训练"],
|
||||||
|
contribution: "ESFT 依据专家相关性只更新任务相关 routed experts,研究 MoE 专长怎样进入低成本微调。",
|
||||||
|
spotlight: "DeepSeek",
|
||||||
|
verified: true,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
year: 2024,
|
||||||
|
title: "DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search",
|
||||||
|
url: "https://arxiv.org/abs/2408.08152",
|
||||||
|
topics: ["推理", "后训练", "评测"],
|
||||||
|
contribution: "把 Lean proof assistant 的可验证反馈用于 formal proof RL,并结合 MCTS 扩展证明搜索。",
|
||||||
|
spotlight: "DeepSeek",
|
||||||
|
verified: true,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
year: 2025,
|
||||||
|
title: "DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition",
|
||||||
|
url: "https://arxiv.org/abs/2504.21801",
|
||||||
|
topics: ["推理", "后训练", "评测"],
|
||||||
|
contribution: "用递归子目标分解构造 cold-start 数据,再以可验证证明奖励强化 formal reasoning。",
|
||||||
|
spotlight: "DeepSeek",
|
||||||
|
verified: true,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
year: 2026,
|
||||||
|
title: "Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models",
|
||||||
|
url: "https://arxiv.org/abs/2601.07372",
|
||||||
|
topics: ["基础", "MoE", "数据", "推理服务"],
|
||||||
|
contribution: "Engram 把可查表的静态模式与动态计算分离,提出不同于 MoE 与 attention 的条件记忆稀疏轴。",
|
||||||
|
spotlight: "DeepSeek",
|
||||||
|
verified: true,
|
||||||
|
},
|
||||||
];
|
];
|
||||||
|
|
||||||
export const paperTopics: PaperTopic[] = [
|
export const paperTopics: PaperTopic[] = [
|
||||||
|
|||||||
+953
-61
File diff suppressed because it is too large
Load Diff
+27
-1
@@ -112,7 +112,7 @@ const paths = [
|
|||||||
<div class="hero-stats">
|
<div class="hero-stats">
|
||||||
<div><b>17</b><span>核心专题</span></div>
|
<div><b>17</b><span>核心专题</span></div>
|
||||||
<div><b>151</b><span>K3 报告来源</span></div>
|
<div><b>151</b><span>K3 报告来源</span></div>
|
||||||
<div><b>480</b><span>关键论文索引</span></div>
|
<div><b>486</b><span>关键论文索引</span></div>
|
||||||
<div><b>47p</b><span>K3 技术报告</span></div>
|
<div><b>47p</b><span>K3 技术报告</span></div>
|
||||||
</div>
|
</div>
|
||||||
</aside>
|
</aside>
|
||||||
@@ -126,6 +126,22 @@ const paths = [
|
|||||||
|
|
||||||
<section class="section compact release-section" id="new-chapters">
|
<section class="section compact release-section" id="new-chapters">
|
||||||
<div class="release-grid">
|
<div class="release-grid">
|
||||||
|
<a class="release-card deepseek-release" href="/deepseek/">
|
||||||
|
<div>
|
||||||
|
<p class="eyebrow"><span>NEW / DEEPSEEK ROUND 02</span> CAPACITY · STATE · SYSTEM · REASONING</p>
|
||||||
|
<h2>从 Dense 到百万上下文:每次创新都在偿还上一代最贵的一张账</h2>
|
||||||
|
<p>
|
||||||
|
用二十四张问题账和十次技术转向,从 DeepSeek LLM、MoE、V2 的 MLA 权重吸收,
|
||||||
|
走到 V3 的 FP8 / DualPipe / MTP、R1 与 DAPO / Dr.GRPO 反查、V3.2 Agent 环境和 V4 异构长状态。
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
<dl>
|
||||||
|
<div><dt>LINEAGE</dt><dd>1991 → 2026 · 10 次转向</dd></div>
|
||||||
|
<div><dt>NODES</dt><dd>60 个一手 / 官方节点</dd></div>
|
||||||
|
<div><dt>LAB</dt><dd>MoE · MLA · V3 协同 · RL 偏差</dd></div>
|
||||||
|
</dl>
|
||||||
|
<span class="release-arrow" aria-hidden="true">进入 DeepSeek 完整技术谱系 →</span>
|
||||||
|
</a>
|
||||||
<a class="release-card representation-release" href="/architecture/representation/">
|
<a class="release-card representation-release" href="/architecture/representation/">
|
||||||
<div>
|
<div>
|
||||||
<p class="eyebrow"><span>NEW / CHAPTER 03</span> TOKEN · POSITION · DEPTH · FFN</p>
|
<p class="eyebrow"><span>NEW / CHAPTER 03</span> TOKEN · POSITION · DEPTH · FFN</p>
|
||||||
@@ -599,6 +615,7 @@ const paths = [
|
|||||||
transition: transform 180ms ease, border-color 180ms ease;
|
transition: transform 180ms ease, border-color 180ms ease;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
.deepseek-release,
|
||||||
.representation-release,
|
.representation-release,
|
||||||
.inference-release,
|
.inference-release,
|
||||||
.agent-release,
|
.agent-release,
|
||||||
@@ -614,6 +631,14 @@ const paths = [
|
|||||||
min-height: 510px;
|
min-height: 510px;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
.deepseek-release {
|
||||||
|
background:
|
||||||
|
radial-gradient(circle at 80% 18%, rgba(159, 91, 52, .24), transparent 30%),
|
||||||
|
radial-gradient(circle at 61% 72%, rgba(35, 86, 84, .2), transparent 28%),
|
||||||
|
repeating-linear-gradient(90deg, transparent 0 48px, rgba(159, 91, 52, .04) 48px 49px),
|
||||||
|
var(--paper-raised);
|
||||||
|
}
|
||||||
|
|
||||||
.representation-release {
|
.representation-release {
|
||||||
background:
|
background:
|
||||||
radial-gradient(circle at 82% 18%, rgba(159, 91, 52, 0.22), transparent 31%),
|
radial-gradient(circle at 82% 18%, rgba(159, 91, 52, 0.22), transparent 31%),
|
||||||
@@ -760,6 +785,7 @@ const paths = [
|
|||||||
padding-bottom: 76px;
|
padding-bottom: 76px;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
.deepseek-release,
|
||||||
.representation-release,
|
.representation-release,
|
||||||
.inference-release,
|
.inference-release,
|
||||||
.alignment-release,
|
.alignment-release,
|
||||||
|
|||||||
@@ -15,7 +15,7 @@ const workstreams = [
|
|||||||
{ label: "表示、位置与残差高速公路", value: 81, next: "加入真实 hidden-state / norm traces、长上下文位置外推复现与更多深层稳定性消融" },
|
{ label: "表示、位置与残差高速公路", value: 81, next: "加入真实 hidden-state / norm traces、长上下文位置外推复现与更多深层稳定性消融" },
|
||||||
{ label: "Scaling Laws", value: 74, next: "加入真实拟合复现、置信区间与更多模型族对照" },
|
{ label: "Scaling Laws", value: 74, next: "加入真实拟合复现、置信区间与更多模型族对照" },
|
||||||
{ label: "数据工程与预训练配方", value: 73, next: "逐图精读 FineWeb / DCLM,加入真实去重与 mixture traces" },
|
{ label: "数据工程与预训练配方", value: 73, next: "逐图精读 FineWeb / DCLM,加入真实去重与 mixture traces" },
|
||||||
{ label: "DeepSeek 专题", value: 71, next: "补 R1 / DAPO 的逐图训练轨迹与复现对照" },
|
{ label: "DeepSeek 专题", value: 83, next: "加入真实专家负载、MLA kernel、RL 训练 traces 与独立复现" },
|
||||||
{ label: "指令微调与人类偏好", value: 75, next: "加入真实偏好分歧样本、RM 长度偏置与 PPO/DPO 小模型复现" },
|
{ label: "指令微调与人类偏好", value: 75, next: "加入真实偏好分歧样本、RM 长度偏置与 PPO/DPO 小模型复现" },
|
||||||
{ label: "推理与测试时扩展", value: 76, next: "真实模型采样曲线、PRM 案例与逐篇图表精读" },
|
{ label: "推理与测试时扩展", value: 76, next: "真实模型采样曲线、PRM 案例与逐篇图表精读" },
|
||||||
{ label: "工具使用与长程 Agent", value: 74, next: "补真实环境 traces、cross-harness 对照、Agent RL 训练曲线与安全案例" },
|
{ label: "工具使用与长程 Agent", value: 74, next: "补真实环境 traces、cross-harness 对照、Agent RL 训练曲线与安全案例" },
|
||||||
@@ -50,7 +50,7 @@ const workstreams = [
|
|||||||
<div><dt>OVERALL</dt><dd>专题平均 {average}%</dd></div>
|
<div><dt>OVERALL</dt><dd>专题平均 {average}%</dd></div>
|
||||||
<div><dt>READABLE</dt><dd>{published} 个首版可读专题</dd></div>
|
<div><dt>READABLE</dt><dd>{published} 个首版可读专题</dd></div>
|
||||||
<div><dt>ACTIVE</dt><dd>{researching} 个研究/写作中</dd></div>
|
<div><dt>ACTIVE</dt><dd>{researching} 个研究/写作中</dd></div>
|
||||||
<div><dt>UPDATED</dt><dd>2026-07-29 10:26 CST</dd></div>
|
<div><dt>UPDATED</dt><dd>2026-07-29 11:10 CST</dd></div>
|
||||||
<div><dt>MODE</dt><dd>持续迭代,不锁死版本</dd></div>
|
<div><dt>MODE</dt><dd>持续迭代,不锁死版本</dd></div>
|
||||||
</dl>
|
</dl>
|
||||||
</div>
|
</div>
|
||||||
@@ -97,11 +97,12 @@ const workstreams = [
|
|||||||
<article><span>✓</span><h3>K3 报告已结构化拆解</h3><p>47 页报告目录、151 条参考来源和架构/后训练/系统主线已经提取。</p></article>
|
<article><span>✓</span><h3>K3 报告已结构化拆解</h3><p>47 页报告目录、151 条参考来源和架构/后训练/系统主线已经提取。</p></article>
|
||||||
<article><span>✓</span><h3>17 专题知识图</h3><p>从语言模型基础到评测安全,包含先修依赖和三条贯穿案例。</p></article>
|
<article><span>✓</span><h3>17 专题知识图</h3><p>从语言模型基础到评测安全,包含先修依赖和三条贯穿案例。</p></article>
|
||||||
<article><span>✓</span><h3>编辑式网站系统</h3><p>响应式导航、章节模板、侧栏、进度、论文链和证据提示组件。</p></article>
|
<article><span>✓</span><h3>编辑式网站系统</h3><p>响应式导航、章节模板、侧栏、进度、论文链和证据提示组件。</p></article>
|
||||||
<article><span>✓</span><h3>五十五个原创交互视图</h3><p>K3、语言模型前史、Transformer、表示深度、DeepSeek、长上下文、MoE、推理、Agent、多模态,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
|
<article><span>✓</span><h3>五十九个原创交互视图</h3><p>K3、语言模型前史、Transformer、表示深度、DeepSeek 四联实验、长上下文、MoE、推理、Agent、多模态,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
|
||||||
<article><span>✓</span><h3>十七篇首版长文</h3><p>K3、语言模型前史、Transformer、表示/位置/残差、DeepSeek、Scaling、数据工程、长上下文、MoE、后训练、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全专题。</p></article>
|
<article><span>✓</span><h3>十七篇首版长文</h3><p>K3、语言模型前史、Transformer、表示/位置/残差、DeepSeek、Scaling、数据工程、长上下文、MoE、后训练、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全专题。</p></article>
|
||||||
<article><span>✓</span><h3>语言模型前史深度专题</h3><p>八张独立问题账、33 个正式节点、20 段长文与概率—向量—记忆—对齐四联实验。</p></article>
|
<article><span>✓</span><h3>语言模型前史深度专题</h3><p>八张独立问题账、33 个正式节点、20 段长文与概率—向量—记忆—对齐四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>Transformer 深度专题</h3><p>十张独立问题账、40 个正式节点、21 段正文与 QKV—Mask—多头位置—Block 成本四联实验。</p></article>
|
<article><span>✓</span><h3>Transformer 深度专题</h3><p>十张独立问题账、40 个正式节点、21 段正文与 QKV—Mask—多头位置—Block 成本四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>表示、位置与残差高速公路深度专题</h3><p>二十张问题账、66 个一手节点、DeepSeek/Kimi 双谱系,以及 Token—位置—Norm—Residual/FFN 四联实验。</p></article>
|
<article><span>✓</span><h3>表示、位置与残差高速公路深度专题</h3><p>二十张问题账、66 个一手节点、DeepSeek/Kimi 双谱系,以及 Token—位置—Norm—Residual/FFN 四联实验。</p></article>
|
||||||
|
<article><span>✓</span><h3>DeepSeek 技术谱系二轮深读</h3><p>二十四张问题账、十次技术转向、60 个一手/官方节点,以及稀疏容量—MLA 缓存—V3 协同—RL 偏差四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>Scaling Laws 深度专题</h3><p>九张账、29 个一手节点、DeepSeek/Kimi 双谱系与曲面—部署—复用—涌现四联实验。</p></article>
|
<article><span>✓</span><h3>Scaling Laws 深度专题</h3><p>九张账、29 个一手节点、DeepSeek/Kimi 双谱系与曲面—部署—复用—涌现四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>数据工程深度专题</h3><p>十二张账、31 个一手节点、DeepSeek/Kimi 双谱系与流水线—去重—混合—改写四联实验。</p></article>
|
<article><span>✓</span><h3>数据工程深度专题</h3><p>十二张账、31 个一手节点、DeepSeek/Kimi 双谱系与流水线—去重—混合—改写四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>长上下文深度专题</h3><p>五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。</p></article>
|
<article><span>✓</span><h3>长上下文深度专题</h3><p>五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。</p></article>
|
||||||
@@ -114,7 +115,7 @@ const workstreams = [
|
|||||||
<article><span>✓</span><h3>原生多模态深度专题</h3><p>十六张账、55 个一手节点、DeepSeek 三分支、Kimi 三代 MoonViT,以及 Token—连接器—光学压缩—视觉闭环四联实验。</p></article>
|
<article><span>✓</span><h3>原生多模态深度专题</h3><p>十六张账、55 个一手节点、DeepSeek 三分支、Kimi 三代 MoonViT,以及 Token—连接器—光学压缩—视觉闭环四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>推理服务与低成本部署深度专题</h3><p>十八本账、62 个一手节点、DeepSeek V2→V4 与 Mooncake→K3 双谱系,以及显存—阶段—推测—集群四联实验。</p></article>
|
<article><span>✓</span><h3>推理服务与低成本部署深度专题</h3><p>十八本账、62 个一手节点、DeepSeek V2→V4 与 Mooncake→K3 双谱系,以及显存—阶段—推测—集群四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>评测、安全与“到底强不强”深度专题</h3><p>二十二张账、80 个一手节点、DeepSeek/K3 评测协议谱系,以及指标—Judge—污染—系统安全四联实验。</p></article>
|
<article><span>✓</span><h3>评测、安全与“到底强不强”深度专题</h3><p>二十二张账、80 个一手节点、DeepSeek/K3 评测协议谱系,以及指标—Judge—污染—系统安全四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>480 篇关键论文索引</h3><p>新增 output embedding、ELMo、BLT、Fixup、ReZero、Hyper-Connections、mHC、xPos、FIRE、LongRoPE 等 30 个表示与深度节点。</p></article>
|
<article><span>✓</span><h3>486 篇关键论文索引</h3><p>新增 DeepSeek-Coder/Coder-V2、ESFT、Prover-V1.5/V2 与 Engram 6 个 DeepSeek 旁支节点。</p></article>
|
||||||
<article><span>✓</span><h3>公开仓库与自托管发布</h3><p>源码公开到 git.k1412.top,网站由不可变镜像、Compose Manager 与 HTTPS 交付。</p></article>
|
<article><span>✓</span><h3>公开仓库与自托管发布</h3><p>源码公开到 git.k1412.top,网站由不可变镜像、Compose Manager 与 HTTPS 交付。</p></article>
|
||||||
</div>
|
</div>
|
||||||
</section>
|
</section>
|
||||||
@@ -129,6 +130,7 @@ const workstreams = [
|
|||||||
</div>
|
</div>
|
||||||
<div class="queue-table">
|
<div class="queue-table">
|
||||||
<div class="head"><b>优先级</b><b>专题</b><b>本轮交付</b><b>完成闸门</b></div>
|
<div class="head"><b>优先级</b><b>专题</b><b>本轮交付</b><b>完成闸门</b></div>
|
||||||
|
<div><span>P0</span><strong>DeepSeek 三轮</strong><p>真实 expert load / MLA kernel → FP8 / pipeline traces → R1-like RL 小模型复现</p><em>运行证据 + 独立复现</em></div>
|
||||||
<div><span>P0</span><strong>Transformer 二轮</strong><p>多头电路逐图 → Pre/Post-LN 真实 traces → Flash/KV 配置与 kernel 对照</p><em>逐图笔记 + 实测边界</em></div>
|
<div><span>P0</span><strong>Transformer 二轮</strong><p>多头电路逐图 → Pre/Post-LN 真实 traces → Flash/KV 配置与 kernel 对照</p><em>逐图笔记 + 实测边界</em></div>
|
||||||
<div><span>P0</span><strong>表示、位置与残差二轮</strong><p>真实 hidden-state / norm traces → 长上下文位置外推 → mHC / AttnRes 深层稳定性消融</p><em>可复现实验 + 逐图笔记</em></div>
|
<div><span>P0</span><strong>表示、位置与残差二轮</strong><p>真实 hidden-state / norm traces → 长上下文位置外推 → mHC / AttnRes 深层稳定性消融</p><em>可复现实验 + 逐图笔记</em></div>
|
||||||
<div><span>P0</span><strong>语言模型前史二轮</strong><p>Kneser–Ney / LSTM / Bahdanau 逐图 → 真实小语料复现 → tokenizer 公平性</p><em>可复现实验 + 逐图笔记</em></div>
|
<div><span>P0</span><strong>语言模型前史二轮</strong><p>Kneser–Ney / LSTM / Bahdanau 逐图 → 真实小语料复现 → tokenizer 公平性</p><em>可复现实验 + 逐图笔记</em></div>
|
||||||
|
|||||||
Reference in New Issue
Block a user