feat: publish MoE deep dive
This commit is contained in:
+14
-8
@@ -6,11 +6,12 @@
|
||||
|
||||
| 工作流 | 状态 | 完成度 | 下一检查点 |
|
||||
|---|---:|---:|---|
|
||||
| 研究框架与规范 | 进行中 | 72% | 给 125 篇索引补充逐篇精读层级 |
|
||||
| 研究框架与规范 | 进行中 | 76% | 给 130 篇索引补充逐篇精读层级 |
|
||||
| 网站设计系统 | 进行中 | 89% | 打印样式与更多通用可视化组件 |
|
||||
| Kimi K3 深读 | 进行中 | 55% | 扩写 scaling / infra 逐图笔记 |
|
||||
| Transformer 基础 | 进行中 | 52% | 矩阵形状动画与手算练习 |
|
||||
| DeepSeek 专题 | 进行中 | 61% | GRPO 完整公式与训练轨迹推导 |
|
||||
| 稀疏计算与 MoE | 完成首版 | 74% | 真实负载 traces 与专家特化案例 |
|
||||
| 长上下文专题 | 完成首版 | 72% | 真实模型配置、内核细节与失败案例 |
|
||||
| 引用与事实检查 | 进行中 | 54% | 自动化外链复查与来源等级扩展 |
|
||||
| 开源仓库 | 已完成首版 | 100% | 持续提交研究与网站迭代 |
|
||||
@@ -25,21 +26,24 @@
|
||||
- [x] 提炼参考网站的编辑设计语言。
|
||||
- [x] 确认 `git.k1412.top` 为 Gitea/Forgejo 兼容服务且本机 HTTPS 凭据可用于既有仓库。
|
||||
- [x] 使用 Grok CLI 检索并形成约 95 篇一手论文的补充路线,主代理已回查关键来源。
|
||||
- [x] 完成首批 125 篇关键论文索引,覆盖 12 个专题与 Kimi/DeepSeek 聚光主线。
|
||||
- [x] 完成首批 130 篇关键论文索引,覆盖 12 个专题与 Kimi/DeepSeek 聚光主线。
|
||||
- [x] 完成可检索、可按专题筛选的论文库页面。
|
||||
- [x] 完成 K3、Transformer 基础、DeepSeek 谱系与长上下文四篇首版长文。
|
||||
- [x] 完成 K3 三轴架构、Self-Attention 实验、DeepSeek 谱系与长上下文成本实验室四张原创交互图。
|
||||
- [x] 完成 K3、Transformer 基础、DeepSeek 谱系、长上下文与 MoE 五篇首版长文。
|
||||
- [x] 完成 K3 三轴架构、Self-Attention 实验、DeepSeek 谱系、长上下文成本与 MoE 路由实验室五张原创交互图。
|
||||
- [x] 完成长上下文首版:五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。
|
||||
- [x] 核验 FlashAttention、DeepSeek-V2/V3.2/V4、Kimi Linear/K3 等六份论文原文,并建立长上下文研究账本。
|
||||
- [x] Astro 类型检查、生产构建、8 个内部路由和桌面/移动端视觉检查通过。
|
||||
- [x] 核验 Switch、ST-MoE、DeepSeekMoE、Loss-Free、V3、LatentMoE 与 K3 原文,并建立 MoE 研究账本。
|
||||
- [x] 完成 MoE 首版:六张账、19 篇一手论文、完整 DeepSeek/K3 主线与路由—容量—通信实验室。
|
||||
- [x] Astro 类型检查、生产构建、9 个内部路由和桌面/移动端视觉检查通过。
|
||||
- [x] 创建 `wuyang/llm-atlas` 公开仓库,匿名 API 确认 `private: false`。
|
||||
- [x] 本地生产镜像通过健康检查与全部 8 个页面路由烟雾测试。
|
||||
- [x] 本地生产镜像通过健康检查与全部 9 个页面路由烟雾测试。
|
||||
- [x] 通过 Unraid Compose Manager、Nginx Proxy Manager 与 HTTPS 发布首版。
|
||||
|
||||
## 正在进行
|
||||
|
||||
- [ ] MoE 路由模拟器与通信成本账本。
|
||||
- [ ] 推理模型与测试时扩展:CoT、verifier、GRPO、R1、Kimi k1.5 与 K3 MOPD。
|
||||
- [ ] 长上下文专题的真实模型配置对比、内核细节与失败案例二轮深化。
|
||||
- [ ] MoE 专题的真实集群 traces、专家特化案例与二轮外部证据。
|
||||
|
||||
## 研究账本
|
||||
|
||||
@@ -49,11 +53,13 @@
|
||||
| 2026-07-28 | K3 作为“汇流点”,不是课程起点 | 初学者可以先学基础,高阶读者可以从 K3 反向跳转 |
|
||||
| 2026-07-28 | 优先重绘论文图并标明“简化/改绘” | 图可缩放、可交互,也减少脱离上下文复制论文图片 |
|
||||
| 2026-07-28 | Grok 只用于线索扩展与交叉检查 | 正文事实必须回到论文、官方仓库或正式文档 |
|
||||
| 2026-07-28 | 首批论文库收录 125 篇,按问题与专题多标签组织 | 论文库承担发现入口,专题正文承担深度精读与机制复核 |
|
||||
| 2026-07-28 | 首批论文库收录 130 篇,按问题与专题多标签组织 | 论文库承担发现入口,专题正文承担深度精读与机制复核 |
|
||||
| 2026-07-28 | 源码公开到 `git.k1412.top/wuyang/llm-atlas` | 路线、进度、研究方法和内容变更均可追踪 |
|
||||
| 2026-07-28 | 站点使用不可变镜像与 Compose Manager 部署 | 每次发布保留明确版本、健康检查和回滚点 |
|
||||
| 2026-07-28 | 长上下文按计算、缓存、位置、状态容量、系统五张账单组织 | 避免把 FlashAttention、位置外推和“记住更久”混成同一个问题 |
|
||||
| 2026-07-28 | 交互缓存数字统一标记为教学估算 | 展示增长规律,不冒充任一模型的真实线上显存基准 |
|
||||
| 2026-07-28 | MoE 按六张账组织,路由算法与集群执行分开核算 | 避免用“稀疏所以便宜”跳过容量、负载、通信与权重读取 |
|
||||
| 2026-07-28 | “aux-loss-free”保留论文真实边界 | 区分 selection bias、mixture weight、z-loss 与 V3 的极小 sequence-wise loss |
|
||||
|
||||
## 未决问题
|
||||
|
||||
|
||||
@@ -17,8 +17,8 @@
|
||||
- 持续进度:[PROGRESS.md](./PROGRESS.md)
|
||||
- 证据与写作规范:[research/METHODOLOGY.md](./research/METHODOLOGY.md)
|
||||
|
||||
首个里程碑包含 16 专题学习地图、125 篇关键论文索引、Kimi K3 完整导读、
|
||||
Transformer 基础、DeepSeek 技术谱系、长上下文与高效注意力深度专题,以及四张原创交互可视化。
|
||||
当前里程碑包含 16 专题学习地图、130 篇关键论文索引、Kimi K3 完整导读、
|
||||
Transformer 基础、DeepSeek 技术谱系、长上下文与 MoE 深度专题,以及五张原创交互可视化。
|
||||
其余专题按进度账本持续扩建。
|
||||
|
||||
## 本地开发
|
||||
|
||||
@@ -45,6 +45,9 @@ GPT 系列 → Kaplan scaling laws → Chinchilla compute-optimal → 数据质
|
||||
|
||||
Conditional Computation → Sparsely-Gated MoE → GShard/Switch → DeepSeekMoE → LatentMoE → K3 Stable LatentMoE。重点解释路由、专家特化、负载均衡与通信。
|
||||
|
||||
首版已发布:以“容量、激活计算、路由、负载、通信、稳定性”六张账串起 1991–2026 的 19 篇一手论文,
|
||||
包含 DeepSeekMoE / Loss-Free / LatentMoE / K3 重点推导,以及 8 种架构、4 类平衡策略的交互实验室。
|
||||
|
||||
### 07. 长上下文与高效注意力
|
||||
|
||||
稀疏注意力、线性注意力、FlashAttention、MQA/GQA、MLA、状态空间模型、Delta Rule、Kimi Linear/KDA、混合注意力与 1M 上下文。
|
||||
|
||||
+3
-1
@@ -8,7 +8,9 @@
|
||||
"dev": "astro dev --host 0.0.0.0",
|
||||
"build": "astro build",
|
||||
"preview": "astro preview --host 0.0.0.0",
|
||||
"check": "astro check"
|
||||
"check": "astro check",
|
||||
"check:site": "node scripts/check-site.mjs",
|
||||
"check:moe-browser": "node scripts/check-moe-browser.mjs"
|
||||
},
|
||||
"dependencies": {
|
||||
"@astrojs/sitemap": "3.7.3",
|
||||
|
||||
@@ -0,0 +1,159 @@
|
||||
# 稀疏计算与 MoE 研究账本
|
||||
|
||||
最后核验:2026-07-28
|
||||
|
||||
## 教学主线
|
||||
|
||||
MoE 不能只写成“更多参数、较少计算”。本专题固定拆成六张账单:
|
||||
|
||||
1. **容量**:模型一共存了多少参数。
|
||||
2. **激活计算**:每个 Token 实际经过多少专家参数。
|
||||
3. **路由**:谁决定 Token 去哪里,选择是 top-1、top-2 还是 top-k。
|
||||
4. **负载**:专家收到的 Token 是否均匀,溢出时是否丢 Token。
|
||||
5. **通信**:专家分散在不同设备后,dispatch / combine 两次 All-to-All 搬多少数据。
|
||||
6. **稳定性**:硬路由、router logits、极稀疏潜空间链怎样影响训练。
|
||||
|
||||
核心因果链:
|
||||
|
||||
> 稠密 FFN 把容量与每 Token 算力绑死
|
||||
> → 软门控专家学会分工,但所有专家仍可能参与
|
||||
> → 稀疏 top-k 让条件计算进入大规模模型
|
||||
> → Transformer MoE 把 FFN 专家分到设备上
|
||||
> → 容量、掉 Token、负载和 All-to-All 成为新瓶颈
|
||||
> → DeepSeek 细分专家、隔离共享知识,并限制跨设备路由
|
||||
> → 无辅助损失 bias 把均衡移出主梯度
|
||||
> → LatentMoE 压缩路由载荷与专家宽度
|
||||
> → K3 在 896 选 16 的极稀疏区间进一步解决激活爆炸、bias 更新和执行不均。
|
||||
|
||||
## 本地一手材料
|
||||
|
||||
PDF 与文本只作本地研究缓存,受 `.gitignore` 排除;公开仓库仅提交本账本和 canonical URL。
|
||||
|
||||
| ID | 来源 | 本地页数 | 本轮用途 |
|
||||
|---|---|---:|---|
|
||||
| `1701.06538` | Sparsely-Gated MoE | 19 | Noisy Top-k 与现代稀疏门控起点 |
|
||||
| `2006.16668` | GShard | 35 | Transformer MoE、自动切分与 Top-2 |
|
||||
| `2101.03961` | Switch Transformers | 40 | Top-1、capacity factor、drop 与 EP |
|
||||
| `2202.08906` | ST-MoE | 38 | router z-loss、稳定性与迁移 |
|
||||
| `2202.09368` | Expert Choice | 14 | 专家选 Token 的负载均衡分支 |
|
||||
| `2401.04088` | Mixtral of Experts | 13 | 社区可运行的 8 选 2 MoE |
|
||||
| `2401.06066` | DeepSeekMoE | 33 | 细粒度专家与共享专家 |
|
||||
| `2408.15664` | Auxiliary-Loss-Free Balancing | 14 | expert bias 与干扰梯度 |
|
||||
| `2412.19437` | DeepSeek-V3 | 53 | 256 选 8、node limit、系统协同 |
|
||||
| `2601.18089` | LatentMoE | 18 | 潜空间专家、带宽与通信 |
|
||||
| `2607.24653` | Kimi K3 | 47 | Stable LatentMoE、QB 与 MoonEP |
|
||||
|
||||
补充复用:
|
||||
|
||||
- DeepSeek-V2 `2405.04434`:本地缓存位于 `research/sources/long-context/`。
|
||||
- Kimi K3 原报告:`research/sources/kimi-k3/k3_tech_report.*`。
|
||||
|
||||
## 已核验的关键结论
|
||||
|
||||
### Switch:capacity factor 不是模型容量
|
||||
|
||||
Switch 的专家容量定义为:
|
||||
|
||||
```text
|
||||
expert capacity = tokens per batch / number of experts × capacity factor
|
||||
```
|
||||
|
||||
- 它是每个专家这一批最多能处理多少 Token,不是专家参数量。
|
||||
- 大于 1 的 capacity factor 提供负载缓冲。
|
||||
- 专家溢出时,Switch 跳过专家计算并让表示沿残差路径进入下一层。
|
||||
- capacity factor 增大也会增加 padding、计算、激活内存和通信。
|
||||
- 论文报告在其主要实验中 dropped tokens 通常低于 1%,不可外推为所有 MoE。
|
||||
|
||||
### ST-MoE:稳定不等于均衡
|
||||
|
||||
- Load-balancing loss 管“专家是否被均匀使用”。
|
||||
- Router z-loss 管“进入 router softmax 的 logits 是否过大”。
|
||||
- ST-MoE 的公式惩罚每个 Token router log-partition 的平方。
|
||||
- 论文的稳定性研究使用训练 CF=1.25、评估 CF=2.0,并令 z-loss 系数 `0.001`。
|
||||
- 因而 z-loss 不能被写成另一种负载均衡损失。
|
||||
|
||||
### DeepSeekMoE:同计算预算下切得更细
|
||||
|
||||
- 传统配置有 `N` 个专家、激活 `K` 个。
|
||||
- DeepSeekMoE 把每个 FFN 专家沿中间维切成 `m` 个更小专家,总数成为 `mN`,激活数成为 `mK`,保持总专家参数与激活计算近似不变。
|
||||
- 论文示例:`N=16, K=2` 只有 `C(16,2)=120` 种组合;切成 `64` 个小专家并激活 `8` 个后,组合数为 `C(64,8)=4,426,165,368`。
|
||||
- Shared expert isolation 让共享专家始终执行,承载共性变换;路由专家更专注于差异知识。
|
||||
- 2B 验证模型为 1 个共享专家 + 63 个路由专家,其中每 Token 激活 1+7。
|
||||
- 禁用共享专家并多激活一个路由专家时,论文的 Pile loss 从 1.808 上升到 2.414;这是该实验设置下的证据,不是通用常数。
|
||||
|
||||
### Loss-Free / V3:selection 与 mixture weight 分开
|
||||
|
||||
- 原始 affinity score 为 `s`,expert bias 为 `b`。
|
||||
- `s+b` 只用于决定 top-k;最终混合专家输出的权重仍来自原始 `s`。
|
||||
- 因此 bias 调整 dispatch,不向语言建模参数引入 auxiliary-loss 的干扰梯度。
|
||||
- 原方法按上一步负载以固定步长更新 bias;步长过小反应慢、过大会振荡。
|
||||
- DeepSeek-V3:671B 总参数、37B 激活;每个 MoE 层 1 shared + 256 routed,激活 8 routed;每 Token 最多发往 4 个节点。
|
||||
- V3 主要使用 auxiliary-loss-free 策略,但仍保留系数极小的 sequence-wise balance loss 防止单序列极端失衡。正文必须保留这个边界,不能简写为“完全没有任何辅助损失”。
|
||||
- V3 报告称训练和推理都不 drop Token;这是其负载均衡与部署策略下的模型报告事实。
|
||||
|
||||
### LatentMoE:省下的是路由宽度与专家权重流量
|
||||
|
||||
- 标准 MoE 以模型宽度 `d` dispatch Token,并让 routed expert 在 `d` 宽度上计算。
|
||||
- LatentMoE 先下投影到 `ℓ<d`,在 latent space dispatch、计算和 combine,再上投影回 `d`。
|
||||
- 共享专家仍可保留全宽路径。
|
||||
- 原论文称 routed parameter load 与 All-to-All traffic 可按 `d/ℓ` 比例降低。
|
||||
- `ℓ-MoE_eff` 主要把节省换成更多总专家;`ℓ-MoE_acc` 同时扩大总专家数与 top-k,以近似固定推理成本换精度,论文推荐后者作为 Pareto 方案。
|
||||
- 论文在 16B/2B active 与 95B/8B active 模型上验证,并把压缩比 `α=d/ℓ=4` 作为主要后续设置;“4× 可行”仍是该实验范围内的结论。
|
||||
- 论文中万亿参数 Kimi-K2 serving 比较来自性能模拟器与 effective-parameter construction,不是真实生产 A/B 基准。
|
||||
|
||||
### K3:Stable LatentMoE 的“Stable”有明确内容
|
||||
|
||||
K3 报告给出的配置:
|
||||
|
||||
- 2.78T 总参数,104.2B 激活参数。
|
||||
- 896 个 routed experts,每 Token 激活 16 个。
|
||||
- 2 个 full-width shared experts。
|
||||
- hidden dimension 7168,latent MoE dimension 3584。
|
||||
- 每个 expert 的 MoE hidden dimension 3072。
|
||||
|
||||
从 LatentMoE 到 Stable LatentMoE 的三项增量:
|
||||
|
||||
1. routed expert 聚合后、上投影前加入 RMSNorm;
|
||||
2. 用有界的 SiTU-GLU 抑制连续矩阵乘链中的内部激活爆炸;
|
||||
3. 用 Quantile Balancing 替代固定步长 bias 更新。
|
||||
|
||||
Quantile Balancing:
|
||||
|
||||
- 每批 `m` 个 Token、`n` 个专家、top-k 时,目标负载为 `q=mk/n`。
|
||||
- 先算 Top-(k+1),第 `k+1` 个分数给出每个 Token 进入 top-k 的 cutoff。
|
||||
- 对每个专家计算其 score 相对各 Token cutoff 的 margin,并取 `(1-k/n)` quantile 得到下一步 bias。
|
||||
- 更新只在下一训练步生效,避免用当前批的未来 Token 改当前路由。
|
||||
- 大规模实现不聚集全部 margin,而用每专家直方图 + 一次 all-reduce 估计全局 quantile。
|
||||
- 最终 bias 在推理阶段冻结。
|
||||
|
||||
### MoonEP:算法均衡之后还有执行均衡
|
||||
|
||||
- QB 使专家层面的目标负载更均衡,但专家部署到 EP ranks 后仍要解决设备执行不均。
|
||||
- MoonEP 为当前 micro-batch 与层在线规划 redundant experts。
|
||||
- K3 报告证明每 rank 预留至多 `E/R` 个冗余专家槽即可保证存在平衡计划。
|
||||
- 每个 rank 最终收到完全相同的 `S×K` Token 槽,形成 static shapes,消除每层 host sync。
|
||||
- fused permute/unpermute 让 Token 直接进入远端 expert-grouped buffer,减少中间 copy。
|
||||
- 这是模型路由、通信、内存与 kernel 调度的系统协同,不能只归功于 router 算法。
|
||||
|
||||
## 关键边界
|
||||
|
||||
- 总参数与激活参数口径可能是否包含 embedding、attention、shared expert;跨模型比较必须沿用各报告口径。
|
||||
- “激活专家比例 K/N”不等于“激活参数比例”,因为 attention、共享专家、投影和 embedding 仍是稠密路径。
|
||||
- 理论 FLOPs 不直接等于端到端延迟;小 batch 常受权重带宽限制,大 EP 常受 All-to-All 限制。
|
||||
- `C(N,K)` 只展示组合空间上限,不证明模型真的学到同等数量的有意义专业化模式。
|
||||
- Expert Choice 可实现专家侧固定容量,但对 causal LM 可能泄露同一 chunk 中未来 Token 对前面 Token 路由的影响;Loss-Free 论文给出明确讨论。
|
||||
- K3 的 104B activated parameters 不应被误写成 16/896×2.8T。
|
||||
- Grok CLI 只用于扩大候选论文与检查遗漏;以上机制和数字均回查原论文或 K3 官方报告。
|
||||
|
||||
## 首版正文完成闸门
|
||||
|
||||
- [x] 用一张图区分 total parameters、active parameters、FLOPs、weight traffic。
|
||||
- [x] 画出 route → dispatch → expert → combine 的两次 All-to-All。
|
||||
- [x] 让读者能手算 capacity factor 与 overflow。
|
||||
- [x] 把 aux loss、router z-loss、loss-free bias、Quantile Balancing 放在同一因果链。
|
||||
- [x] 单独重绘 DeepSeekMoE 细粒度与 shared expert。
|
||||
- [x] 单独推导 LatentMoE 的 `d→ℓ→d`。
|
||||
- [x] 交互实验室可切换 Switch、Mixtral、DeepSeekMoE/V3、LatentMoE 与 K3。
|
||||
- [x] 讲清 K3 的 RMSNorm、SiTU-GLU、QB 和 MoonEP,不把 Stable LatentMoE 缩成一行配置。
|
||||
- [x] 论文链全部指向一手来源。
|
||||
- [x] 类型、链接、交互、桌面/移动端和容器检查通过。
|
||||
@@ -0,0 +1,187 @@
|
||||
import { writeFileSync } from "node:fs";
|
||||
|
||||
const cdpPort = process.env.CDP_PORT ?? "9223";
|
||||
const baseUrl = process.env.SITE_URL ?? "http://127.0.0.1:4322";
|
||||
const pages = await fetch(`http://127.0.0.1:${cdpPort}/json/list`).then((response) => response.json());
|
||||
const page = pages.find((entry) => entry.type === "page");
|
||||
if (!page) throw new Error(`CDP ${cdpPort} 没有可用页面`);
|
||||
|
||||
const socket = new WebSocket(page.webSocketDebuggerUrl);
|
||||
await new Promise((resolve, reject) => {
|
||||
socket.addEventListener("open", resolve, { once: true });
|
||||
socket.addEventListener("error", reject, { once: true });
|
||||
});
|
||||
|
||||
let nextId = 0;
|
||||
const pending = new Map();
|
||||
const exceptions = [];
|
||||
socket.addEventListener("message", (event) => {
|
||||
const message = JSON.parse(event.data);
|
||||
if (message.id && pending.has(message.id)) {
|
||||
const { resolve, reject } = pending.get(message.id);
|
||||
pending.delete(message.id);
|
||||
if (message.error) reject(new Error(message.error.message));
|
||||
else resolve(message.result);
|
||||
}
|
||||
if (message.method === "Runtime.exceptionThrown") {
|
||||
exceptions.push(message.params.exceptionDetails.text);
|
||||
}
|
||||
});
|
||||
|
||||
const command = (method, params = {}) => new Promise((resolve, reject) => {
|
||||
const id = ++nextId;
|
||||
pending.set(id, { resolve, reject });
|
||||
socket.send(JSON.stringify({ id, method, params }));
|
||||
});
|
||||
const pause = (milliseconds) => new Promise((resolve) => setTimeout(resolve, milliseconds));
|
||||
const evaluate = async (expression) => {
|
||||
const result = await command("Runtime.evaluate", {
|
||||
expression,
|
||||
returnByValue: true,
|
||||
awaitPromise: true,
|
||||
});
|
||||
if (result.exceptionDetails) throw new Error(result.exceptionDetails.text);
|
||||
return result.result.value;
|
||||
};
|
||||
const navigate = async (path) => {
|
||||
await command("Page.navigate", { url: `${baseUrl}${path}` });
|
||||
for (let attempt = 0; attempt < 30; attempt += 1) {
|
||||
await pause(100);
|
||||
if (await evaluate("document.readyState === 'complete'")) return;
|
||||
}
|
||||
throw new Error(`${path} 加载超时`);
|
||||
};
|
||||
const screenshot = async (path) => {
|
||||
const result = await command("Page.captureScreenshot", { format: "png", captureBeyondViewport: false });
|
||||
writeFileSync(path, Buffer.from(result.data, "base64"));
|
||||
};
|
||||
|
||||
await command("Page.enable");
|
||||
await command("Runtime.enable");
|
||||
await command("Emulation.setDeviceMetricsOverride", {
|
||||
width: 1440,
|
||||
height: 1100,
|
||||
deviceScaleFactor: 1,
|
||||
mobile: false,
|
||||
});
|
||||
await navigate("/moe/");
|
||||
|
||||
const k3 = await evaluate(`(() => {
|
||||
const root = document.querySelector("[data-moe-lab]");
|
||||
const started = performance.now();
|
||||
root.querySelector('[data-preset-button="k3"]').click();
|
||||
const duration = performance.now() - started;
|
||||
const skew = root.querySelector("[data-skew]");
|
||||
const capacity = root.querySelector("[data-capacity]");
|
||||
skew.value = "100";
|
||||
capacity.value = "0.75";
|
||||
skew.dispatchEvent(new Event("input", { bubbles: true }));
|
||||
capacity.dispatchEvent(new Event("input", { bubbles: true }));
|
||||
root.querySelector('[data-balance="none"]').click();
|
||||
const none = {
|
||||
ratio: root.querySelector("[data-load-ratio]").textContent,
|
||||
overflow: root.querySelector("[data-drop-rate]").textContent,
|
||||
};
|
||||
root.querySelector('[data-balance="quantile"]').click();
|
||||
return {
|
||||
duration: Number(duration.toFixed(2)),
|
||||
title: root.querySelector("[data-route-title]").textContent,
|
||||
fraction: root.querySelector("[data-active-fraction]").textContent,
|
||||
communication: root.querySelector("[data-comm-value]").textContent,
|
||||
communicationNote: root.querySelector("[data-comm-note]").textContent,
|
||||
combination: root.querySelector("[data-combination-value]").textContent,
|
||||
none,
|
||||
quantile: {
|
||||
ratio: root.querySelector("[data-load-ratio]").textContent,
|
||||
overflow: root.querySelector("[data-drop-rate]").textContent,
|
||||
},
|
||||
metrics: root.querySelectorAll(".metric-grid > article").length,
|
||||
};
|
||||
})()`);
|
||||
|
||||
const latent = await evaluate(`(() => {
|
||||
const root = document.querySelector("[data-moe-lab]");
|
||||
const started = performance.now();
|
||||
root.querySelector('[data-preset-button="latent"]').click();
|
||||
return {
|
||||
duration: Number((performance.now() - started).toFixed(2)),
|
||||
fraction: root.querySelector("[data-active-fraction]").textContent,
|
||||
communication: root.querySelector("[data-comm-value]").textContent,
|
||||
note: root.querySelector("[data-comm-note]").textContent,
|
||||
};
|
||||
})()`);
|
||||
|
||||
const layout = await evaluate(`(() => {
|
||||
const header = document.querySelector(".site-header");
|
||||
const nav = document.querySelector(".top-nav");
|
||||
const meta = document.querySelector(".header-meta");
|
||||
return {
|
||||
documentOverflow: document.documentElement.scrollWidth - document.documentElement.clientWidth,
|
||||
navGap: Number((meta.getBoundingClientRect().left - nav.getBoundingClientRect().right).toFixed(1)),
|
||||
articleSections: document.querySelectorAll(".article-section").length,
|
||||
};
|
||||
})()`);
|
||||
|
||||
await evaluate(`(() => {
|
||||
document.documentElement.style.scrollBehavior = "auto";
|
||||
document.querySelector("[data-moe-lab]").scrollIntoView({ block: "start", behavior: "instant" });
|
||||
return { scrollY, top: document.querySelector("[data-moe-lab]").getBoundingClientRect().top };
|
||||
})()`);
|
||||
await pause(200);
|
||||
await screenshot("/tmp/llm-atlas-moe-lab-desktop.png");
|
||||
|
||||
await command("Emulation.setDeviceMetricsOverride", {
|
||||
width: 390,
|
||||
height: 844,
|
||||
deviceScaleFactor: 1,
|
||||
mobile: true,
|
||||
});
|
||||
await navigate("/moe/");
|
||||
const mobile = await evaluate(`({
|
||||
documentOverflow: document.documentElement.scrollWidth - document.documentElement.clientWidth,
|
||||
menuVisible: getComputedStyle(document.querySelector("#menu-toggle")).display !== "none",
|
||||
title: document.querySelector("h1").innerText,
|
||||
})`);
|
||||
await screenshot("/tmp/llm-atlas-moe-mobile.png");
|
||||
|
||||
await command("Emulation.setDeviceMetricsOverride", {
|
||||
width: 1440,
|
||||
height: 1100,
|
||||
deviceScaleFactor: 1,
|
||||
mobile: false,
|
||||
});
|
||||
await navigate("/");
|
||||
const home = await evaluate(`({
|
||||
documentOverflow: document.documentElement.scrollWidth - document.documentElement.clientWidth,
|
||||
releaseCards: document.querySelectorAll(".release-card").length,
|
||||
navLinks: document.querySelectorAll(".top-nav a").length,
|
||||
})`);
|
||||
await screenshot("/tmp/llm-atlas-home-desktop.png");
|
||||
await evaluate(`(() => {
|
||||
document.documentElement.style.scrollBehavior = "auto";
|
||||
document.querySelector("#new-chapters").scrollIntoView({ block: "start", behavior: "instant" });
|
||||
})()`);
|
||||
await pause(100);
|
||||
await screenshot("/tmp/llm-atlas-home-releases.png");
|
||||
|
||||
const report = { k3, latent, layout, mobile, home, exceptions };
|
||||
console.log(JSON.stringify(report, null, 2));
|
||||
|
||||
const failures = [];
|
||||
if (!k3.title.includes("896 选 16")) failures.push("K3 预设未生效");
|
||||
if (k3.metrics !== 6) failures.push(`指标卡数量异常:${k3.metrics}`);
|
||||
if (!k3.communicationNote.includes("2× latent")) failures.push("K3 latent 通信说明缺失");
|
||||
if (!latent.note.includes("4× latent")) failures.push("LatentMoE 4× 通信说明缺失");
|
||||
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) {
|
||||
failures.push("页面存在横向溢出");
|
||||
}
|
||||
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
||||
if (!mobile.menuVisible) failures.push("移动端菜单按钮未显示");
|
||||
if (home.releaseCards !== 2) failures.push(`首页新章卡数量异常:${home.releaseCards}`);
|
||||
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
|
||||
|
||||
socket.close();
|
||||
if (failures.length) {
|
||||
failures.forEach((failure) => console.error(`- ${failure}`));
|
||||
process.exit(1);
|
||||
}
|
||||
@@ -0,0 +1,67 @@
|
||||
import { existsSync, readdirSync, readFileSync, statSync } from "node:fs";
|
||||
import { extname, join, normalize } from "node:path";
|
||||
|
||||
const root = new URL("../dist/", import.meta.url);
|
||||
const rootPath = root.pathname;
|
||||
|
||||
function walk(directory) {
|
||||
return readdirSync(directory).flatMap((name) => {
|
||||
const path = join(directory, name);
|
||||
return statSync(path).isDirectory() ? walk(path) : [path];
|
||||
});
|
||||
}
|
||||
|
||||
function routeFile(pathname) {
|
||||
const clean = decodeURIComponent(pathname).replace(/^\/+/, "");
|
||||
if (!clean) return join(rootPath, "index.html");
|
||||
const local = normalize(join(rootPath, clean));
|
||||
if (!local.startsWith(normalize(rootPath))) return null;
|
||||
if (extname(local)) return local;
|
||||
return join(local, "index.html");
|
||||
}
|
||||
|
||||
if (!existsSync(rootPath)) {
|
||||
console.error("dist/ 不存在;请先运行 npm run build。");
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
const htmlFiles = walk(rootPath).filter((file) => file.endsWith(".html"));
|
||||
const anchors = new Map();
|
||||
for (const file of htmlFiles) {
|
||||
const html = readFileSync(file, "utf8");
|
||||
anchors.set(file, new Set([...html.matchAll(/\sid=["']([^"']+)["']/g)].map((match) => match[1])));
|
||||
}
|
||||
|
||||
let references = 0;
|
||||
let anchorReferences = 0;
|
||||
const failures = [];
|
||||
|
||||
for (const source of htmlFiles) {
|
||||
const html = readFileSync(source, "utf8");
|
||||
const hrefs = [...html.matchAll(/\shref=["']([^"']+)["']/g)].map((match) => match[1]);
|
||||
for (const href of hrefs) {
|
||||
if (!href.startsWith("/") || href.startsWith("//")) continue;
|
||||
references += 1;
|
||||
const url = new URL(href, "https://llm-atlas.local");
|
||||
const target = routeFile(url.pathname);
|
||||
if (!target || !existsSync(target)) {
|
||||
failures.push(`${source.replace(rootPath, "/")} → ${href}(目标不存在)`);
|
||||
continue;
|
||||
}
|
||||
if (url.hash) {
|
||||
anchorReferences += 1;
|
||||
const id = decodeURIComponent(url.hash.slice(1));
|
||||
if (!anchors.get(target)?.has(id)) {
|
||||
failures.push(`${source.replace(rootPath, "/")} → ${href}(锚点不存在)`);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
console.log(
|
||||
`${htmlFiles.length} 个页面,${references} 个站内引用,${anchorReferences} 个跨页锚点,${failures.length} 个失败。`,
|
||||
);
|
||||
if (failures.length) {
|
||||
failures.forEach((failure) => console.error(`- ${failure}`));
|
||||
process.exit(1);
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -5,6 +5,7 @@
|
||||
</div>
|
||||
<div class="footer-links">
|
||||
<a href="/roadmap/">学习地图</a>
|
||||
<a href="/moe/">MoE 专题</a>
|
||||
<a href="/long-context/">长上下文专题</a>
|
||||
<a href="/progress/">研究进度</a>
|
||||
<a href="https://git.k1412.top/wuyang/llm-atlas" rel="noreferrer">开放源码</a>
|
||||
|
||||
@@ -10,6 +10,7 @@ const items = [
|
||||
{ id: "k3", href: "/k3/", label: "K3 解剖" },
|
||||
{ id: "deepseek", href: "/deepseek/", label: "DeepSeek" },
|
||||
{ id: "foundations", href: "/foundations/", label: "基础原理" },
|
||||
{ id: "moe", href: "/moe/", label: "MoE" },
|
||||
{ id: "long-context", href: "/long-context/", label: "长上下文" },
|
||||
{ id: "papers", href: "/papers/", label: "论文库" },
|
||||
{ id: "progress", href: "/progress/", label: "进度" },
|
||||
|
||||
@@ -95,14 +95,14 @@ export const chapters: Chapter[] = [
|
||||
},
|
||||
{
|
||||
number: "06",
|
||||
slug: "architecture/moe",
|
||||
slug: "moe",
|
||||
title: "稀疏计算与 MoE",
|
||||
kicker: "SPARSE EXPERTS",
|
||||
question: "怎样让模型装下更多知识,却不让每个 Token 都付全部算力?",
|
||||
summary: "从条件计算到 DeepSeekMoE、LatentMoE 与 K3 Stable LatentMoE,解释路由、特化和负载均衡。",
|
||||
status: "researching",
|
||||
progress: 28,
|
||||
papers: 17,
|
||||
status: "published",
|
||||
progress: 74,
|
||||
papers: 19,
|
||||
prerequisites: ["02", "04"],
|
||||
highlights: ["专家路由", "细粒度专家", "896 选 16"],
|
||||
},
|
||||
|
||||
@@ -71,6 +71,14 @@ export const papers: Paper[] = [
|
||||
contribution: "把 BPE 引入神经机器翻译,形成现代子词分词主线。",
|
||||
verified: true,
|
||||
},
|
||||
{
|
||||
year: 1991,
|
||||
title: "Adaptive Mixtures of Local Experts",
|
||||
url: "https://doi.org/10.1162/neco.1991.3.1.79",
|
||||
topics: ["MoE"],
|
||||
contribution: "以门控网络让多个局部专家竞争分工,是现代 MoE 的概念起点。",
|
||||
verified: true,
|
||||
},
|
||||
{
|
||||
year: 2017,
|
||||
title: "Attention Is All You Need",
|
||||
@@ -401,6 +409,22 @@ export const papers: Paper[] = [
|
||||
contribution: "系统化稀疏专家的稳定训练与迁移配方。",
|
||||
verified: true,
|
||||
},
|
||||
{
|
||||
year: 2022,
|
||||
title: "Mixture-of-Experts with Expert Choice Routing",
|
||||
url: "https://arxiv.org/abs/2202.09368",
|
||||
topics: ["MoE"],
|
||||
contribution: "把 token 选专家反转为专家选固定容量的 token,直接控制专家负载。",
|
||||
verified: true,
|
||||
},
|
||||
{
|
||||
year: 2023,
|
||||
title: "From Sparse to Soft Mixtures of Experts",
|
||||
url: "https://arxiv.org/abs/2308.00951",
|
||||
topics: ["MoE"],
|
||||
contribution: "用连续 slot 分配替代硬 top-k,探索完全可微的专家混合。",
|
||||
verified: true,
|
||||
},
|
||||
{
|
||||
year: 2024,
|
||||
title: "Mixtral of Experts",
|
||||
@@ -427,6 +451,23 @@ export const papers: Paper[] = [
|
||||
spotlight: "DeepSeek",
|
||||
verified: true,
|
||||
},
|
||||
{
|
||||
year: 2024,
|
||||
title: "Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language Models",
|
||||
url: "https://arxiv.org/abs/2404.02258",
|
||||
topics: ["MoE", "Transformer"],
|
||||
contribution: "把条件计算从专家宽度扩到网络深度,让不同 Token 动态跳过或进入层。",
|
||||
verified: true,
|
||||
},
|
||||
{
|
||||
year: 2024,
|
||||
title: "Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts",
|
||||
url: "https://arxiv.org/abs/2408.15664",
|
||||
topics: ["MoE"],
|
||||
contribution: "以只影响 top-k 选择的 expert bias 调负载,避免均衡辅助损失干扰主梯度。",
|
||||
spotlight: "DeepSeek",
|
||||
verified: true,
|
||||
},
|
||||
{
|
||||
year: 2024,
|
||||
title: "DeepSeek-V3 Technical Report",
|
||||
|
||||
@@ -140,6 +140,7 @@ const toc = [
|
||||
因此模型侧路由与系统侧 Expert Parallel 从来不能分开看。
|
||||
</p>
|
||||
</div>
|
||||
<a class="button primary" href="/moe/#deepseek">进入 MoE 专题:比较粗专家、细粒度专家与共享专家 →</a>
|
||||
</section>
|
||||
|
||||
<section class="article-section" id="mla">
|
||||
@@ -212,6 +213,10 @@ const toc = [
|
||||
根据近期负载上调冷门专家、下调热门专家;bias 只影响路由选择,不直接进入最终门控权重,
|
||||
从而把“系统要均衡”和“模型要准确”更松地解耦。
|
||||
</p>
|
||||
<p>
|
||||
<a href="/moe/#balance">专题版推导</a>会进一步区分 z-loss、辅助平衡损失、Loss-Free expert bias
|
||||
与 K3 Quantile Balancing;“aux-loss-free”并不等于系统里不存在任何平衡约束。
|
||||
</p>
|
||||
|
||||
<h3>FP8 训练真正难在哪</h3>
|
||||
<p>
|
||||
|
||||
+57
-25
@@ -7,6 +7,7 @@ import { chapters, statusLabel } from "@/data/chapters";
|
||||
const routes: Record<string, string> = {
|
||||
roadmap: "/roadmap/",
|
||||
foundations: "/foundations/",
|
||||
moe: "/moe/",
|
||||
"long-context": "/long-context/",
|
||||
};
|
||||
|
||||
@@ -75,6 +76,7 @@ const paths = [
|
||||
<a class="button primary" href="/roadmap/">选择学习路径 <span aria-hidden="true">↓</span></a>
|
||||
<a class="button" href="/k3/">直接解剖 K3</a>
|
||||
<a class="button" href="/deepseek/">DeepSeek 专题</a>
|
||||
<a class="button" href="/moe/">MoE 专题</a>
|
||||
<a class="button" href="/long-context/">长上下文专题</a>
|
||||
</div>
|
||||
</div>
|
||||
@@ -85,7 +87,7 @@ const paths = [
|
||||
<div class="hero-stats">
|
||||
<div><b>16</b><span>核心专题</span></div>
|
||||
<div><b>151</b><span>K3 报告来源</span></div>
|
||||
<div><b>125</b><span>关键论文索引</span></div>
|
||||
<div><b>130</b><span>关键论文索引</span></div>
|
||||
<div><b>47p</b><span>K3 技术报告</span></div>
|
||||
</div>
|
||||
</aside>
|
||||
@@ -97,23 +99,41 @@ const paths = [
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section class="section compact release-section" id="new-long-context">
|
||||
<a class="release-card" href="/long-context/">
|
||||
<div>
|
||||
<p class="eyebrow"><span>NEW / CHAPTER 07</span> LONG CONTEXT</p>
|
||||
<h2>一百万 Token,不是一扇更大的窗</h2>
|
||||
<p>
|
||||
新专题把长上下文拆成计算、缓存、位置、状态容量与系统五张账单,
|
||||
沿 26 篇一手论文走完 FlashAttention、MLA、Delta Rule、KDA、Kimi K3 与 DeepSeek-V4。
|
||||
</p>
|
||||
</div>
|
||||
<dl>
|
||||
<div><dt>LINEAGE</dt><dd>2019 → 2026</dd></div>
|
||||
<div><dt>VISUALS</dt><dd>10+ 机制图</dd></div>
|
||||
<div><dt>LAB</dt><dd>8 种策略 · 4 档长度</dd></div>
|
||||
</dl>
|
||||
<span class="release-arrow" aria-hidden="true">进入专题 →</span>
|
||||
</a>
|
||||
<section class="section compact release-section" id="new-chapters">
|
||||
<div class="release-grid">
|
||||
<a class="release-card moe-release" href="/moe/">
|
||||
<div>
|
||||
<p class="eyebrow"><span>NEW / CHAPTER 06</span> SPARSE EXPERTS</p>
|
||||
<h2>2.8T 参数,不等于每个 Token 跑 2.8T</h2>
|
||||
<p>
|
||||
把 MoE 拆成容量、激活计算、路由、负载、通信与稳定性六张账,
|
||||
从 Switch、DeepSeekMoE、Loss-Free、LatentMoE 一路走到 K3 Stable LatentMoE。
|
||||
</p>
|
||||
</div>
|
||||
<dl>
|
||||
<div><dt>LINEAGE</dt><dd>1991 → 2026</dd></div>
|
||||
<div><dt>PAPERS</dt><dd>19 篇一手来源</dd></div>
|
||||
<div><dt>LAB</dt><dd>8 架构 · 4 平衡策略</dd></div>
|
||||
</dl>
|
||||
<span class="release-arrow" aria-hidden="true">进入 MoE 专题 →</span>
|
||||
</a>
|
||||
<a class="release-card context-release" href="/long-context/">
|
||||
<div>
|
||||
<p class="eyebrow"><span>CHAPTER 07</span> LONG CONTEXT</p>
|
||||
<h2>一百万 Token,不是一扇更大的窗</h2>
|
||||
<p>
|
||||
把长上下文拆成计算、缓存、位置、状态容量与系统五张账,
|
||||
沿 26 篇一手论文走完 FlashAttention、MLA、Delta Rule、KDA、Kimi K3 与 DeepSeek-V4。
|
||||
</p>
|
||||
</div>
|
||||
<dl>
|
||||
<div><dt>LINEAGE</dt><dd>2019 → 2026</dd></div>
|
||||
<div><dt>VISUALS</dt><dd>10+ 机制图</dd></div>
|
||||
<div><dt>LAB</dt><dd>8 种策略 · 4 档长度</dd></div>
|
||||
</dl>
|
||||
<span class="release-arrow" aria-hidden="true">进入长上下文 →</span>
|
||||
</a>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section class="section" id="why">
|
||||
@@ -323,12 +343,20 @@ const paths = [
|
||||
padding-bottom: 0;
|
||||
}
|
||||
|
||||
.release-grid {
|
||||
display: grid;
|
||||
grid-template-columns: repeat(2, minmax(0, 1fr));
|
||||
gap: 22px;
|
||||
}
|
||||
|
||||
.release-card {
|
||||
position: relative;
|
||||
display: grid;
|
||||
grid-template-columns: minmax(0, 1.45fr) minmax(280px, 0.65fr);
|
||||
gap: clamp(34px, 6vw, 90px);
|
||||
padding: clamp(28px, 4vw, 58px);
|
||||
grid-template-rows: 1fr auto;
|
||||
gap: 28px;
|
||||
min-height: 560px;
|
||||
padding: clamp(28px, 3.6vw, 50px);
|
||||
padding-bottom: clamp(74px, 7vw, 92px);
|
||||
border: 1px solid var(--line-strong);
|
||||
color: inherit;
|
||||
text-decoration: none;
|
||||
@@ -346,8 +374,8 @@ const paths = [
|
||||
.release-card h2 {
|
||||
max-width: 760px;
|
||||
margin: 20px 0 18px;
|
||||
font-size: clamp(2rem, 4vw, 4.2rem);
|
||||
line-height: 1.03;
|
||||
font-size: clamp(2rem, 3.2vw, 3.5rem);
|
||||
line-height: 1.05;
|
||||
}
|
||||
|
||||
.release-card p:not(.eyebrow) {
|
||||
@@ -358,7 +386,7 @@ const paths = [
|
||||
}
|
||||
|
||||
.release-card dl {
|
||||
margin: 0 0 28px;
|
||||
margin: 0;
|
||||
border-top: 1px solid var(--line);
|
||||
}
|
||||
|
||||
@@ -391,8 +419,12 @@ const paths = [
|
||||
}
|
||||
|
||||
@media (max-width: 760px) {
|
||||
.release-card {
|
||||
.release-grid {
|
||||
grid-template-columns: 1fr;
|
||||
}
|
||||
|
||||
.release-card {
|
||||
min-height: 0;
|
||||
padding-bottom: 76px;
|
||||
}
|
||||
|
||||
|
||||
@@ -286,6 +286,7 @@ const toc = [
|
||||
这条技术线与 DeepSeekMoE 紧密相连:shared experts 保存公共知识,细粒度 routed experts 促进专业化。
|
||||
K3 再借 LatentMoE 让“选 16 个专家”的通信和权重读取可承受,并为极端稀疏补上稳定性机制。
|
||||
</p>
|
||||
<a class="button primary" href="/moe/#k3">进入 MoE 专题:逐步推导 LatentMoE 与 Quantile Balancing →</a>
|
||||
</section>
|
||||
|
||||
<section class="article-section" id="vision">
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -7,11 +7,12 @@ const published = chapters.filter((chapter) => chapter.status === "published").l
|
||||
const researching = chapters.filter((chapter) => ["researching", "drafting"].includes(chapter.status)).length;
|
||||
|
||||
const workstreams = [
|
||||
{ label: "研究框架与规范", value: 72, next: "给 125 篇索引补充逐篇精读层级" },
|
||||
{ label: "研究框架与规范", value: 76, next: "给 130 篇索引补充逐篇精读层级" },
|
||||
{ label: "网站设计系统", value: 89, next: "打印样式与更多通用可视化组件" },
|
||||
{ label: "Kimi K3 深读", value: 55, next: "扩写 scaling / infra 逐图笔记" },
|
||||
{ label: "Transformer 基础", value: 52, next: "加入矩阵形状动画与手算练习" },
|
||||
{ label: "DeepSeek 专题", value: 61, next: "GRPO 完整公式与训练轨迹推导" },
|
||||
{ label: "稀疏计算与 MoE", value: 74, next: "补充真实集群 traces 与专家特化案例" },
|
||||
{ label: "长上下文专题", value: 72, next: "加入更多论文逐图笔记与真实模型配置对比" },
|
||||
{ label: "引用与事实检查", value: 54, next: "自动化外链复查与来源等级扩展" },
|
||||
{ label: "开源与部署", value: 100, next: "每轮保留不可变镜像、提交与回滚点" },
|
||||
@@ -37,7 +38,7 @@ const workstreams = [
|
||||
<div><dt>OVERALL</dt><dd>专题平均 {average}%</dd></div>
|
||||
<div><dt>READABLE</dt><dd>{published} 个首版可读专题</dd></div>
|
||||
<div><dt>ACTIVE</dt><dd>{researching} 个研究/写作中</dd></div>
|
||||
<div><dt>UPDATED</dt><dd>2026-07-28 22:47 CST</dd></div>
|
||||
<div><dt>UPDATED</dt><dd>2026-07-28 23:24 CST</dd></div>
|
||||
<div><dt>MODE</dt><dd>持续迭代,不锁死版本</dd></div>
|
||||
</dl>
|
||||
</div>
|
||||
@@ -47,7 +48,7 @@ const workstreams = [
|
||||
<div class="section-heading">
|
||||
<div>
|
||||
<p class="eyebrow"><span>01</span> WORKSTREAMS</p>
|
||||
<h2>八条工作流同时推进,但不混淆“有页面”和“已核验”</h2>
|
||||
<h2>九条工作流同时推进,但不混淆“有页面”和“已核验”</h2>
|
||||
</div>
|
||||
<p class="section-lead">
|
||||
内容首版优先打通全局脉络;随后每轮迭代选择一个专题推进到论文/工程层,并做独立事实复核。
|
||||
@@ -84,10 +85,11 @@ const workstreams = [
|
||||
<article><span>✓</span><h3>K3 报告已结构化拆解</h3><p>47 页报告目录、151 条参考来源和架构/后训练/系统主线已经提取。</p></article>
|
||||
<article><span>✓</span><h3>16 专题知识图</h3><p>从语言模型基础到评测安全,包含先修依赖和三条贯穿案例。</p></article>
|
||||
<article><span>✓</span><h3>编辑式网站系统</h3><p>响应式导航、章节模板、侧栏、进度、论文链和证据提示组件。</p></article>
|
||||
<article><span>✓</span><h3>四张原创交互图</h3><p>K3 三轴架构、Self-Attention Query、DeepSeek 谱系与长上下文成本实验室。</p></article>
|
||||
<article><span>✓</span><h3>四篇首版长文</h3><p>K3 完整导读、Transformer 基础、DeepSeek 论文谱系与长上下文专题。</p></article>
|
||||
<article><span>✓</span><h3>五张原创交互图</h3><p>K3 三轴架构、Self-Attention Query、DeepSeek 谱系、长上下文成本与 MoE 路由实验室。</p></article>
|
||||
<article><span>✓</span><h3>五篇首版长文</h3><p>K3 导读、Transformer 基础、DeepSeek 谱系、长上下文与 MoE 专题。</p></article>
|
||||
<article><span>✓</span><h3>长上下文深度专题</h3><p>五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。</p></article>
|
||||
<article><span>✓</span><h3>125 篇关键论文索引</h3><p>覆盖 12 个专题,支持全文搜索、标签筛选与 Kimi/DeepSeek 聚光主线。</p></article>
|
||||
<article><span>✓</span><h3>MoE 深度专题</h3><p>六张账、19 篇一手论文、DeepSeek/K3 主线与路由—容量—通信交互实验室。</p></article>
|
||||
<article><span>✓</span><h3>130 篇关键论文索引</h3><p>覆盖 12 个专题,支持全文搜索、标签筛选与 Kimi/DeepSeek 聚光主线。</p></article>
|
||||
<article><span>✓</span><h3>公开仓库与自托管发布</h3><p>源码公开到 git.k1412.top,网站由不可变镜像、Compose Manager 与 HTTPS 交付。</p></article>
|
||||
</div>
|
||||
</section>
|
||||
@@ -102,9 +104,9 @@ const workstreams = [
|
||||
</div>
|
||||
<div class="queue-table">
|
||||
<div class="head"><b>优先级</b><b>专题</b><b>本轮交付</b><b>完成闸门</b></div>
|
||||
<div><span>P0</span><strong>稀疏计算与 MoE</strong><p>Switch → DeepSeekMoE → LatentMoE → Stable LatentMoE</p><em>路由模拟器 + 通信账本</em></div>
|
||||
<div><span>P0</span><strong>推理模型与测试时扩展</strong><p>CoT → verifier → GRPO → R1 → k1.5 → K3 MOPD</p><em>奖励/预算交互图</em></div>
|
||||
<div><span>P1</span><strong>长上下文二轮深化</strong><p>真实模型配置 → 内核细节 → 长上下文评测与失败案例</p><em>配置比较器 + 逐图论文笔记</em></div>
|
||||
<div><span>P1</span><strong>推理模型与测试时扩展</strong><p>CoT → verifier → GRPO → R1 → k1.5 → K3 MOPD</p><em>奖励/预算交互图</em></div>
|
||||
<div><span>P1</span><strong>MoE 二轮深化</strong><p>真实负载 traces → 专家特化可解释性 → 共享专家语义</p><em>案例库 + 集群证据</em></div>
|
||||
<div><span>P1</span><strong>大规模训练系统</strong><p>ZeRO/Megatron → Expert/Context Parallel → DualPipe/MoonEP</p><em>显存与通信计算器</em></div>
|
||||
<div><span>P2</span><strong>原生多模态</strong><p>ViT/CLIP → connector VLM → Kimi-VL/MoonViT-V2</p><em>视觉 Token 流程图</em></div>
|
||||
</div>
|
||||
|
||||
Reference in New Issue
Block a user