feat: launch LLM Atlas research course

This commit is contained in:
wuyang
2026-07-28 21:55:19 +08:00
commit 7d5c5e4c54
33 changed files with 13285 additions and 0 deletions
+535
View File
@@ -0,0 +1,535 @@
# LLM Atlas — Grok 辅助检索路线(K3-Anchored
> 生成来源:本机 Grok CLI Headless 模式,2026-07-28。
> 用途:扩展论文线索与检查遗漏,不直接作为正文证据。主代理需回到每篇一手来源核验后,才能把结论写入网站。
**Site:** Chinese educational site *LLM Atlas*
**Anchor primary source:** Kimi Team, *Kimi K3: Open Frontier Intelligence*, arXiv:[2607.24653](https://arxiv.org/abs/2607.24653) (2026-07) · DOI:[10.48550/arXiv.2607.24653](https://doi.org/10.48550/arXiv.2607.24653)
**Companion official sources:** [Kimi Linear](https://arxiv.org/abs/2510.26692) · [Attention Residuals](https://arxiv.org/abs/2603.15031) · [Kimi K2](https://arxiv.org/abs/2507.20534) · [Kimi K2.5](https://arxiv.org/abs/2602.02276) · [Kimi-VL / MoonViT](https://arxiv.org/abs/2504.07491) · [Muon is Scalable](https://arxiv.org/abs/2502.16982)
**Scope rules**
- Landmark *primary* papers only; year | title | canonical arXiv/DOI | problem solved | bridge to next idea.
- ≥80 papers; DeepSeek lineage given dense coverage; post-V3.2 only if **official** DeepSeek release.
- Uncertain / underspecified claims marked **`[UNCERTAIN]`**.
- No chapter prose — notes + tables only.
**Paper count (this doc):** ~95 listed entries (some multi-cited across modules).
---
## 0. K3 Spec Snapshot (from official report)
| Spec | Value (official abstract / Table 1) |
|---|---|
| Total / activated params | 2.8T MoE / **104B** activated |
| Context | **1M** tokens |
| Attention pattern | Hybrid **3 KDA : 1 Gated MLA** per block; final layer Gated MLA; **69 KDA + 24 MLA** |
| Depth connectivity | **Block Attention Residuals** (blocks of 12 layers) |
| FFN | **Stable LatentMoE**: **16 of 896** routed experts (+ shared experts) |
| Vision | **MoonViT-V2** (~401M), trained **from scratch** with NTP |
| Optimizer | **Per-Head Muon** (+ weight clipping) |
| Positional | **NoPE** on MLA; position via KDA decay/gating |
| Deploy quant (post-train QAT) | MoE expert weights **MXFP4**, activations **MXFP8** |
| Scaling claim | ~**2.5×** overall scaling efficiency vs Kimi K2 **`[claim from authors]`** |
| Post-train pillars | Multi-effort RL · Multi-Teacher On-Policy Distillation (MOPD) · agent environments · 1M agentic RL |
---
## 1. K3 Ingredient → Research Ancestry Map
| K3 ingredient | Direct primary | Immediate ancestry | Module |
|---|---|---|---|
| **KDA / Kimi Linear** | [2510.26692](https://arxiv.org/abs/2510.26692) | Gated DeltaNet → DeltaNet → GLA → Linear Transformer | M04M05 |
| **Gated MLA** | K3 §2.1.2 + DeepSeek-V2 MLA [2405.04434](https://arxiv.org/abs/2405.04434) | MQA/GQA → KV compression → latent KV | M03, M05 |
| **Attention Residuals** | [2603.15031](https://arxiv.org/abs/2603.15031) | ResNet residuals → PreNorm dilution problem | M02, M06 |
| **Stable LatentMoE** | K3 §2.3 + LatentMoE [2601.18089](https://arxiv.org/abs/2601.18089) | Switch/GShard → DeepSeekMoE → aux-loss-free routing → QB | M07, M09 |
| **MoonViT-V2** | K3 §2.4; prior MoonViT in Kimi-VL [2504.07491](https://arxiv.org/abs/2504.07491) | ViT → native-res VL encoders; K3 drops SigLIP init | M14 |
| **Per-head Muon** | K3 §2.5; Muon scale-up [2502.16982](https://arxiv.org/abs/2502.16982); MuonClip in K2 [2507.20534](https://arxiv.org/abs/2507.20534) | AdamW → matrix orthogonalization optimizers | M08 |
| **Scaling / data** | K3 §3; Kaplan/Hoffmann; K2 data rephrase | Scaling laws + domain mix + rephrase | M08 |
| **Long context (1M)** | K3 §3 progressive extension + KDA CP | RoPE/YaRN → Ring/Ulysses → NoPE+KDA | M11 |
| **Multi-effort RL** | K3 §4 effort ∈ {low, high, max} | o1-style test-time compute; k1.5 | M13 |
| **Multi-teacher on-policy distillation (MOPD)** | K3 §4.1.3; Thinking Machines on-policy distill | R1 distillation; on-policy KD | M13 |
| **Agent environments** | K3 §4.2 (pluggable harness, AET, knowledge graph) | Tool-use / SWE / OSWorld lineage | M14 |
| **KDA systems** | K3 §5.1 FlashKDA, KDA Context Parallelism | FLA / DeltaNet CP | M11, M15 |
| **MoonEP** | K3 §5 (perfect balance EP; vs DeepEP) | DeepEP / expert parallel | M11, M15 |
| **1M agentic RL** | K3 §5 partial rollouts, external KV, resumable sandboxes | Long-horizon RL infra | M14M15 |
| **MXFP4** | K3 QAT §4.1.4; OCP MX [2310.10537](https://arxiv.org/abs/2310.10537) | FP8 training → microscaling | M10 |
---
## 2. Module Graph & Prerequisites
```
M01 Foundations ──► M02 Residual/Norm/Pos
│ │
▼ ▼
M03 Softmax KV path M04 Linear/Delta/SSM path
│ │
└────────┬───────────┘
M05 Hybrid Attention (KDA+Gated MLA, NoPE)
M06 Depth Connectivity (AttnRes)
M07 Sparse MoE width (Stable LatentMoE)
M08 Scale / Data / Muon
M09 DeepSeek full stack (MLA→V3→R1→V3.2→V4)
┌───────────┼───────────┐
▼ ▼ ▼
M10 Quant M11 Long-ctx M12 Alignment (PPO/DPO)
│ systems │
└───────────┬───────────┘
M13 Reasoning RL + MOPD + multi-effort
M14 Agents + Vision (MoonViT-V2)
M15 K3 Capstone Systems (MoonEP, 1M RL, MXFP4 serve)
```
**Suggested learning order:** M01→M02→(M03∥M04)→M05→M06→M07→M08→M09→(M10∥M11∥M12)→M13→M14→M15.
---
## M01 — Foundations: Sequence Models & Softmax Attention
**Prerequisites:** undergrad ML, basic seq modeling
**Why for K3:** defines the quadratic attention baseline that KDA/MLA hybridize.
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2014 | Sequence to Sequence Learning with Neural Networks | https://arxiv.org/abs/1409.3215 | End-to-end neural transduction without alignments | Need better alignment / attention |
| 2015 | Neural Machine Translation by Jointly Learning to Align and Translate | https://arxiv.org/abs/1409.0473 | Soft attention over encoder states | Attention as differentiable lookup |
| 2017 | Attention Is All You Need | https://arxiv.org/abs/1706.03762 | Drop recurrence; multi-head self-attention + PE | Canonical Transformer stack |
| 2018 | BERT: Pre-training of Deep Bidirectional Transformers | https://arxiv.org/abs/1810.04805 | Bidirectional pretrain for understanding | Pretrainfinetune paradigm |
| 2018 | Improving Language Understanding by Generative Pre-Training (GPT) | https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf | Decoder-only LM pretrain | GPT lineage |
| 2019 | Language Models are Unsupervised Multitask Learners (GPT-2) | https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf | Zero-shot transfer via scale | Scaling narrative |
| 2020 | Language Models are Few-Shot Learners (GPT-3) | https://arxiv.org/abs/2005.14165 | In-context learning at 175B | ICL as emergent interface |
---
## M02 — Residual Stacks, Normalization, Position
**Prerequisites:** M01
**Why for K3:** PreNorm residual dilution motivates **AttnRes**; NoPE choice; RMSNorm in MoonViT-V2.
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2015 | Deep Residual Learning for Image Recognition | https://arxiv.org/abs/1512.03385 | Train very deep nets via identity skip | Residual highway becomes default |
| 2016 | Layer Normalization | https://arxiv.org/abs/1607.06450 | Stabilize RNN/Transformer activations | LN → PreNorm stacks |
| 2019 | Root Mean Square Layer Normalization | https://arxiv.org/abs/1910.07467 | Cheaper, stable norm without mean centering | Modern LLM default (incl. MoonViT-V2) |
| 2021 | RoFormer: Enhanced Transformer with Rotary Position Embedding | https://arxiv.org/abs/2104.09864 | Relative PE via rotations | Dominant PE until long-ctx hacks |
| 2023 | YaRN: Efficient Context Window Extension of LLMs | https://arxiv.org/abs/2309.00071 | Extend RoPE LMs beyond train length | Positional interpolation family |
| 2024 | The NoPE Hypothesis (and related NoPE studies) **`[select primary carefully; see also K3 NoPE practice]`** | https://arxiv.org/abs/2404.12224 | When explicit PE is unnecessary | K3: NoPE on MLA; position via KDA |
**K3 note:** MLA layers use **NoPE**; positional signal comes from KDAs channel-wise decay/gating (K3 §2.1, §3).
---
## M03 — Softmax Attention Efficiency & KV Compression → MLA
**Prerequisites:** M01M02
**Why for K3:** **Gated MLA** is the global-attention “anchor” layer in the 3:1 hybrid.
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2019 | Fast Transformer Decoding: One Write-Head is All You Need (MQA) | https://arxiv.org/abs/1911.02150 | Share KV across heads → smaller cache | Multi-query attention |
| 2023 | GQA: Training Generalized Multi-Query Transformer Models | https://arxiv.org/abs/2305.13245 | Interpolate MHA↔MQA | Production GQA default |
| 2022 | FlashAttention: Fast and Memory-Efficient Exact Attention | https://arxiv.org/abs/2205.14135 | IO-aware exact attention | Hardware-aware attention kernels |
| 2023 | FlashAttention-2 | https://arxiv.org/abs/2307.08691 | Higher occupancy / better parallelism | Training/prefill baseline |
| 2023 | LongNet: Scaling Transformers to 1B Tokens (dilated attn) | https://arxiv.org/abs/2307.02486 | Sparse patterns for extreme length | Sparse attention design space |
| 2024 | DeepSeek-V2: Strong Economical Efficient MoE LM (**MLA**) | https://arxiv.org/abs/2405.04434 | Compress KV into latent vector (MLA) | Latent cache for huge MoEs |
| 2024 | DeepSeek-V3 Technical Report (MLA at 671B) | https://arxiv.org/abs/2412.19437 | Scale MLA + DeepSeekMoE production-grade | Gated MLA inherits MLA |
| 2025 | DeepSeek-V3.2 (**DSA** on MLA) | https://arxiv.org/abs/2512.02556 | Fine-grained sparse attention + lightning indexer | Sparse selection over latent KV |
**K3 Gated MLA (official):** ungated MLA output gated by full-rank \(W_g\); gate matches full-rank KDA output gate (K3 §2.1.2).
---
## M04 — Linear Attention, Delta Rule, SSMs → KDA Ancestry
**Prerequisites:** M01
**Why for K3:** **KDA** = gated delta-rule recurrence with finer channel-wise forget + full-rank output gate.
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2020 | Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention | https://arxiv.org/abs/2006.16236 | Kernelize attention → linear time / RNN form | Linear Transformer baseline |
| 2021 | Linear Transformers Are Secretly Fast Weight Programmers | https://arxiv.org/abs/2102.11174 | Delta-rule / fast weights view of linear attn | **Delta rule** update idea |
| 2023 | Mamba: Linear-Time Sequence Modeling with Selective SSMs | https://arxiv.org/abs/2312.00752 | Input-dependent SSM selection | Modern linear-time competitor |
| 2024 | Gated Linear Attention Transformers (GLA) | https://arxiv.org/abs/2312.06635 | Data-dependent gates + FlashLinearAttention | Hardware-efficient gated linear attn |
| 2024 | Parallelizing Linear Transformers with the Delta Rule (DeltaNet) | https://arxiv.org/abs/2406.06484 | Parallel train delta-rule linear transformers | Chunkwise delta training |
| 2024 | Transformers are SSMs / Mamba-2 (SSD) | https://arxiv.org/abs/2405.21060 | Unify Transformer & SSM via structured duality | State-space dual algorithms |
| 2025 | Gated Delta Networks: Improving Mamba2 with Delta Rule | https://arxiv.org/abs/2412.06464 | Combine **gating + delta rule** | Direct parent of **KDA** |
| 2025 | Kimi Linear: Expressive Efficient Attention (**KDA**) | https://arxiv.org/abs/2510.26692 | Channel-wise fine-grained KDA + hybrid with MLA | Productionized in K3 (with K3-specific gates) |
**K3 KDA deltas vs Kimi Linear (official §2.1.1):** lower-bounded log-decay (tile-stable chunkwise form); **full-rank** output gate (vs low-rank in Linear).
---
## M05 — Hybrid Attention Architectures (KDA + Gated MLA + NoPE)
**Prerequisites:** M03, M04
**Why for K3:** core sequence mixer: **3×KDA + 1×Gated MLA**, final Gated MLA.
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2023 | Retentive Network (RetNet) | https://arxiv.org/abs/2307.08621 | Multi-scale retention recurrence | Hybrid linear/global motivation |
| 2024 | Griffin: Mixing Gated Linear Recurrences with Local Attention | https://arxiv.org/abs/2402.19427 | Mix local attention + linear recurrence | Hybrid design pattern |
| 2024 | Jamba: Hybrid Transformer-Mamba | https://arxiv.org/abs/2403.19887 | Production hybrid MoE | Hybrid at scale |
| 2024 | DeltaNet hybrids (sliding / global layers) *in* DeltaNet paper | https://arxiv.org/abs/2406.06484 | Sparse global layers restore quality | 3:1-style interleaving idea |
| 2025 | Kimi Linear (3:1 KDA:MLA hybrid) | https://arxiv.org/abs/2510.26692 | Match/beat full MLA quality at lower KV | Direct K3 attention recipe |
| 2026 | Kimi K3 (Gated MLA + refined KDA) | https://arxiv.org/abs/2607.24653 | Scale hybrid to 2.8T / 1M ctx | Capstone hybrid |
**Reading target:** derive why periodic full latent attention repairs linear-state information bottlenecks.
---
## M06 — Depth-wise Connectivity: PreNorm Dilution → Attention Residuals
**Prerequisites:** M02, M05
**Why for K3:** **Block AttnRes** is the depth information-flow upgrade.
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2015 | Deep Residual Learning (again as residual baseline) | https://arxiv.org/abs/1512.03385 | Identity mapping depth | Fixed unit residual weights |
| 2016 | Highway Networks | https://arxiv.org/abs/1505.00387 | Learned gates on residual paths | Content-dependent depth flow |
| 2017 | DenseNet | https://arxiv.org/abs/1608.06993 | Concatenate all prior feature maps | Dense cross-layer reuse |
| 2021 | DeepNet: Stabilizing Extremely Deep Transformers | https://arxiv.org/abs/2203.00555 | Scale residuals for 1000-layer Transformers | Deep PreNorm issues |
| 2026 | Attention Residuals (AttnRes / Block AttnRes) | https://arxiv.org/abs/2603.15031 | Softmax over prior layer outputs; block variant for memory | **K3 uses Block AttnRes** (S≈12) |
| 2026 | DeepSeek-V4 mHC (manifold-constrained hyper-connections) *official later residual variant* | https://arxiv.org/abs/2606.19348 | Alternative residual upgrade at 1M-ctx MoE | Parallel residual research thread |
**K3 detail:** pseudo-query \(q_l=w_l\); Block AttnRes reduces overhead \(O(Ld)\to O(Nd)\).
---
## M07 — Sparse Width: MoE → DeepSeekMoE → LatentMoE → Stable LatentMoE
**Prerequisites:** M01, M08 (scaling intuition helps)
**Why for K3:** **Stable LatentMoE** width path (16/896 + Quantile Balancing).
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2017 | Outrageously Large Neural Networks (Sparsely-Gated MoE) | https://arxiv.org/abs/1701.06538 | Conditional computation at scale | MoE modern start |
| 2020 | GShard: Scaling Giant Models with Conditional Computation | https://arxiv.org/abs/2006.16668 | Automatic sharding + MoE at TPU scale | Systems MoE |
| 2021 | Switch Transformers | https://arxiv.org/abs/2101.03961 | Simplify to top-1 routing; trillion-param | Load-balance auxiliaries |
| 2022 | ST-MoE | https://arxiv.org/abs/2202.08906 | Stable training recipes for sparse models | Stability folklore |
| 2024 | DeepSeekMoE: Towards Ultimate Expert Specialization | https://arxiv.org/abs/2401.06066 | Fine-grained experts + shared experts | DeepSeek MoE DNA |
| 2024 | Mixtral of Experts | https://arxiv.org/abs/2401.04088 | Strong open sparse LM | Sparse quality reference |
| 2024 | DeepSeek-V2 (DeepSeekMoE @ production) | https://arxiv.org/abs/2405.04434 | Economical training with sparse FFN | MLA+MoE combo |
| 2024 | DeepSeek-V3 (aux-loss-free routing) | https://arxiv.org/abs/2412.19437 | Bias-based load balance without aux loss | Aux-loss-free lineage |
| 2026 | LatentMoE: Optimal Accuracy per FLOP/Parameter | https://arxiv.org/abs/2601.18089 | Route in latent cheaper expert path | Namesake of **LatentMoE** |
| 2026 | Kimi K3 Stable LatentMoE + **Quantile Balancing** | https://arxiv.org/abs/2607.24653 | Stabilize extreme sparsity (896 experts, k=16) | Production K3 FFN |
**K3 QB (official):** expert bias from margin quantiles via global histogram all-reduce (not token-side exact quantiles).
**Also cite:** Jianlin Su blog *Travels in MoE* (load balance via optimal assignment) — referenced by K3 as [111] (Chinese primary blog, not arXiv).
---
## M08 — Scaling Laws, Data Recipes, Optimizers (→ Per-Head Muon)
**Prerequisites:** M01
**Why for K3:** 2.5× scaling-efficiency claim; Per-Head Muon; cosine > WSD in their search.
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2020 | Scaling Laws for Neural Language Models (Kaplan et al.) | https://arxiv.org/abs/2001.08361 | Power-law loss vs N, D, C | Classical scaling |
| 2022 | Training Compute-Optimal LLMs (Chinchilla / Hoffmann) | https://arxiv.org/abs/2203.15556 | Tokensparams balance | TPP retuning practice |
| 2014 | Adam: A Method for Stochastic Optimization | https://arxiv.org/abs/1412.6980 | Adaptive moments | Default baseline |
| 2017 | Decoupled Weight Decay Regularization (AdamW) | https://arxiv.org/abs/1711.05101 | Correct weight decay | Pre-Muon default |
| 2024 | Muon: An Optimizer for Hidden Layers (Keller Jordan et al.) | https://kellerjordan.github.io/posts/muon/ | NewtonSchulz orthogonalized momentum | Matrix-wise updates |
| 2025 | Muon is Scalable for LLM Training (Moonshot / Moonlight) | https://arxiv.org/abs/2502.16982 | Weight decay + update scale → Muon at LLM scale | Moonshot optimizer stack |
| 2025 | Kimi K2 (MuonClip / QK-Clip) | https://arxiv.org/abs/2507.20534 | Stabilize attention logits under Muon | Pretrain 15.5T zero spike claim |
| 2026 | Kimi K3 Per-Head Muon | https://arxiv.org/abs/2607.24653 | Orthogonalize QKV **per head** for balanced head scales | K3 default optimizer |
**Data / schedule landmarks**
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2020 | Exploring the Limits of Transfer Learning (T5) | https://arxiv.org/abs/1910.10683 | Unified text-to-text + C4 data | Data quality culture |
| 2022 | Training Compute-Optimal… (data side of Chinchilla) | https://arxiv.org/abs/2203.15556 | More tokens for given compute | Overtrain regimes |
| 2024 | DeepSeek LLM: Scaling Open-Source LMs with Longtermism | https://arxiv.org/abs/2401.02954 | Open scaling-law study + 67B | DeepSeek program start |
| 2025 | MiniCPM / WSD schedule discussion | https://arxiv.org/abs/2404.06395 | WarmupStableDecay alternative | K3 finds **cosine better** after independent HPO |
| 2025 | Kimi K2 data rephrasing recipe | https://arxiv.org/abs/2507.20534 | Knowledge/math rephrase + fidelity checks | Inherited by K3 §3.1 |
---
## M09 — DeepSeek Lineage (Dense → MoE → MLA → V3 → R1 → V3.2 → V4)
**Prerequisites:** M03, M07, M08, M12M13 for post-train pieces
**Policy:** include only **official** DeepSeek papers/reports.
### 9.1 Pretrain architecture stack
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2024 | DeepSeek LLM: Scaling Open-Source Language Models with Longtermism | https://arxiv.org/abs/2401.02954 | Open dense 7B/67B + scaling study | Foundation |
| 2024 | DeepSeekMoE: Ultimate Expert Specialization | https://arxiv.org/abs/2401.06066 | Fine-grained + shared experts | Sparse specialization |
| 2024 | DeepSeek-V2: MLA + DeepSeekMoE | https://arxiv.org/abs/2405.04434 | KV latent compression + economical MoE | **MLA birth** |
| 2024 | DeepSeek-V3 Technical Report | https://arxiv.org/abs/2412.19437 | 671B/37B act; **FP8**; **MTP**; **aux-loss-free routing**; **DualPipe** | Flagship pretrain stack |
| 2025 | DeepSeek-V3.2: DSA + scalable RL + agentic synthesis | https://arxiv.org/abs/2512.02556 | **DeepSeek Sparse Attention**; agent post-train scale | Long-ctx efficiency |
| 2026 | DeepSeek-V4: Million-Token Context Intelligence | https://arxiv.org/abs/2606.19348 | CSA+HCA hybrid attn; **mHC**; Muon; 1M ctx Pro/Flash | Official later 1M-ctx peer to K3 |
### 9.2 Math, RL, reasoning distillation
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2024 | DeepSeekMath (+ **GRPO**) | https://arxiv.org/abs/2402.03300 | Math continual pretrain + critic-free group RL | **GRPO** algorithm |
| 2025 | DeepSeek-R1: Incentivizing Reasoning via RL | https://arxiv.org/abs/2501.12948 | **R1-Zero** pure RL; **R1** multi-stage; **distillation** to dense 1.5B70B | Reasoning RL template |
| 2025 | DeepSeek-R1 Nature version (same line) | https://doi.org/10.1038/s41586-025-09422-z | Peer-reviewed presentation of R1 | Archival citation |
### 9.3 V3 systems keywords → primary sections
| Keyword | Where (primary) | Problem solved | Bridge |
|---|---|---|---|
| **FP8 mixed precision** | V3 report §3.3 | Low-precision train at 671B | → MX / QAT serving |
| **MTP** (multi-token prediction) | V3 report | Stronger pretrain objective / draft head | Speculative / denser signal |
| **Aux-loss-free routing** | V3 report | Balance without aux loss interference | → K3 QB different mechanism |
| **DualPipe** | V3 report §3.2.1 | Overlap PP bubbles + MoE dispatch/combine | → MoonEP alternative EP design |
| **DSA** | V3.2 report | Sparse token selection under MLA | Contrast with KDA hybrid |
| **DeepEP** | https://github.com/deepseek-ai/DeepEP | Expert-parallel comm library | K3 MoonEP builds on / contrasts |
**Atlas teaching note:** treat DeepSeek as the *dense reference stack* for MLA/MoE/FP8/RL; treat K3 as the *hybrid linear + AttnRes + Stable LatentMoE* fork at larger sparsity and 1M agentic RL.
---
## M10 — Low-Precision Training & Deployment (FP8 → MXFP4)
**Prerequisites:** M08M09
**Why for K3:** post-train **QAT** with **MXFP4 weights / MXFP8 activations** on MoE experts.
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2018 | Quantization and Training of Neural Networks for Integer-Arithmetic-Only Inference | https://arxiv.org/abs/1712.05877 | QAT for int inference | QAT paradigm |
| 2020 | Training with Quantization Noise | https://arxiv.org/abs/2004.07320 | Noise injection for robust quant | QAT variants |
| 2022 | FP8 Formats for Deep Learning (Micikevicius et al.) | https://arxiv.org/abs/2209.05433 | FP8 train/infer formats | Hardware FP8 |
| 2023 | Microscaling Data Formats for Deep Learning (**MX / MXFP4**) | https://arxiv.org/abs/2310.10537 | Block scales + narrow element types | **OCP MX** family |
| 2024 | DeepSeek-V3 FP8 training framework | https://arxiv.org/abs/2412.19437 | Validate FP8 at extreme MoE scale | Pretrain low-prec |
| 2026 | Kimi K3 MXFP4 QAT from SFT onward | https://arxiv.org/abs/2607.24653 | Deploy-time memory for 2.8T experts | Serving recipe |
**OCP primary (non-arXiv):** OCP Microscaling Formats (MX) Specification v1.0 — https://www.opencompute.org/documents/ocp-microscaling-formats-mx-v1-0-spec-final-pdf
---
## M11 — Long Context: Algorithms & Parallelism (→ KDA CP, 1M)
**Prerequisites:** M03M05
**Why for K3:** progressive 8K→64K→256K→**1M**; **KDA Context Parallelism**; state-aware prefix cache.
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2023 | ALiBi | https://arxiv.org/abs/2108.12409 | Length extrapolation via linear biases | PE-free-ish long ctx |
| 2023 | RoPE / YaRN (see M02) | https://arxiv.org/abs/2309.00071 | Extend rotary models | Common production path |
| 2023 | Ring Attention with Blockwise Transformers | https://arxiv.org/abs/2310.01889 | Sequence parallel for near-infinite context | Distributed long attn |
| 2023 | DeepSpeed Ulysses | https://arxiv.org/abs/2309.14509 | Sequence parallelism system opts | Cluster long-ctx train |
| 2024 | Linear Attention Sequence Parallelism (LASP) | https://arxiv.org/abs/2404.02882 | SP specialized for linear attn | Linear-SP ancestors |
| 2025 | LASP-2 / hybrid linear SP | https://arxiv.org/abs/2502.07864 **`[verify id if citing; K3 cites related]`** | Hybrid linear+softmax SP | Closer to KDA hybrids |
| 2025 | Context Parallelism for DeltaNet (Wang) | https://yywangcs.notion.site/DeltaNet-2a9fc9f5d8058013a498f34e0b25bd52 | CP for delta recurrence | Direct ancestor of **KCP** |
| 2026 | Kimi K3 KDA Context Parallelism + FlashKDA | https://arxiv.org/abs/2607.24653 · FlashKDA https://github.com/MoonshotAI/FlashKDA | Correct state transport under delta update \(M_t S_{t-1}\) | 1M train/prefill |
**K3 progressive extension (official):** pretrain grows 8K→64K; cooldown 256K→1M; needle/synthetic scatter tasks force full-window use.
---
## M12 — Alignment Foundations (RLHF / PPO / DPO)
**Prerequisites:** M01
**Why for K3:** base of modern post-train before GRPO / multi-effort / MOPD.
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2017 | Proximal Policy Optimization Algorithms | https://arxiv.org/abs/1707.06347 | Stable policy gradient updates | RL workhorse |
| 2022 | Training Language Models to Follow Instructions with Human Feedback (InstructGPT) | https://arxiv.org/abs/2203.02155 | RLHF pipeline for LLMs | Industry alignment template |
| 2022 | Constitutional AI | https://arxiv.org/abs/2212.08073 | Principle-based AI feedback | RLAIF direction |
| 2023 | Direct Preference Optimization (DPO) | https://arxiv.org/abs/2305.18290 | Preference learning without RL loop | Offline preference |
| 2023 | Llama 2: Open Foundation and Fine-Tuned Chat Models | https://arxiv.org/abs/2307.09288 | Open RLHF stack details | Open alignment recipes |
| 2024 | SimPO / ORPO family (optional shortlist) | https://arxiv.org/abs/2405.14734 | Simpler preference objectives | Alt to DPO |
---
## M13 — Reasoning RL, Multi-Effort, On-Policy Distillation
**Prerequisites:** M12, M09
**Why for K3:** multi-domain RL × effort levels; **MOPD** consolidation.
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2022 | Chain-of-Thought Prompting | https://arxiv.org/abs/2201.11903 | Elicit intermediate reasoning | Test-time reasoning culture |
| 2023 | Lets Verify Step by Step (process reward) | https://arxiv.org/abs/2305.20050 | Process vs outcome supervision | PRM path |
| 2024 | DeepSeekMath / **GRPO** | https://arxiv.org/abs/2402.03300 | Group-relative baseline, no critic | Memory-efficient RL |
| 2024 | OpenAI o1 announcement / “Learning to Reason with LLMs” | https://openai.com/index/learning-to-reason-with-llms/ | Scale RL + test-time compute | Multi-effort ancestor **`[blog primary; not full paper]`** |
| 2025 | Kimi k1.5: Scaling RL with LLMs | https://arxiv.org/abs/2501.12599 | Long-CoT RL scaling (Moonshot) | Kimi reasoning line |
| 2025 | DeepSeek-R1 / R1-Zero / distillation | https://arxiv.org/abs/2501.12948 | Pure RL emergence; then multi-stage + distill | Distill reasoning patterns |
| 2025 | On-policy distillation (Thinking Machines Lab) | https://thinkingmachines.ai/blog/on-policy-distillation/ **`[blog; K3 cites as Connectionism note]`** | Student on-policy w.r.t teacher | **MOPD** conceptual parent |
| 2026 | MiMo-V2-Flash Technical Report (multi-teacher mention lineage) | https://arxiv.org/abs/2601.02780 | Related multi-teacher post-train | Peer system cited by K3 |
| 2026 | Kimi K3 multi-effort RL + **MOPD** | https://arxiv.org/abs/2607.24653 | Domain experts {general, agentic, coding} × {low, high, max} → single policy | Capstone post-train |
**K3 MOPD (official):** on-policy student samples; multi-teacher logits/rewards consolidate specialists; top-k distill ablations showed no clear gain.
---
## M14 — Agents, Environments, Multimodal (MoonViT-V2)
**Prerequisites:** M13, M05
**Why for K3:** agent harness modularization; AET; native vision-in-the-loop; MoonViT-V2 from scratch.
### 14.1 Agent / tool / coding environments
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2023 | Toolformer | https://arxiv.org/abs/2302.04761 | Self-supervised tool API calls | Tool-use pretrain |
| 2023 | ReAct | https://arxiv.org/abs/2210.03629 | Interleave reason + act | Agent loop pattern |
| 2024 | SWE-bench | https://arxiv.org/abs/2310.06770 | Real GitHub issue resolution | Coding agent benchmark |
| 2024 | OSWorld | https://arxiv.org/abs/2404.07972 | Computer-use agents | GUI agent env |
| 2025 | Terminal-Bench | https://arxiv.org/abs/2502.14045 **`[confirm version K3 cites]`** | Hard CLI agent tasks | Terminal agents |
| 2025 | BrowseComp | https://arxiv.org/abs/2504.12516 | Browsing agents | Web agents |
| 2025 | Kimi K2 agentic data synthesis + joint RL | https://arxiv.org/abs/2507.20534 | Large synthetic tool trajectories | Kimi agent stack |
| 2026 | Kimi K2.5 Visual Agentic Intelligence | https://arxiv.org/abs/2602.02276 | Multimodal agents / swarm **`[swarm details in K2.5]`** | Vision agents |
| 2026 | Kimi K3 agent envs (pluggable harness, AET, knowledge graph) | https://arxiv.org/abs/2607.24653 | Avoid harness overfitting; long-horizon verifiers | Production agent RL data |
### 14.2 Vision encoders → MoonViT-V2
| Year | Title | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2020 | An Image is Worth 16x16 Words (ViT) | https://arxiv.org/abs/2010.11929 | Transformer vision backbone | ViT era |
| 2021 | Learning Transferable Visual Models From Natural Language Supervision (CLIP) | https://arxiv.org/abs/2103.00020 | Contrastive visionlanguage pretrain | SigLIP-style inits |
| 2023 | SigLIP | https://arxiv.org/abs/2303.15343 | Sigmoid contrastive VL | Prior MoonViT init |
| 2024 | LLaVA-OneVision | https://arxiv.org/abs/2408.03326 | Unified single/multi-image/video tasks | Native-res VL practice |
| 2025 | Kimi-VL Technical Report (**MoonViT**) | https://arxiv.org/abs/2504.07491 | Native-resolution MoonViT + MoE LM | Moonshot VL stack |
| 2026 | Kimi K3 **MoonViT-V2** from-scratch NTP | https://arxiv.org/abs/2607.24653 | Drop SigLIP init for stability; match quality | K3 vision path |
**K3 vision claim (official):** MoonViT-V2 ≈0.4B, 27 layers, RMSNorm, bias-free; shared image/video; 2×2 pixel-shuffle; up to 3584² in 1M context.
---
## M15 — Capstone Systems: KDA Kernels, MoonEP, 1M Agentic RL, Serve
**Prerequisites:** M05M07, M10M11, M13M14
**Why for K3:** makes 2.8T hybrid + 1M agentic RL trainable and deployable.
| Year | Title / artifact | URL | Problem solved | Bridge to next |
|---|---|---|---|---|
| 2019 | Triton: Intermediate Language for Tiled NN Kernels | https://www.eecs.harvard.edu/~htk/publication/2019-mapl-tillet-kung-cox.pdf | Productive GPU kernels | FLA / custom attn |
| 2021 | GPipe | https://arxiv.org/abs/1811.06965 | Pipeline parallel | PP lineage |
| 2021 | Megatron-LM / efficient large-scale training | https://arxiv.org/abs/2104.04473 | 3D parallelism | Pretrain systems baseline |
| 2020 | ZeRO | https://arxiv.org/abs/1910.02054 | Shard optimizer states | Memory for huge models |
| 2024 | FLA: Flash Linear Attention library | https://github.com/fla-org/flash-linear-attention | Kernel zoo for linear/delta | KDA ops host |
| 2025 | DeepEP | https://github.com/deepseek-ai/DeepEP | Expert-parallel communication | MoonEP contrast |
| 2025 | DualPipe (in V3 report) | https://arxiv.org/abs/2412.19437 | Overlap MoE+PP | Pipeline bubbles |
| 2026 | FlashKDA | https://github.com/MoonshotAI/FlashKDA | CUTLASS chunkwise KDA | Train/prefill speed |
| 2026 | Kimi K3 **MoonEP** | https://arxiv.org/abs/2607.24653 | Perfect balance EP, static shapes, zero-copy, bounded redundant experts | 2.8T MoE train |
| 2026 | Kimi K3 1M agentic RL co-located system | https://arxiv.org/abs/2607.24653 | Partial rollouts, external KV retention, resumable sandboxes | Long-horizon RL |
| 2026 | Kimi K3 state-aware KDA prefix caching + MXFP4 serve | https://arxiv.org/abs/2607.24653 | Decode across hybrid state + quant experts | Production deployment |
---
## 3. Module → K3 Ingredient Checklist
| Module | Delivers toward K3 |
|---|---|
| M01M02 | Transformer residual + PE literacy |
| M03 | MLA / Gated MLA |
| M04M05 | KDA + hybrid 3:1 |
| M06 | Attention Residuals |
| M07 | Stable LatentMoE + QB vs aux-loss-free |
| M08 | Scaling 2.5× story + Per-Head Muon |
| M09 | DeepSeek dense comparator (MLA, FP8, MTP, DualPipe, GRPO, R1, DSA, V4) |
| M10 | MXFP4/MXFP8 QAT |
| M11 | 1M context + KCP |
| M12M13 | Multi-effort RL + MOPD |
| M14 | Agent envs + MoonViT-V2 |
| M15 | MoonEP + 1M agentic RL systems |
---
## 4. DeepSeek Coverage Matrix (Atlas “strong track”)
| Topic | Primary | Module | Must-read density |
|---|---|---|---|
| DeepSeek LLM | 2401.02954 | M08/M09 | high |
| DeepSeekMoE | 2401.06066 | M07/M09 | high |
| V2 / MLA | 2405.04434 | M03/M09 | **core** |
| Math / GRPO | 2402.03300 | M13/M09 | **core** |
| V3 FP8 / MTP / aux-free / DualPipe | 2412.19437 | M09/M10/M15 | **core** |
| R1 / R1-Zero / distill | 2501.12948 (+ Nature DOI) | M13/M09 | **core** |
| V3.2 / DSA | 2512.02556 | M03/M09 | high |
| V4 (official later) | 2606.19348 | M06/M09/M11 | high (1M peer) |
| DeepEP | GitHub deepseek-ai/DeepEP | M15 | systems |
---
## 5. Uncertain / Careful Claims
| Claim | Status |
|---|---|
| K3 **2.5×** scaling efficiency vs K2 | Author-reported fit on OOD val curves (Fig. 7); **not independent replication** |
| Benchmarks vs Claude Fable 5 / GPT-5.6 Sol | Author suite; **external harness variance** possible |
| “First open 3T-class” marketing phrasing | Product language; verify total-param definitions (shared vs routed, embeddings, vision) |
| Exact MoonEP algorithm vs DeepEP | Described in K3 §5; **no standalone MoonEP paper** at time of roadmap |
| FlashKDA paper | Primarily **code artifact** + K3 description; limited standalone theory paper |
| LASP-2 arXiv number if used in syllabus | Double-check id before publishing student links |
| On-policy distillation “Connectionism” note | Blog-level primary; treat as concept source, not peer-reviewed algorithm paper |
| DeepSeek-V4 CSA/HCA vs K3 KDA | Parallel 1M-ctx designs; **do not equate** without side-by-side study |
| MoonViT-V2 “matches SigLIP init quality” | Author ablation claim (Fig. 6 + text) |
| Multi-effort exact reward formulas | Partially specified (thinking-token thresholds); full reward suite may be incomplete in report |
---
## 6. Compact Reading Paths (for site UX)
### Path A — “Understand K3 architecture in 12 papers”
1. Attention Is All You Need
2. FlashAttention-2
3. DeepSeek-V2 (MLA)
4. Gated DeltaNet
5. Kimi Linear (KDA)
6. Attention Residuals
7. DeepSeekMoE
8. LatentMoE
9. Muon is Scalable
10. Microscaling (MXFP4)
11. DeepSeek-V3 (systems)
12. **Kimi K3**
### Path B — “DeepSeek full stack (official only)”
DeepSeek LLM → DeepSeekMoE → DeepSeekMath/GRPO → V2/MLA → V3 → R1 → V3.2/DSA → V4
### Path C — “Post-train / agents”
InstructGPT → DPO → GRPO → R1 → k1.5 → K2 → K2.5 → K3 (MOPD + multi-effort + 1M agentic RL)
---
## 7. Suggested Atlas Module Metadata (CMS fields)
```yaml
module_id: M05
title_zh: 混合注意力:KDA 与 Gated MLA
prerequisites: [M03, M04]
k3_tags: [KDA, GatedMLA, NoPE, hybrid-3to1]
primary_papers: 6
estimated_hours: 8
```
---
## 8. Source Log (what this roadmap used)
| Source | Role |
|---|---|
| arXiv:2607.24653 K3 PDF text extract | Architecture, post-train, systems, bibliography |
| arXiv abstracts for DeepSeek V2/V3/V3.2/V4, R1, Math, MoE, LLM | Official lineage |
| arXiv:2510.26692 Kimi Linear | KDA definition |
| arXiv:2603.15031 AttnRes | Depth residual redesign |
| arXiv:2502.16982 / 2507.20534 | Muon / K2 |
| arXiv:2504.07491 / 2602.02276 | MoonViT / K2.5 |
| arXiv:2310.10537 + OCP MX | MXFP4 |
| arXiv:2601.18089 LatentMoE | Latent MoE naming/ancestry |
---
*End of structured research notes — LLM Atlas / K3-anchored roadmap.*
+46
View File
@@ -0,0 +1,46 @@
# 研究、引用与写作规范
## 来源优先级
1. **P0 一手论文**:arXiv、会议论文、期刊论文、作者正式技术报告。
2. **P1 一手实现**:作者官方仓库、模型卡、训练/评测代码。
3. **P2 官方说明**:实验室技术博客、产品文档、系统卡。
4. **P3 独立复现**:有方法、代码和数据的第三方复现或评测。
5. **P4 二手解释**:只用于发现线索,不承载关键事实。
正文关键结论原则上至少有一个 P0/P1 来源。模型发布当日尚无论文时,可以临时使用 P2,并清楚标注。
## 三类句子
- **事实**:论文明确报告的架构、数据、实验或结论,可直接引用。
- **解释**:为了教学而做的类比、拆解和重绘,标为“直觉解释”。
- **推断**:跨论文比较或报告未直接说明的因果判断,标为“我们的推断”并写出依据。
## 论文卡字段
- 稳定 ID、标题、作者/机构、年份;
- canonical URL、arXiv/DOI、发表状态;
- 所属专题、先修概念、后继工作;
- 它解决的问题、核心机制、关键证据;
- 一句话直觉、关键公式、图表索引;
- 限制、复现状态、与 K3/DeepSeek 的关系;
- 核验状态、核验日期、核验人/工具。
## 图表规范
- 优先原创重绘 SVG/HTML,可缩放并支持键盘阅读。
- 图题必须说明“原创示意”“依据某图改绘”或“原图引用”。
- 简化图不能伪装成模型的逐算子精确实现。
- 图中颜色始终保持语义一致:蓝色表示信息流,铜色表示稀疏选择,绿色表示训练/验证反馈,紫色表示系统状态。
## 质量闸门
章节从草稿到发布依次通过:
1. 结构检查:问题链是否完整;
2. 事实检查:数字、版本、时间、公式;
3. 来源检查:链接可达、引用贴近结论;
4. 教学检查:术语首次出现有解释,图能独立读懂;
5. 对照检查:K3/DeepSeek 映射是否准确;
6. 可访问性检查:键盘、对比度、移动端、减少动画;
7. 构建与链接检查。
+12
View File
@@ -0,0 +1,12 @@
# 本地来源缓存
这里用于研究期间缓存论文 PDF、网页快照与文本抽取。二进制和全文缓存默认被 `.gitignore` 排除;公开仓库只提交 canonical URL、校验信息、研究笔记和允许再分发的原创内容。
当前锚点:
- Kimi K3 Technical Report
<https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf>
- Kimi K3 official repository
<https://github.com/MoonshotAI/Kimi-K3>
- arXiv:2607.24653
<https://arxiv.org/abs/2607.24653>