feat: publish language model origins chapter

This commit is contained in:
wuyang
2026-07-29 04:47:36 +08:00
parent 2997434f7f
commit a54152dc24
22 changed files with 5521 additions and 35 deletions
@@ -0,0 +1,126 @@
# 语言模型前史:Grok 候选线索(未核验区)
> 生成时间:2026-07-29
> 生成方式:Grok CLI headless`--no-subagents`
> 状态:只用于扩大检索召回率,不作为网站正文证据。标题、年份、作者、DOI、实验数字都必须回到原论文或出版社页面核验。
> 正式证据:见 `LANGUAGE_MODEL_HISTORY_RESEARCH.md`。
## 本轮提示词要求
围绕八张彼此独立的账,寻找 40–50 个候选一手节点:
1. 预测单位:字母、词、子词、字符与大词表;
2. 概率与信息:链式法则、熵、交叉熵、困惑度;
3. 上下文:Markov / N-gram、固定窗口、循环状态;
4. 稀疏性:未见事件、折扣、回退、插值;
5. 分布式表示:共现、矩阵分解、预测式词向量;
6. 循环记忆:SRN、长程梯度、LSTM、GRU;
7. 序列转导:encoder–decoder、固定向量、软对齐;
8. 评测与成本:基准、词表 Softmax、并行性、系统约束。
每个候选需要给出“旧瓶颈、核心贡献、遗留问题、适合重绘的图、最容易误传的边界”,并单列 Kimi K3、DeepSeek-V3/V4 对 next-token prediction 与 MTP 的继承关系。
## 候选池摘要
### 概率、信息与预测
- Shannon 1948 — *A Mathematical Theory of Communication*
- Shannon 1951 — *Prediction and Entropy of Printed English*
- Jelinek et al. 1977 — *Perplexity—a measure of the difficulty of speech recognition tasks*
- Brown et al. 1992 — *Class-Based n-gram Models of Natural Language*
- Chelba et al. 2013 — *One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling*
### 稀疏性与平滑
- Good 1953 — *The Population Frequencies of Species and the Estimation of Population Parameters*
- Katz 1987 — *Estimation of Probabilities from Sparse Data for the Language Model Component of a Speech Recognizer*
- Witten & Bell 1991 — *The Zero-Frequency Problem*
- Kneser & Ney 1995 — *Improved Backing-off for M-gram Language Modeling*
- Chen & Goodman 1996/1998/1999 — *An Empirical Study of Smoothing Techniques for Language Modeling*
- Goodman 2001 — *A Bit of Progress in Language Modeling*
### 分布与表示
- Harris 1954 — *Distributional Structure*
- Firth 1957 — *A Synopsis of Linguistic Theory 19301955*
- Deerwester et al. 1990 — *Indexing by Latent Semantic Analysis*
- Bengio et al. 2003 — *A Neural Probabilistic Language Model*
- Morin & Bengio 2005 — *Hierarchical Probabilistic Neural Network Language Model*
- Collobert & Weston 2008 / Collobert et al. 2011
- Mikolov et al. 2013 — CBOW / Skip-gram
- Mikolov et al. 2013 — negative sampling / phrases
- Pennington et al. 2014 — GloVe
- Levy & Goldberg 2014 — SGNS as implicit matrix factorization
### 循环记忆
- Elman 1990 — *Finding Structure in Time*
- Bengio, Simard & Frasconi 1994 — *Learning long-term dependencies with gradient descent is difficult*
- Hochreiter & Schmidhuber 1997 — *Long Short-Term Memory*
- Gers, Schmidhuber & Cummins 2000 — forget gate
- Mikolov et al. 2010 — RNN language model
- Graves 2013 — recurrent sequence generation
- Cho et al. 2014 — encoderdecoder / GRU
### 序列转导与 Attention 前夜
- Kalchbrenner & Blunsom 2013 — recurrent continuous translation
- Cho et al. 2014 — RNN encoderdecoder
- Sutskever, Vinyals & Le 2014 — Seq2Seq
- Bahdanau, Cho & Bengio 2014/2015 — additive attention
- Luong, Pham & Manning 2015 — global/local attention
- Sennrich et al. 2016 — BPE for rare words
- Gehring et al. 2017 — convolutional Seq2Seq
- Vaswani et al. 2017 — Transformer(本章终点,不在本章展开)
### 词表、评测与系统成本
- Papineni et al. 2002 — BLEU
- Wu et al. 2016 — GNMT
- Jozefowicz et al. 2016 — large-scale RNN language modeling
- Grave et al. 2016 — adaptive softmax
- Press & Wolf 2016 — input/output embedding weight tying
- Dauphin et al. 2016 — gated convolutional language modeling
- Shazeer et al. 2017 — sparsely-gated MoE
- Merity et al. 2017 — AWD-LSTM
### 当代回声
- Gloeckle et al. 2024 — multi-token prediction
- DeepSeek-V3 Technical Report
- DeepSeek-V4 Technical Report
- Kimi K3 Technical Report
## Grok 输出里需要特别警惕的地方
1. 它把若干“历史联系”写成了直接因果;正式正文只称“概念桥”或“后见之明下的连续性”。
2. Brown 1992 的作者名和若干旧论文页码有拼写风险,不能照抄。
3. ChenGoodman 有 1996 会议版、1998 技术报告和 1999 期刊版,正式引用必须固定版本。
4. 现代课堂常见的三门 LSTM 含后来的 forget gate,不能全部算到 1997 原论文。
5. Word2Vec 是词表示训练方法,不是完整句子语言模型。
6. Bahdanau additive attention 与 Transformer scaled dot-product attention 不是同一公式。
7. “RNN 理论上能看无限历史”不能写成“实践中记得无限历史”。
8. MTP 是辅助训练/推测解码桥梁,不等于取消单步自回归生成。
9. K3 的视觉与文本统一 next-token objective 和它的一层 MTP 同时存在,二者不矛盾。
10. DeepSeek-V4 的版本号、标题和细节必须以本地保存的官方技术报告为准。
## 已经被正式核准的高价值线索
- Good 1953 DOI`10.1093/biomet/40.3-4.237`
- Elman 1990 DOI`10.1207/s15516709cog1402_1`
- LSTM 1997 DOI`10.1162/neco.1997.9.8.1735`
- Morin & Bengio 2005PMLR R5:246252
- Luong et al. 2015ACL Anthology D15-1166DOI `10.18653/v1/D15-1166`
- Mikolov et al. 2010ISCA DOI `10.21437/Interspeech.2010-343`
## 留在发现队列、暂不进入正文的候选
- Firth 1957:重要但书章元数据不稳定,且不是计算模型;
- Gage 1994:BPE 压缩起源的网页副本与正式书目信息需要额外核查;
- WittenBell 1991:可作为平滑旁支,但会分散主线;
- Collobert 系列:更偏通用 NLP 表示,不是本章最短因果链;
- KalchbrennerBlunsom 2013:重要前驱,但正文空间优先给 Cho / Sutskever / Bahdanau
- BLEU / WMT:属于序列转导评测,正文只保留边界说明;
- GNMT:系统史重要,首版暂放扩展阅读;
- Gloeckle MTP:本章当代回声可引用,MTP 机制主体仍留给推理服务章节。