Files
llm-atlas/research/LANGUAGE_MODEL_HISTORY_GROK_LEADS.md
T
2026-07-29 04:47:36 +08:00

127 lines
6.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 语言模型前史:Grok 候选线索(未核验区)
> 生成时间:2026-07-29
> 生成方式:Grok CLI headless`--no-subagents`
> 状态:只用于扩大检索召回率,不作为网站正文证据。标题、年份、作者、DOI、实验数字都必须回到原论文或出版社页面核验。
> 正式证据:见 `LANGUAGE_MODEL_HISTORY_RESEARCH.md`。
## 本轮提示词要求
围绕八张彼此独立的账,寻找 40–50 个候选一手节点:
1. 预测单位:字母、词、子词、字符与大词表;
2. 概率与信息:链式法则、熵、交叉熵、困惑度;
3. 上下文:Markov / N-gram、固定窗口、循环状态;
4. 稀疏性:未见事件、折扣、回退、插值;
5. 分布式表示:共现、矩阵分解、预测式词向量;
6. 循环记忆:SRN、长程梯度、LSTM、GRU;
7. 序列转导:encoder–decoder、固定向量、软对齐;
8. 评测与成本:基准、词表 Softmax、并行性、系统约束。
每个候选需要给出“旧瓶颈、核心贡献、遗留问题、适合重绘的图、最容易误传的边界”,并单列 Kimi K3、DeepSeek-V3/V4 对 next-token prediction 与 MTP 的继承关系。
## 候选池摘要
### 概率、信息与预测
- Shannon 1948 — *A Mathematical Theory of Communication*
- Shannon 1951 — *Prediction and Entropy of Printed English*
- Jelinek et al. 1977 — *Perplexity—a measure of the difficulty of speech recognition tasks*
- Brown et al. 1992 — *Class-Based n-gram Models of Natural Language*
- Chelba et al. 2013 — *One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling*
### 稀疏性与平滑
- Good 1953 — *The Population Frequencies of Species and the Estimation of Population Parameters*
- Katz 1987 — *Estimation of Probabilities from Sparse Data for the Language Model Component of a Speech Recognizer*
- Witten & Bell 1991 — *The Zero-Frequency Problem*
- Kneser & Ney 1995 — *Improved Backing-off for M-gram Language Modeling*
- Chen & Goodman 1996/1998/1999 — *An Empirical Study of Smoothing Techniques for Language Modeling*
- Goodman 2001 — *A Bit of Progress in Language Modeling*
### 分布与表示
- Harris 1954 — *Distributional Structure*
- Firth 1957 — *A Synopsis of Linguistic Theory 19301955*
- Deerwester et al. 1990 — *Indexing by Latent Semantic Analysis*
- Bengio et al. 2003 — *A Neural Probabilistic Language Model*
- Morin & Bengio 2005 — *Hierarchical Probabilistic Neural Network Language Model*
- Collobert & Weston 2008 / Collobert et al. 2011
- Mikolov et al. 2013 — CBOW / Skip-gram
- Mikolov et al. 2013 — negative sampling / phrases
- Pennington et al. 2014 — GloVe
- Levy & Goldberg 2014 — SGNS as implicit matrix factorization
### 循环记忆
- Elman 1990 — *Finding Structure in Time*
- Bengio, Simard & Frasconi 1994 — *Learning long-term dependencies with gradient descent is difficult*
- Hochreiter & Schmidhuber 1997 — *Long Short-Term Memory*
- Gers, Schmidhuber & Cummins 2000 — forget gate
- Mikolov et al. 2010 — RNN language model
- Graves 2013 — recurrent sequence generation
- Cho et al. 2014 — encoderdecoder / GRU
### 序列转导与 Attention 前夜
- Kalchbrenner & Blunsom 2013 — recurrent continuous translation
- Cho et al. 2014 — RNN encoderdecoder
- Sutskever, Vinyals & Le 2014 — Seq2Seq
- Bahdanau, Cho & Bengio 2014/2015 — additive attention
- Luong, Pham & Manning 2015 — global/local attention
- Sennrich et al. 2016 — BPE for rare words
- Gehring et al. 2017 — convolutional Seq2Seq
- Vaswani et al. 2017 — Transformer(本章终点,不在本章展开)
### 词表、评测与系统成本
- Papineni et al. 2002 — BLEU
- Wu et al. 2016 — GNMT
- Jozefowicz et al. 2016 — large-scale RNN language modeling
- Grave et al. 2016 — adaptive softmax
- Press & Wolf 2016 — input/output embedding weight tying
- Dauphin et al. 2016 — gated convolutional language modeling
- Shazeer et al. 2017 — sparsely-gated MoE
- Merity et al. 2017 — AWD-LSTM
### 当代回声
- Gloeckle et al. 2024 — multi-token prediction
- DeepSeek-V3 Technical Report
- DeepSeek-V4 Technical Report
- Kimi K3 Technical Report
## Grok 输出里需要特别警惕的地方
1. 它把若干“历史联系”写成了直接因果;正式正文只称“概念桥”或“后见之明下的连续性”。
2. Brown 1992 的作者名和若干旧论文页码有拼写风险,不能照抄。
3. ChenGoodman 有 1996 会议版、1998 技术报告和 1999 期刊版,正式引用必须固定版本。
4. 现代课堂常见的三门 LSTM 含后来的 forget gate,不能全部算到 1997 原论文。
5. Word2Vec 是词表示训练方法,不是完整句子语言模型。
6. Bahdanau additive attention 与 Transformer scaled dot-product attention 不是同一公式。
7. “RNN 理论上能看无限历史”不能写成“实践中记得无限历史”。
8. MTP 是辅助训练/推测解码桥梁,不等于取消单步自回归生成。
9. K3 的视觉与文本统一 next-token objective 和它的一层 MTP 同时存在,二者不矛盾。
10. DeepSeek-V4 的版本号、标题和细节必须以本地保存的官方技术报告为准。
## 已经被正式核准的高价值线索
- Good 1953 DOI`10.1093/biomet/40.3-4.237`
- Elman 1990 DOI`10.1207/s15516709cog1402_1`
- LSTM 1997 DOI`10.1162/neco.1997.9.8.1735`
- Morin & Bengio 2005PMLR R5:246252
- Luong et al. 2015ACL Anthology D15-1166DOI `10.18653/v1/D15-1166`
- Mikolov et al. 2010ISCA DOI `10.21437/Interspeech.2010-343`
## 留在发现队列、暂不进入正文的候选
- Firth 1957:重要但书章元数据不稳定,且不是计算模型;
- Gage 1994:BPE 压缩起源的网页副本与正式书目信息需要额外核查;
- WittenBell 1991:可作为平滑旁支,但会分散主线;
- Collobert 系列:更偏通用 NLP 表示,不是本章最短因果链;
- KalchbrennerBlunsom 2013:重要前驱,但正文空间优先给 Cho / Sutskever / Bahdanau
- BLEU / WMT:属于序列转导评测,正文只保留边界说明;
- GNMT:系统史重要,首版暂放扩展阅读;
- Gloeckle MTP:本章当代回声可引用,MTP 机制主体仍留给推理服务章节。