# 语言模型前史:Grok 候选线索(未核验区) > 生成时间:2026-07-29 > 生成方式:Grok CLI headless(`--no-subagents`) > 状态:只用于扩大检索召回率,不作为网站正文证据。标题、年份、作者、DOI、实验数字都必须回到原论文或出版社页面核验。 > 正式证据:见 `LANGUAGE_MODEL_HISTORY_RESEARCH.md`。 ## 本轮提示词要求 围绕八张彼此独立的账,寻找 40–50 个候选一手节点: 1. 预测单位:字母、词、子词、字符与大词表; 2. 概率与信息:链式法则、熵、交叉熵、困惑度; 3. 上下文:Markov / N-gram、固定窗口、循环状态; 4. 稀疏性:未见事件、折扣、回退、插值; 5. 分布式表示:共现、矩阵分解、预测式词向量; 6. 循环记忆:SRN、长程梯度、LSTM、GRU; 7. 序列转导:encoder–decoder、固定向量、软对齐; 8. 评测与成本:基准、词表 Softmax、并行性、系统约束。 每个候选需要给出“旧瓶颈、核心贡献、遗留问题、适合重绘的图、最容易误传的边界”,并单列 Kimi K3、DeepSeek-V3/V4 对 next-token prediction 与 MTP 的继承关系。 ## 候选池摘要 ### 概率、信息与预测 - Shannon 1948 — *A Mathematical Theory of Communication* - Shannon 1951 — *Prediction and Entropy of Printed English* - Jelinek et al. 1977 — *Perplexity—a measure of the difficulty of speech recognition tasks* - Brown et al. 1992 — *Class-Based n-gram Models of Natural Language* - Chelba et al. 2013 — *One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling* ### 稀疏性与平滑 - Good 1953 — *The Population Frequencies of Species and the Estimation of Population Parameters* - Katz 1987 — *Estimation of Probabilities from Sparse Data for the Language Model Component of a Speech Recognizer* - Witten & Bell 1991 — *The Zero-Frequency Problem* - Kneser & Ney 1995 — *Improved Backing-off for M-gram Language Modeling* - Chen & Goodman 1996/1998/1999 — *An Empirical Study of Smoothing Techniques for Language Modeling* - Goodman 2001 — *A Bit of Progress in Language Modeling* ### 分布与表示 - Harris 1954 — *Distributional Structure* - Firth 1957 — *A Synopsis of Linguistic Theory 1930–1955* - Deerwester et al. 1990 — *Indexing by Latent Semantic Analysis* - Bengio et al. 2003 — *A Neural Probabilistic Language Model* - Morin & Bengio 2005 — *Hierarchical Probabilistic Neural Network Language Model* - Collobert & Weston 2008 / Collobert et al. 2011 - Mikolov et al. 2013 — CBOW / Skip-gram - Mikolov et al. 2013 — negative sampling / phrases - Pennington et al. 2014 — GloVe - Levy & Goldberg 2014 — SGNS as implicit matrix factorization ### 循环记忆 - Elman 1990 — *Finding Structure in Time* - Bengio, Simard & Frasconi 1994 — *Learning long-term dependencies with gradient descent is difficult* - Hochreiter & Schmidhuber 1997 — *Long Short-Term Memory* - Gers, Schmidhuber & Cummins 2000 — forget gate - Mikolov et al. 2010 — RNN language model - Graves 2013 — recurrent sequence generation - Cho et al. 2014 — encoder–decoder / GRU ### 序列转导与 Attention 前夜 - Kalchbrenner & Blunsom 2013 — recurrent continuous translation - Cho et al. 2014 — RNN encoder–decoder - Sutskever, Vinyals & Le 2014 — Seq2Seq - Bahdanau, Cho & Bengio 2014/2015 — additive attention - Luong, Pham & Manning 2015 — global/local attention - Sennrich et al. 2016 — BPE for rare words - Gehring et al. 2017 — convolutional Seq2Seq - Vaswani et al. 2017 — Transformer(本章终点,不在本章展开) ### 词表、评测与系统成本 - Papineni et al. 2002 — BLEU - Wu et al. 2016 — GNMT - Jozefowicz et al. 2016 — large-scale RNN language modeling - Grave et al. 2016 — adaptive softmax - Press & Wolf 2016 — input/output embedding weight tying - Dauphin et al. 2016 — gated convolutional language modeling - Shazeer et al. 2017 — sparsely-gated MoE - Merity et al. 2017 — AWD-LSTM ### 当代回声 - Gloeckle et al. 2024 — multi-token prediction - DeepSeek-V3 Technical Report - DeepSeek-V4 Technical Report - Kimi K3 Technical Report ## Grok 输出里需要特别警惕的地方 1. 它把若干“历史联系”写成了直接因果;正式正文只称“概念桥”或“后见之明下的连续性”。 2. Brown 1992 的作者名和若干旧论文页码有拼写风险,不能照抄。 3. Chen–Goodman 有 1996 会议版、1998 技术报告和 1999 期刊版,正式引用必须固定版本。 4. 现代课堂常见的三门 LSTM 含后来的 forget gate,不能全部算到 1997 原论文。 5. Word2Vec 是词表示训练方法,不是完整句子语言模型。 6. Bahdanau additive attention 与 Transformer scaled dot-product attention 不是同一公式。 7. “RNN 理论上能看无限历史”不能写成“实践中记得无限历史”。 8. MTP 是辅助训练/推测解码桥梁,不等于取消单步自回归生成。 9. K3 的视觉与文本统一 next-token objective 和它的一层 MTP 同时存在,二者不矛盾。 10. DeepSeek-V4 的版本号、标题和细节必须以本地保存的官方技术报告为准。 ## 已经被正式核准的高价值线索 - Good 1953 DOI:`10.1093/biomet/40.3-4.237` - Elman 1990 DOI:`10.1207/s15516709cog1402_1` - LSTM 1997 DOI:`10.1162/neco.1997.9.8.1735` - Morin & Bengio 2005:PMLR R5:246–252 - Luong et al. 2015:ACL Anthology D15-1166,DOI `10.18653/v1/D15-1166` - Mikolov et al. 2010:ISCA DOI `10.21437/Interspeech.2010-343` ## 留在发现队列、暂不进入正文的候选 - Firth 1957:重要但书章元数据不稳定,且不是计算模型; - Gage 1994:BPE 压缩起源的网页副本与正式书目信息需要额外核查; - Witten–Bell 1991:可作为平滑旁支,但会分散主线; - Collobert 系列:更偏通用 NLP 表示,不是本章最短因果链; - Kalchbrenner–Blunsom 2013:重要前驱,但正文空间优先给 Cho / Sutskever / Bahdanau; - BLEU / WMT:属于序列转导评测,正文只保留边界说明; - GNMT:系统史重要,首版暂放扩展阅读; - Gloeckle MTP:本章当代回声可引用,MTP 机制主体仍留给推理服务章节。