Files
llm-atlas/research/LANGUAGE_MODEL_HISTORY_GROK_LEADS.md
T
2026-07-29 04:47:36 +08:00

6.2 KiB
Raw Blame History

语言模型前史:Grok 候选线索(未核验区)

生成时间:2026-07-29
生成方式:Grok CLI headless--no-subagents
状态:只用于扩大检索召回率,不作为网站正文证据。标题、年份、作者、DOI、实验数字都必须回到原论文或出版社页面核验。
正式证据:见 LANGUAGE_MODEL_HISTORY_RESEARCH.md

本轮提示词要求

围绕八张彼此独立的账,寻找 40–50 个候选一手节点:

  1. 预测单位:字母、词、子词、字符与大词表;
  2. 概率与信息:链式法则、熵、交叉熵、困惑度;
  3. 上下文:Markov / N-gram、固定窗口、循环状态;
  4. 稀疏性:未见事件、折扣、回退、插值;
  5. 分布式表示:共现、矩阵分解、预测式词向量;
  6. 循环记忆:SRN、长程梯度、LSTM、GRU;
  7. 序列转导:encoder–decoder、固定向量、软对齐;
  8. 评测与成本:基准、词表 Softmax、并行性、系统约束。

每个候选需要给出“旧瓶颈、核心贡献、遗留问题、适合重绘的图、最容易误传的边界”,并单列 Kimi K3、DeepSeek-V3/V4 对 next-token prediction 与 MTP 的继承关系。

候选池摘要

概率、信息与预测

  • Shannon 1948 — A Mathematical Theory of Communication
  • Shannon 1951 — Prediction and Entropy of Printed English
  • Jelinek et al. 1977 — Perplexity—a measure of the difficulty of speech recognition tasks
  • Brown et al. 1992 — Class-Based n-gram Models of Natural Language
  • Chelba et al. 2013 — One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling

稀疏性与平滑

  • Good 1953 — The Population Frequencies of Species and the Estimation of Population Parameters
  • Katz 1987 — Estimation of Probabilities from Sparse Data for the Language Model Component of a Speech Recognizer
  • Witten & Bell 1991 — The Zero-Frequency Problem
  • Kneser & Ney 1995 — Improved Backing-off for M-gram Language Modeling
  • Chen & Goodman 1996/1998/1999 — An Empirical Study of Smoothing Techniques for Language Modeling
  • Goodman 2001 — A Bit of Progress in Language Modeling

分布与表示

  • Harris 1954 — Distributional Structure
  • Firth 1957 — A Synopsis of Linguistic Theory 19301955
  • Deerwester et al. 1990 — Indexing by Latent Semantic Analysis
  • Bengio et al. 2003 — A Neural Probabilistic Language Model
  • Morin & Bengio 2005 — Hierarchical Probabilistic Neural Network Language Model
  • Collobert & Weston 2008 / Collobert et al. 2011
  • Mikolov et al. 2013 — CBOW / Skip-gram
  • Mikolov et al. 2013 — negative sampling / phrases
  • Pennington et al. 2014 — GloVe
  • Levy & Goldberg 2014 — SGNS as implicit matrix factorization

循环记忆

  • Elman 1990 — Finding Structure in Time
  • Bengio, Simard & Frasconi 1994 — Learning long-term dependencies with gradient descent is difficult
  • Hochreiter & Schmidhuber 1997 — Long Short-Term Memory
  • Gers, Schmidhuber & Cummins 2000 — forget gate
  • Mikolov et al. 2010 — RNN language model
  • Graves 2013 — recurrent sequence generation
  • Cho et al. 2014 — encoderdecoder / GRU

序列转导与 Attention 前夜

  • Kalchbrenner & Blunsom 2013 — recurrent continuous translation
  • Cho et al. 2014 — RNN encoderdecoder
  • Sutskever, Vinyals & Le 2014 — Seq2Seq
  • Bahdanau, Cho & Bengio 2014/2015 — additive attention
  • Luong, Pham & Manning 2015 — global/local attention
  • Sennrich et al. 2016 — BPE for rare words
  • Gehring et al. 2017 — convolutional Seq2Seq
  • Vaswani et al. 2017 — Transformer(本章终点,不在本章展开)

词表、评测与系统成本

  • Papineni et al. 2002 — BLEU
  • Wu et al. 2016 — GNMT
  • Jozefowicz et al. 2016 — large-scale RNN language modeling
  • Grave et al. 2016 — adaptive softmax
  • Press & Wolf 2016 — input/output embedding weight tying
  • Dauphin et al. 2016 — gated convolutional language modeling
  • Shazeer et al. 2017 — sparsely-gated MoE
  • Merity et al. 2017 — AWD-LSTM

当代回声

  • Gloeckle et al. 2024 — multi-token prediction
  • DeepSeek-V3 Technical Report
  • DeepSeek-V4 Technical Report
  • Kimi K3 Technical Report

Grok 输出里需要特别警惕的地方

  1. 它把若干“历史联系”写成了直接因果;正式正文只称“概念桥”或“后见之明下的连续性”。
  2. Brown 1992 的作者名和若干旧论文页码有拼写风险,不能照抄。
  3. ChenGoodman 有 1996 会议版、1998 技术报告和 1999 期刊版,正式引用必须固定版本。
  4. 现代课堂常见的三门 LSTM 含后来的 forget gate,不能全部算到 1997 原论文。
  5. Word2Vec 是词表示训练方法,不是完整句子语言模型。
  6. Bahdanau additive attention 与 Transformer scaled dot-product attention 不是同一公式。
  7. “RNN 理论上能看无限历史”不能写成“实践中记得无限历史”。
  8. MTP 是辅助训练/推测解码桥梁,不等于取消单步自回归生成。
  9. K3 的视觉与文本统一 next-token objective 和它的一层 MTP 同时存在,二者不矛盾。
  10. DeepSeek-V4 的版本号、标题和细节必须以本地保存的官方技术报告为准。

已经被正式核准的高价值线索

  • Good 1953 DOI10.1093/biomet/40.3-4.237
  • Elman 1990 DOI10.1207/s15516709cog1402_1
  • LSTM 1997 DOI10.1162/neco.1997.9.8.1735
  • Morin & Bengio 2005PMLR R5:246252
  • Luong et al. 2015ACL Anthology D15-1166DOI 10.18653/v1/D15-1166
  • Mikolov et al. 2010ISCA DOI 10.21437/Interspeech.2010-343

留在发现队列、暂不进入正文的候选

  • Firth 1957:重要但书章元数据不稳定,且不是计算模型;
  • Gage 1994:BPE 压缩起源的网页副本与正式书目信息需要额外核查;
  • WittenBell 1991:可作为平滑旁支,但会分散主线;
  • Collobert 系列:更偏通用 NLP 表示,不是本章最短因果链;
  • KalchbrennerBlunsom 2013:重要前驱,但正文空间优先给 Cho / Sutskever / Bahdanau
  • BLEU / WMT:属于序列转导评测,正文只保留边界说明;
  • GNMT:系统史重要,首版暂放扩展阅读;
  • Gloeckle MTP:本章当代回声可引用,MTP 机制主体仍留给推理服务章节。