# Attention / Transformer:Grok 候选线索 > 状态:**未核验发现队列,不是正式证据。** > > 生成方式:2026-07-29 使用本机 Grok CLI Headless 检索;主代理只保留可回查的候选节点,正文、数字、年份和归因不得直接引用本文件。 ## Grok 建议的十张问题账 1. 信息路径与可微检索:固定向量瓶颈怎样变成按需读取。 2. Q / K / V 匹配几何:加性打分、点积、缩放与 softmax。 3. 信息可见性:encoder self-attention、causal self-attention、cross-attention 与 prefix mask。 4. Multi-Head:多子空间表达、头冗余与后验电路分析。 5. 位置:绝对、相对、RoPE、ALiBi 与跨段记忆。 6. 局部计算:FFN、GLU、SwiGLU 与 MoE。 7. 深度路径:Residual、LayerNorm、Pre/Post-Norm、RMSNorm 与 AttnRes。 8. 架构与目标:encoder-only、decoder-only、encoder–decoder、MLM、causal LM、span corruption。 9. 系统成本:训练并行、Attention IO、KV Cache、MQA / GQA。 10. 当代映射:DeepSeek MLA / MTP / CSA-HCA 与 Kimi KDA / Gated MLA / AttnRes。 ## 候选节点(进入正式账本前必须回到原文) ### 信息路径 - Bahdanau et al. — *Neural Machine Translation by Jointly Learning to Align and Translate* - Luong et al. — *Effective Approaches to Attention-based Neural Machine Translation* - Graves et al. — *Neural Turing Machines* - Sukhbaatar et al. — *End-To-End Memory Networks* - Xu et al. — *Show, Attend and Tell* - Vinyals et al. — *Pointer Networks* ### QKV、Mask 与架构家族 - Vaswani et al. — *Attention Is All You Need* - Radford et al. — GPT-1 / GPT-2 technical reports - Devlin et al. — BERT - Raffel et al. — T5 - Dong et al. — UniLM - Child et al. — Sparse Transformer ### 多头与解释边界 - Michel et al. — *Are Sixteen Heads Really Better than One?* - Clark et al. — *What Does BERT Look at?* - Voita et al. — *Analyzing Multi-Head Self-Attention* - Jain & Wallace — *Attention is not Explanation* - Wiegreffe & Pinter — *Attention is not not Explanation* - Elhage et al. — *A Mathematical Framework for Transformer Circuits* ### 位置 - Shaw et al. — relative position representations - Dai et al. — Transformer-XL - Su et al. — RoFormer / RoPE - Press et al. — ALiBi ### FFN、残差与归一化 - He et al. — ResNet - Ba et al. — Layer Normalization - Xiong et al. — Pre-LN / Post-LN analysis - Zhang & Sennrich — RMSNorm - Dauphin et al. — GLU - Shazeer — GLU Variants / SwiGLU - Shazeer et al. — sparsely-gated MoE - Fedus et al. — Switch Transformer ### 系统效率 - Shazeer — Multi-Query Attention - Ainslie et al. — Grouped-Query Attention - Dao et al. — FlashAttention / FlashAttention-2 - Shoeybi et al. — Megatron-LM - Rajbhandari et al. — ZeRO ### DeepSeek / Kimi - DeepSeek-V2 — MLA + DeepSeekMoE - DeepSeek-V3 — MLA + MTP + auxiliary-loss-free balancing - DeepSeek-V4 — CSA / HCA + mHC - Kimi Linear — KDA + MLA hybrid - Attention Residuals — Full / Block AttnRes - Kimi K3 — 3:1 KDA / Gated MLA + AttnRes + Stable LatentMoE ## 明确拒绝直接采信的 Grok 表述 - “Attention 权重就是解释”:只可作为一个中间变量或诊断线索,不能自动升级为因果归因。 - “Transformer 就是 GPT”:原始论文是 encoder–decoder 机器翻译系统。 - “FlashAttention 把 Attention 变成线性复杂度”:FlashAttention 是 IO-aware 的精确实现,数学上的稠密注意力仍是二次计算。 - “MQA / GQA / MLA 是同一种方法”:它们都影响 KV 成本,但共享、分组与低秩潜变量压缩是不同机制。 - “KDA 已全面取代 softmax attention”:Kimi Linear / K3 使用混合层;结论限定于报告实验设置。 - “DeepSeek MTP 一定用于多 Token 推理”:V3 明确说推理可丢弃 MTP 模块,也可把它用于 speculative decoding。