Files
llm-atlas/research/K3_ATTNRES_LOCAL_PATH_SCOPING.md

198 lines
7.9 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# K3 Attention Residuals 局部 mixer 路径:Round 07 前置定位
研究日期:2026-07-30
阶段身份:**定向 scoping,不是 Round 07 预注册结果**
上游协议:`llm-atlas-k3-attnres-spike-path-v1`
## 1. 已知到什么程度
Round 06 在同一个 depth-32 / Block AttnRes 缩小模型上,把前向保持为 learned
weights,只改写 mixer 的反向规则。三 seed 的最终 `post_mlp_state / element_rms`
结果为:
| backward rule | spike contrast(3-seed mean) | 相对 detached learned | peak normalized(3-seed mean) | 相对 detached learned |
|---|---:|---:|---:|---:|
| learned | 2.754 | — | 4.803 | — |
| detached learned | 2.812 | reference | 4.956 | reference |
| uniform value / all 65 mixers | 0.837 | **−70.2%** | 3.122 | **−37.0%** |
其中:
- `learned → detached learned` 没有降低尖峰,contrast 反而平均增加约 2.0%;
- `detached learned → uniform value backward` 在 contrast 和 peak 上都 3 / 3 seed
超过预注册的 20% material threshold;
- 三种模式的 logits、loss、六位置 activation 和 mixer forward 摘要全部 exact;
- 这证明的是**全局 backward-rule sensitivity**,不是训练变体,也不是局部归因。
因此 Round 07 不再重复问“learned value coefficients 是否重要”,而是问:
> 65 个 mixer 全局改写带来的下降,主要能否由尖峰邻近的 group 6 / 7
> depth mixers 复现,并能否从反方向恢复?
## 2. 固定拓扑,而不是结果后挑层
depth 32 的 Block AttnRes 有 8 个 aggregation groups,每组 4 个 Transformer
blocks。每层有 attention 和 MLP 两个 depth mixers,合计 64 个;模型末尾还有一个
独立 output mixer,合计 65 个 intervention nodes。
对 1-based layer `l`:
```text
group = floor((l - 1) / 4) + 1
attention mixer index = 2 × (l - 1) # 0-based
MLP mixer index = 2 × (l - 1) + 1 # 0-based
```
所以:
| scope | layers | 0-based depth mixer indices | mixer count |
|---|---:|---:|---:|
| group 6 | 21–24 | 40–47 | 8 |
| group 7 | 25–28 | 48–55 | 8 |
| groups 6+7 | 21–28 | 40–55 | 16 |
| output | — | separate node | 1 |
Round 05 已在看过数据后冻结尖峰集合 `S = layers 21–25`。它覆盖完整 group 6 和
group 7 的首层。因此 Round 07 明确是**定向邻域追踪**,不能称为盲发现;group 6 / 7
也不能结果后替换成更好看的范围。
## 3. 为什么需要两个方向
只在 learned 背景把 group 6 / 7 改成 uniform,回答的是:
> 只改这段是否足以复现全局干预的一大部分下降?
但 mixer 路径有串联、分流和 nonlinear interaction,单侧结果可能被其他 learned
路径补偿。反过来,在 all-uniform 背景只把 group 6 / 7 恢复为 learned,回答的是:
> 只恢复这段是否足以让尖峰朝 reference 回升?
两种值都不是“贡献百分比”,也不要求相加为 100%。Round 07 用相同的 global log gap
归一化两种方向,只把双向、跨 seed 稳定的结果称为 localization evidence。
## 4. 冻结候选范围
从 `detached_learned` 背景出发的 sufficiency scopes:
1. group 6;
2. group 7;
3. groups 6+7;
4. group 6 attention-only;
5. group 6 MLP-only;
6. group 7 attention-only;
7. group 7 MLP-only;
8. output-only;
9. all 64 depth mixers;
10. all 65 mixers。
从 all-uniform 背景出发的 restoration scopes:
1. restore group 6 to detached-learned value coefficients;
2. restore group 7 to detached-learned value coefficients;
3. restore groups 6+7 to detached-learned value coefficients。
加上 `detached_learned` reference,共 14 种模式。预期 uniform selector census 为:
| mode | uniform mixers |
|---|---:|
| detached reference | 0 |
| group branch only | 4 |
| group only | 8 |
| groups 6+7 | 16 |
| output only | 1 |
| all depth | 64 |
| all | 65 |
| uniform except group 6 / 7 | 57 |
| uniform except groups 6+7 | 49 |
每次 diagnostic 都必须保存实际选中的 mixer identity;不能只信 mode 名称。
group 7 的 scope 包含 layers 26–28,它们不在固定尖峰集合 `S=21–25` 中,所以
group 7 是预先定义的**完整邻接 group intervention**,不是 spike-layer-only
intervention。
## 5. 不把局部 intervention 误译成什么
即使 groups 6+7 双向通过,结论也只限于:
- 固定训练状态;
- 固定 diagnostic batch 与 loss;
- 固定 `S = 21–25` 指标;
- 同前向、替代 source-gradient coefficient 的 diagnostic backward。
它不等于:
- 这些层“产生”了尖峰;
- group 6 / 7 是唯一原因;
- 真实 K3 checkpoint 有相同梯度路径;
- 把 mixer 训练成 uniform 会有同样结果;
- 局部 effect 可加,或可解释成方差分解;
- 论文 Figure 5(c) 的未公开 telemetry 已被复现。
保留 output-only 和 all-depth 两个控制,是为了看清最终 readout 与 depth path 的关系;
它们不进入 group 6 / 7 localization 的主判定。
## 6. 同期 artifact 状态审计:`A_log`
这一问题与缩小实验的局部梯度机制**相互独立**,但会限制任何真实 K3 checkpoint
验证,因此在冻结 Round 07 前重新检查官方模型仓库。
截至 **2026-07-30 12:35 CST**:
- 官方 Hugging Face main commit 仍为
`9f62e4e9fffbd0a83ddd60e1c209d828994b3569`;
- main 的 `modeling_kimi_linear.py` 仍以 `num_heads=96` 初始化 `A_log`;
- 已发布 checkpoint 中该张量的公开 shape 是 `[128]`,与 main 存在加载不匹配;
- 官方 main 尚未合并修复或给出 conversion contract。
同时出现了两个**未合并、互相竞争的社区 PR**:
### PR #144:把参数改成 128
- 一行把初始化从 `self.num_heads` 改为 `self.head_dim`;
- 提交者报告所有 shards 能加载;
- 提交者明确说没有独立验证 forward;
- 它把 checkpoint shape 当作权威语义。
### PR #150:保留 96,加载时验证并裁零尾
- 保持模型参数为 `[num_heads]=[96]`;
- `_load_from_state_dict` 检查 `[96:128]` 全为零后再裁掉;
- 提交者报告检查了 69 个 KDA 层,所有 32 项尾部都 exact zero;
- 提交者还报告经过 disk-offloaded MoE 的完整生成;
- 这些 checkpoint 全量扫描和生成是**提交者报告**,本项目没有下载约 1.56 TB
权重独立复核;本项目只核对了 PR diff、main 代码路径和 PR 状态。
PR #150 进一步指出,forward 中 `v` 被 reshape 为 96 heads,kernel 随后接收
`A_log`;若直接采用 #144 的 128 元素参数,现有 `view(H, 1)` 路径会在 96 heads 下
失败。这个论证比单看 checkpoint shape 更完整,但在官方合并或独立复核前,仍必须标成
高可信社区解释,而不是 Kimi 官方结论。
当前准确状态应写成:
> official main 仍然不匹配;社区已有两个竞争性候选修复,其中 #150 提供了更完整的
> checkpoint-tail 与 end-to-end 证据,但尚无官方裁决。
来源:
- [Kimi-K3 official main](https://huggingface.co/moonshotai/Kimi-K3/tree/main)
- [main `modeling_kimi_linear.py`](https://huggingface.co/moonshotai/Kimi-K3/blob/main/modeling_kimi_linear.py)
- [community PR #144](https://huggingface.co/moonshotai/Kimi-K3/discussions/144)
- [community PR #150](https://huggingface.co/moonshotai/Kimi-K3/discussions/150)
## 7. Round 07 的可证伪问题
Round 07 将:
1. exact replay Round 06 的三 seed 训练;
2. exact reproduce Round 06 的 `detached_learned` 与 `uniform_all` 两个端点;
3. 在 step 0 对 14 种模式做负控制,在 step 8,000 做正式矩阵;
4. 同时测 `spike_contrast` 与 `peak_normalized`;
5. 用 groups 6+7 的 sufficiency 与 restoration 两个方向预注册 50% log-gap
localization threshold;
6. 用单 group 的 20% 阈值和 attention-vs-MLP 的 15 percentage-point margin
作更细分的层级判定;
7. 从初始化完整 replay seed 2026073001。
完整模式、公式、失败规则和复现合同见
`research/K3_ATTNRES_LOCAL_PATH_PROTOCOL.md`。