7.4 KiB
K3 Attention Residuals 训练期前向干预:Round 08 前置定位
研究日期:2026-07-30
阶段身份:Round 07 后的定向 scoping,不是 Round 08 结果
上游协议:llm-atlas-k3-attnres-local-path-v1
1. 为什么还需要一次训练期实验
Round 06 / 07 都保持 learned forward 完全不变,只在固定 diagnostic 的 backward 中 替换 source value coefficients。它们回答的是:
- 全部 65 个 mixer 的 uniform value backward 能不能压低固定尖峰;
- group 6 / 7 的 16 个 depth mixer 在 learned 背景上是否足以复现全局下降;
- 从 all-uniform 背景恢复这些 mixer 是否能反向恢复尖峰。
Round 07 的正式结论是:
groups 6+7 sufficiency:6 / 6 seed×metric cells 通过
groups 6+7 restoration:3 / 6 cells 通过
formal status:one_sided_evidence_localization_not_established
这已经足以排除“局部 mask 完全没有反应”,但还不能回答:
如果训练的每一次 forward 都真的把这段 depth routing 改成算术平均,模型会怎样适应?
Round 08 把 intervention 放进 optimizer path。它不再追求 forward-identical,而是让 选中 mixer 的输出在训练、验证与诊断中始终为所有 source states 的等权平均。
2. 这不是“只改变 forward”
选中 mixer 的 learned 路径原本是:
keys = RMSNorm(sources)
logits = query · keys
weights = softmax(logits over source-depth)
output = Σ weights_i × source_i
Round 08 的选中路径是:
output = (1 / N) × Σ source_i
因此 intervention 同时改变:
- forward 的 branch input;
- 由新 forward 自然产生的 source gradients;
- 下游 activation、loss 与所有后续 optimizer updates;
- 选中 mixer 的参数可达性:
query与key_norm.weight不参与图,不得到梯度。
这是一项训练期架构消融,不是“只改变 forward、不改变 backward”的可分离因果实验。 结果不能被翻译成 query/key 路径的纯因果效应。
3. 为什么只选四个新变体
固定 depth-32 / Block AttnRes 拓扑:
| scope | layers | 0-based depth mixer indices | count |
|---|---|---|---|
| group 6 | 21–24 | 40–47 | 8 |
| group 7 | 25–28 | 48–55 | 8 |
| groups 6+7 | 21–28 | 40–55 | 16 |
| group 7 MLP | 25–28 | 49, 51, 53, 55 | 4 |
四个新训练变体固定为:
uniform_group_6_forwarduniform_group_7_forwarduniform_groups_6_7_forward(主变体)uniform_group_7_mlp_forward
选择依据不是 Round 08 结果:
- joint 6+7 是 Round 07 的固定主 scope;
- 单 group 6 / 7 用来构成交互图;
- group 7 MLP-only 是 Round 07 的 branch-level 次级线索;
- output mixer 保持 learned,避免把 local depth intervention 扩成全局 readout 改写。
不加入 attention-only、output-only、all-depth 或 all-65,是为了把本轮正式算力集中在 “局部训练适应”而不是重新展开 14-mode backward census。
4. 冻结参考,而不是重新挑基线
paired learned reference 复用 Round 05 已公开的三个:
formal / depth-32 / block / seeds 2026073001..3
它们与新变体使用相同:
- 初始化过程;
- 8,000-step deterministic byte schedule;
- 64 validation windows;
- 16 diagnostic windows;
- optimizer、LR schedule、BF16 forward 和 FP32 residual accumulation;
- 六个 diagnostic checkpoints。
runner 的 learned_reference smoke mode 必须逐字段复现父 runner 数值。正式 analyzer
还必须验证新旧 cell 的 initial public/mixer hashes、输入 gate tensor hashes、
validation tensor hash、diagnostic tensor hash 与 deterministic / autocast 数值合同。
GPU 名称、驱动 minor 与 timing 只记录,不作为配对 hard gate。
复用参考的意义是避免为已冻结、可复现的三格再消耗 196,608,000 target bytes。报告必须 把它写成历史配对参考,不能写成同期随机对照。
5. 主问题与质量护栏
主对象沿用 Round 05–07,并且主判定只读取最终 step 8,000:
step = 8,000
position = post-MLP Transformer-block output
reduction = activation-gradient element RMS
S = layers 21–25
R = other 27 layers
C = mean(g[S]) / mean(g[R])
P = max(g) / mean(g)
对变体 v 和同 seed learned reference:
D_X(v) = (X_ref - X_v) / X_ref
主变体 uniform_groups_6_7_forward 只有在:
C与P的D_X ≥ 0.20;- 三个 seed 的六个 cell 全部通过;
- 每个 seed 的 final validation BPC 相对 reference 不劣化超过
+0.05; - 三 seed mean BPC delta 不超过
+0.03;
时才得到 “forward-training attenuation established within reduced protocol”。
质量阈值在结果前固定。它不是“能力保持”的通用标准,只用于阻止把明显训练失败造成的 梯度塌缩计成正结果。即使通过,也不能推出下游能力等价。
6. 非加性局部交互图
对每个 seed、checkpoint 和指标:
E6 = ln(X_ref / X_group6)
E7 = ln(X_ref / X_group7)
E67 = ln(X_ref / X_groups6+7)
I67 = E67 - E6 - E7
I67 > 0:joint log attenuation 大于两个 single effects 的和;I67 < 0:joint log attenuation 小于两个 single effects 的和;I67 = 0:只是在这个定义下恰好 log-additive。
I67 没有预注册显著性阈值,不是 Shapley value、方差分解、独立性检验或因果交互估计。
三个 effect 来自三套独立训练,它只是跨 run 的 log-attenuation residual。它的用途是把
训练轨迹中的补偿/放大关系画清楚,而不是制造一个新的“通过/失败”结论。
groups 6+7 覆盖 layers 21–28,而固定尖峰窗只到 layer 25;layers 26–28 落在 R。
所以 C 的变化可能同时来自 S 下降与 R 上升。正式结果必须把两者拆开报告,不能把
contrast 下降单独翻译成“尖峰层被关闭”。
7. 真实 K3 checkpoint 的同期边界
截至本轮预检,官方 Kimi-K3 Hugging Face main 仍停在 revision
9f62e4e9fffbd0a83ddd60e1c209d828994b3569,remote code 仍按 96 heads 初始化
A_log,公开 checkpoint header 仍为 [128]。社区 PR #144 / #150 仍是两个未合并、
语义不同的候选修复;没有官方裁决。
所以本轮不下载约 1.56 TB 权重,不声称对真实 K3 forward 做了验证。缩小实验只继承 Block AttnRes 的拓扑动机,不是 K3 checkpoint 的数值替身。
一手状态页:
8. 本轮可证伪交付
Round 08 将在查看正式结果前完成:
- 冻结协议与 machine-readable manifest;
- 实现一个 selector,而不是四份分叉 forward;
- 通过 empty-selector 父等价、step-0 uniform identity、selector census、参数不可达性、 loss-scale 与输入 hash 闸门;
- 运行 4 variants × 3 seeds × 8,000 steps;
- 从初始化 replay 主变体 seed 2026073001;
- 由单一 analyzer 生成主判定、质量闸门、轨迹与非加性交互;
- 独立审阅机器可读结果;
- 以五视图交互实验接入网站、开源并发布。