Files
llm-atlas/research/K3_ATTNRES_LOCAL_PATH_PROTOCOL.md

19 KiB
Raw Permalink Blame History

K3 Attention Residuals 局部 mixer 路径干预协议

协议 ID:llm-atlas-k3-attnres-local-path-v1
冻结日期:2026-07-30
协议状态:结果前预注册 frozen;任何语义变更必须更换 protocol ID
父协议:llm-atlas-k3-attnres-spike-path-v1

0. 研究身份

本轮是 Round 06 结果后的定向机制追踪,不是盲发现。

已知:

  • depth-32 / Block 的固定尖峰集合是 layers 21–25;
  • 全部 65 个 mixer 的 source value backward coefficients 从 learned 改成 uniform, 在三 seed 上 material 地降低最终 post_mlp_state 的 spike contrast 和 peak;
  • 该全局干预保持 forward exact;
  • group 6 覆盖 layers 21–24,group 7 覆盖 layers 25–28。

未知:

  • 全局下降是否主要集中在 group 6 / 7 的 16 个 depth mixers;
  • group 6 与 group 7 是否各自有稳定 effect;
  • attention 与 MLP mixer 是否能在预注册阈值下区分;
  • output mixer 是否解释了大量全局 effect;
  • 局部 sufficiency 与反向 restoration 是否给出一致证据。

前置证据、拓扑和同期 artifact audit 固定在 research/K3_ATTNRES_LOCAL_PATH_SCOPING.md。任何结果不得倒写成事前未知。

1. 允许回答的问题

  1. 在相同模型、batch、loss、activation 与 learned forward weights 下,只改变某个固定 mixer scope 的 source-gradient coefficients,能复现多少全局 log gap?
  2. 从 all-uniform 背景只恢复 group 6 / 7 的 learned coefficients,能恢复多少 global log gap?
  3. groups 6+7 是否在 sufficiency 与 restoration 两个方向、两个 spike 指标、三 seed 同时通过 50% 阈值?
  4. 单独 group 6 或 group 7 是否在两个指标、三 seed 通过 20% 阈值?
  5. attention-only 与 MLP-only 是否达到预注册的 branch dominance 规则?
  6. output-only 与 all-depth 控制是否显示 effect 主要来自最终 output mixer?

2. 明确不回答的问题

  • Kimi K3 2.8T checkpoint 的真实训练梯度;
  • 论文 Figure 5(c) 未公开 telemetry 的精确定义;
  • 哪个 layer、source 或 operator “产生”尖峰;
  • learned source weight 的语义归因;
  • 重新训练局部 uniform variant 的最终能力;
  • intervention effect 的可加性、Shapley value 或方差分解;
  • 三 seed 外的总体显著性、置信区间或 p-value;
  • K3 A_log 两个社区修复中哪个已经得到官方认可;
  • checkpoint conversion、推理正确性或部署可用性。

3. 冻结训练与数据合同

完整复用 Round 06:

字段 固定值
architecture Block AttnRes
Transformer depth 32
aggregation groups 8
blocks / group 4
depth mixers / output mixers 64 / 1
width / heads / FFN 192 / 6 / 768
context / vocabulary 256 / byte-256
seeds 2026073001 / 2026073002 / 2026073003
steps / batch 8,000 / 32
target bytes / formal cell 65,536,000
optimizer AdamW
peak / min LR 3e-4 / 3e-5
warmup 400
weight decay 0.1 for ndim ≥ 2
betas / epsilon 0.9, 0.95 / 1e-8
clip global norm 1.0
forward CUDA BF16 autocast
residual accumulation explicit FP32
diagnostic CE fixed 16 × 256 token-mean FP32 CE
parent diagnostic steps 0 / 100 / 500 / 2,000 / 4,000 / 8,000
local matrix steps 0 / 8,000

训练路径必须逐调用父 runner 原始 forward;局部 custom autograd 只能存在于 optimizer step 外的 diagnostic。数据 bytes、schedule、validation tensor、diagnostic tensor 与三 seed 的每个 optimizer input 都必须与 Round 06 exact。

诊断调度明确分成两条:

  • 六个 parent diagnostic steps 都运行原始 learned mode,用于 Round 06 等价;
  • step 0 / 8,000 另外运行下述 14-mode local matrix;
  • localization 公式只读取 14-mode matrix,绝不把 full-autograd learned 混入分母。

4. 正式矩阵与 replay

正式运行:

depth-32 / block / seed-2026073001
depth-32 / block / seed-2026073002
depth-32 / block / seed-2026073003

另从初始化完整重跑:

replay / depth-32 / block / seed-2026073001

正式三格处理 196,608,000 target bytes;含 replay 共 262,144,000 bytes。每格使用 全新 Python 进程。最多并行两个进程;不能共享 model、optimizer、RNG 或 CUDA graph。 wall-time 不进入数值复现合同。

5. 固定主对象与指标

主对象固定为最终 step 8,000:

position  = post_mlp_state
reduction = element_rms
S         = layers 21, 22, 23, 24, 25
R         = other 27 layers

对 mode m:

C_m = mean(metric[S]) / mean(metric[R])   # spike contrast
P_m = max(metric) / mean(metric)          # peak normalized

C_m 与 P_m 必须 finite 且严格大于 1e-30。后文统一用 X_ref 表示 X_detached_learned;不用 detached_reference 等其他别名。不允许用其他位置、reduction、layer 集合或 metric 替换主对象。32-layer raw spectrum、normalized spectrum、peak layer 和 top-five layers 全量报告,但不参与主阈值。

6. 14 种冻结模式

所有模式调用同一个 parent learned forward。custom Function 的 forward 直接返回 parent output,只有 backward 对 source tensors 使用选定 coefficients。

14-mode matrix 全部使用 RoutedSourceBackward:每个 mixer 只在 stopgrad(w) 与 1/N 两种 source value coefficients 中选择。完整 query / key / softmax autograd 的 learned 不属于这 14 种模式,只用于训练和父诊断 等价。实现不得把某个 restoration scope 切回 full-autograd learned。

6.1 reference

detached_learned

  • 65 个 mixer 全部使用 learned w 作为 source value backward coefficients;
  • w detach,不走 query / key / softmax derivative path;
  • 必须 exact reproduce Round 06 的同名 mode。

6.2 learned 背景上的局部 uniform:sufficiency family

未选 mixer 使用 detached learned coefficients;选中 mixer 使用 1/N:

  1. uniform_group_6_only
  2. uniform_group_7_only
  3. uniform_groups_6_7_only
  4. uniform_group_6_attention_only
  5. uniform_group_6_mlp_only
  6. uniform_group_7_attention_only
  7. uniform_group_7_mlp_only
  8. uniform_output_only
  9. uniform_depth_all
  10. uniform_all

uniform_all 必须 exact reproduce Round 06 的 uniform_value_backward。

6.3 all-uniform 背景上的 detached-learned restoration family

选中 scope 恢复 detached learned coefficients,其余保持 uniform:

  1. uniform_except_group_6
  2. uniform_except_group_7
  3. uniform_except_groups_6_7

名字中的 except 表示该 scope 不是 uniform。报告和网站必须同时展示人话标签 “restore ... to detached-learned value coefficients”,避免误读。

7. selector 的唯一合同

为了同时满足新 selector audit 和 Round 06 endpoint exactness,保留两个互不混写的 schema:

  1. 父 trace.mixers summary 逐字段不变;output 仍使用父 schema 的 mixer_index=null,并继续参与父 mixer_summary_sha256;
  2. 新增平行 selector_visits,只用于 local mask audit,不写入父 summary。

selector_visits 的 depth mixer identity 用:

(kind="depth", index=0..63, layer=1..32,
 group=1..8, branch in {"attention","mlp"})

selector_visits 的 output mixer identity 用:

(kind="output", index=64, layer=null, group=null, branch="output")

这里 index=64 只是新 selector schema 的稳定别名,不得回写父 summary。

令 D_i 表示 kind=depth,index=i,O 表示 output。14 种 mode 的 uniform identity 集合冻结如下:

mode exact uniform set
detached_learned ∅
uniform_group_6_only {D40,…,D47}
uniform_group_7_only {D48,…,D55}
uniform_groups_6_7_only {D40,…,D55}
uniform_group_6_attention_only {D40,D42,D44,D46}
uniform_group_6_mlp_only {D41,D43,D45,D47}
uniform_group_7_attention_only {D48,D50,D52,D54}
uniform_group_7_mlp_only {D49,D51,D53,D55}
uniform_output_only {O}
uniform_depth_all {D0,…,D63}
uniform_all {D0,…,D63,O}
uniform_except_group_6 {D0,…,D39,D48,…,D63,O}
uniform_except_group_7 {D0,…,D47,D56,…,D63,O}
uniform_except_groups_6_7 {D0,…,D39,D56,…,D63,O}

runner 必须把这些 set 编码为一个 frozen selector 函数;不能散落在 mode-specific if/else 中。manifest 同时保存 machine-readable exact index lists。runner 通过 override _mix 或等价 hook 做 set lookup,并替换父 runner 中只接受三种 global mode 的 mode validation、bundle loop 和相关 gate;训练 forward 继续直接调用父路径。

每次 forward 必须验证:

  • exactly 65 个 mixer visits;
  • identity 不重复;
  • identity 顺序与 reference exact;
  • 父 trace.mixers schema 与 hash 路径没有新字段;
  • uniform census 与 manifest exact;
  • selected identity list 与 selector rule exact;
  • reference 的 uniform count 为 0;
  • branch-only 4,group-only 8,groups 6+7 为 16;
  • output-only 1,all-depth 64,all 65;
  • except-one-group 57,except-two-groups 49。

任一 gate 失败,cell invalid;不得只改结果 JSON。

8. forward identity 与 parent exactness

8.1 所有 14 模式的 forward identity

同 seed / step 相对 detached_learned 必须满足:

  • logits tensor SHA-256 exact;
  • loss FP32 value exact;
  • 六位置 activation tensor hashes exact;
  • 65 个 mixer forward summaries exact。

任一 mode 失败,整格 invalid。

8.2 Round 06 endpoint exactness

对 step 0 / 8,000:

  • detached_learned 的 logits、loss、六位置 activation、未改 schema 的 mixer summaries 和六位置 gradient reductions 必须与对应 Round 06 raw output exact;
  • uniform_all 的同一组字段和六位置 gradient reductions必须与对应 Round 06 uniform_value_backward exact;
  • 正式训练的 final model hash、optimizer hash、六个 validation BPC、training history 与六个 parent learned diagnostics 必须与 Round 06 exact;
  • 新 selector_visits 不参与旧 mixer_summary_sha256,而由独立 canonical hash 和 exact-set gate 管理。

runner / protocol / scoping / manifest 物理 hash 在运行前冻结。父 raw 文件同时检查 physical SHA-256、canonical SHA-256、final model hash 和 final optimizer hash。

9. 初始化负控制

step 0 的 mixer query 为零,learned w 是 uniform。14 模式在六个位置的 element_rms 必须:

  • 32 个 raw values 全部 finite、strictly positive;
  • 相对 reference 的逐层 raw relative error ≤1e-6;
  • normalized absolute error ≤1e-6。

此外:

  • detached_learned 与 uniform_all 必须分别与 Round 06 step-0 endpoint exact;
  • parent full-autograd learned 与 detached_learned 必须按 Round 06 负控制在 1e-6 tolerance 内一致;
  • detached_learned 另执行同一 loss 的 ×1 / ×2 backward,六位置、七 reductions 都必须通过父协议相同的 scale 与 normalized-spectrum gate。

失败表示 selector 或 surrogate 没有隔离预期路径;正式结果无效。

10. global log gap

对每个 seed 和每个指标 X ∈ {C,P}:

G_X = ln(X_ref / X_uniform_all)
relative_drop_X = (X_ref - X_uniform_all) / X_ref

只有同时满足以下条件才允许解释局部比例:

  1. G_C > 0 且 G_P > 0;
  2. 上式 relative_drop_X ≥0.20;
  3. Round 06 endpoint exactness 通过。

seed s 的 metric X 任一条件不满足,则该 (s,X) 称为 global gap not established,不计算该格 S_X / R_X。groups 6+7 的主 gate 要求 3 seed × 2 metrics 全部存在,因此任一 required cell 缺失都会使主 localization 判定失败;仍公开 raw matrix,不使用事后替代分母。

log ratio 用于让相同的乘法变化在两个方向可比。所有归一化值按原值报告,不裁剪到 [0,1];负值表示反方向,超过 1 表示局部 intervention 超过 all-uniform endpoint。

11. sufficiency score

对 sufficiency mode m:

S_X(m) = ln(X_ref / X_m) / G_X

11.1 groups 6+7 主判定

只有 uniform_groups_6_7_only 对 C 和 P 都满足:

S_X(m) ≥ 0.50

且三个 formal seed 6 / 6 全部达标,才记为:

groups 6+7 的 16 个 depth mixers 在本 diagnostic 中,足以复现至少一半 all-65 uniform intervention 的预注册 log-gap reduction。

任一失败记为 not sufficient at the preregistered 50% threshold。mixed 精确定义为: seed 通过/失败不一致、C/P 通过/失败不一致,或 score 的正负号跨 seed 不一致;可同时 附加多个原因,不得降低阈值。

11.2 单 group

group 6 / group 7 分别对 C 和 P、三 seed 全部满足:

S_X(m) ≥ 0.20

才称为 material local sufficiency at the 20% threshold。没过阈值不等于 effect 为零。

12. restoration score

对 restoration mode r:

R_X(r) = ln(X_r / X_uniform_all) / G_X

12.1 groups 6+7 主判定

只有 uniform_except_groups_6_7 对 C 和 P、三 seed全部满足:

R_X(r) ≥ 0.50

才称为:

从 all-uniform 背景只恢复 groups 6+7 的 learned coefficients,恢复了至少一半 预注册 global log gap。

这仍是同前向 backward-rule restoration sensitivity,不是严格 causal necessity。

12.2 单 group

uniform_except_group_6 / uniform_except_group_7 分别以 ≥0.20、两个指标、三 seed 作为 material restoration threshold。

13. localization 总闸门

只有以下两项同时通过:

  1. groups 6+7 sufficiency:S_C,S_P ≥0.50,3 / 3 seeds;
  2. groups 6+7 restoration:R_C,R_P ≥0.50,3 / 3 seeds;

才允许写:

在本缩小模型、固定训练状态和 diagnostic backward 下,全局 value-coefficient sensitivity 的主要部分 localization 到 group 6 / 7 mixer path。

即使通过,也必须紧邻注明:

  • “主要部分”由 50% 双向阈值定义;
  • effect non-additive;
  • 不是唯一来源或 layer-origin;
  • 不是真实 K3 checkpoint 结论。

一侧通过一侧失败,统一写成 one-sided evidence, localization not established。

14. attention vs MLP branch 判定

每个 group 独立比较 attention-only 与 MLP-only sufficiency score。只有某 branch:

  1. S_C ≥0.20 且 S_P ≥0.20;
  2. 在 C 与 P 上都比 sibling 高至少 0.15;
  3. 三 seed 全部满足前两项;

才称为 branch-dominant at the preregistered margin。

若 group-level sufficiency 未通过 20% 阈值,不允许宣称其内部 branch dominance。 branch-only scores 可能交互、超加或相互抵消,不能相加成 group score。 本节只有 sufficiency 方向,没有 branch-level restoration,属于预注册的次级探索性 判定,证据层级低于 §13 双向 localization。

15. output 与 depth 控制

uniform_output_only 和 uniform_depth_all 不进入 group localization 主判定。

探索性报告:

S_X(output)
S_X(depth_all)
interaction_residual_X =
  1 - S_X(output) - S_X(depth_all)

interaction_residual 只是 log-gap bookkeeping,不是统计交互估计或贡献分解。 它不预期接近 0,也不是 hypothesis test。

只有 output-only 对两个指标、三 seed 都 ≥0.50,才标记 output mixer alone captures at least half the global gap。即使如此,也不否定 groups 6+7;两者可能重叠、串联或超加。

16. 报告顺序与反 cherry-picking

固定报告顺序:

  1. input / parent / endpoint exactness;
  2. step-0 negative control;
  3. 每 seed 的 raw C / P 矩阵;
  4. global gaps;
  5. groups 6+7 sufficiency;
  6. groups 6+7 restoration;
  7. localization gate;
  8. single-group scores;
  9. branch scores;
  10. output / depth controls;
  11. full 32-layer spectra;
  12. replay;
  13. limitations。

所有 14 modes、两个指标、三个 seed 都公开。不得只展示通过阈值的 scope。不得用跨 seed 均值替代 3 / 3 gate;均值只用于视觉摘要。

17. replay 与复现闸门

seed 2026073001 从初始化独立 replay,比较去除以下字段后的 canonical content:

  • run_kind;
  • wall-clock timing;
  • output path;
  • self canonical hash。

至少以下字段必须 exact:

  • input tensor hashes;
  • initial/final model 与 optimizer hashes;
  • evaluations / training history;
  • parent diagnostics;
  • 14-mode step-0 / step-8,000 forward hashes;
  • selector census / identities;
  • 六位置 raw gradient reductions;
  • global gaps / local scores / gates。

若正式 seed1 与 replay 不 exact,Round 07 数值结论无效。 replay 固定在与 formal 相同 host、GPU、Python、PyTorch、CUDA 和 CUBLAS_WORKSPACE_CONFIG 环境;本协议不声称跨硬件 bit exact。

18. 预期失败与停止规则

以下任一项使 cell invalid:

  • CUDA deterministic contract 未开启;
  • parent manifest / runner / protocol / scoping / raw hash 不匹配;
  • 训练等价失败;
  • diagnostic 改变 optimizer state;
  • forward identity 失败;
  • selector identity / census 失败;
  • Round 06 endpoint exactness 失败;
  • step-0 negative control 失败;
  • raw gradient missing、non-finite 或 non-positive;
  • global gap denominator 不成立。

程序错误修复必须:

  1. 保存失败日志;
  2. 修改 runner;
  3. 更新 runner hash;
  4. 明确判断协议语义是否改变;
  5. 若改变 selector、mode、metric、threshold 或 aggregation,创建新 protocol ID;
  6. 全部受影响 cell 从初始化重跑。

19. 结果语言边界

允许:

  • “在同前向 diagnostic backward 下,uniformizing scope X 改变了固定尖峰指标”;
  • “groups 6+7 在预注册 50% 双向阈值下建立 / 未建立 localization”;
  • “branch effect mixed / below threshold”;
  • “这是一项 reduced-model mechanism probe”。

禁止:

  • “证明 K3 的尖峰来自第 6 组”;
  • “这些 mixer 贡献了 X% 梯度”;
  • “group effect 加总为 100%”;
  • “uniform mixer 更适合训练”;
  • “复现了 Figure 5(c)”;
  • “验证了 K3 2.8T checkpoint”;
  • “社区 PR #144 或 #150 已成为官方修复”。

20. 冻结清单

在任何 formal 结果产生前必须完成:

  • scoping 文件完成;
  • protocol 状态改为 frozen;
  • 14 modes 与 selector census 写入 manifest;
  • thresholds / formulas 写入 manifest;
  • Round 06 父 artifact physical / canonical hashes 写入 manifest;
  • runner、protocol、scoping、manifest hashes 固定;
  • Grok Headless 对抗审阅完成,采纳/拒绝理由留档;
  • step-0 smoke 全门通过;
  • formal 命令与环境写入 README;
  • 单一 analyze.py 实现所有 score / gate,网站只消费其冻结输出;
  • protocol commit 早于 formal result commit。