Compare commits
4 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| dc8ec30f78 | |||
| 2dbfd7944a | |||
| 39a9ad6215 | |||
| 6911efc6e7 |
+13
-4
@@ -8,7 +8,7 @@
|
||||
|---|---:|---:|---|
|
||||
| 研究框架与规范 | 进行中 | 83% | Scaling Laws 二轮拟合复现与逐图精读 |
|
||||
| 网站设计系统 | 进行中 | 89% | 打印样式与更多通用可视化组件 |
|
||||
| Kimi K3 深读 | 六轮实证进行中 | 99% | 对 group 6 / 7 做局部 mixer backward 干预,并等待 `A_log` 官方转换合同 |
|
||||
| Kimi K3 深读 | 七轮实证进行中 | 99% | 设计前向训练变体,并等待 `A_log` 社区候选的官方裁决 |
|
||||
| 语言模型前史 | 完成首版 | 78% | Kneser–Ney、LSTM、Bahdanau 逐图精读与真实小语料复现 |
|
||||
| Transformer 基础 | 完成首版 | 79% | 多头电路、归一化 traces 与真实 kernel / KV 配置 |
|
||||
| 表示、位置与残差高速公路 | 完成首版 | 81% | 真实 hidden-state / norm traces、长上下文位置外推与深层稳定性消融 |
|
||||
@@ -41,7 +41,7 @@
|
||||
- [x] 完成 486 篇关键论文索引,覆盖 16 个标签专题与 Kimi/DeepSeek 聚光主线。
|
||||
- [x] 完成可检索、可按专题筛选的论文库页面。
|
||||
- [x] 完成 K3、语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全十七篇首版长文。
|
||||
- [x] 完成 K3 三轴架构、八联报告实验、四联开放工件实验、Round 04 / 05 / 06 各五联 AttnRes 独立实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 二十二联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等一百零四个原创交互视图。
|
||||
- [x] 完成 K3 三轴架构、八联报告实验、四联开放工件实验、Round 04 / 05 / 06 / 07 各五联 AttnRes 独立实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 二十二联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等一百零九个原创交互视图。
|
||||
- [x] 完成长上下文首版:五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。
|
||||
- [x] 核验 FlashAttention、DeepSeek-V2/V3.2/V4、Kimi Linear/K3 等六份论文原文,并建立长上下文研究账本。
|
||||
- [x] 核验 Switch、ST-MoE、DeepSeekMoE、Loss-Free、V3、LatentMoE 与 K3 原文,并建立 MoE 研究账本。
|
||||
@@ -298,10 +298,16 @@
|
||||
- [x] K3 Round 06 五视图实验室完成:训练轨迹、六位置谱、12 格 reduction 稳健性、same-forward 三规则干预与 mixer 散点/证据阶梯可交互;protocol、scoping、audit、runner、analyzer、raw/aggregate/compact/reproduction 全部进入公开树。
|
||||
- [x] Round 06 本地闸门通过:97 个 Astro 文件零诊断,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;Round 04/05/06 三套冻结数据、三套专项、K3 全量与全站 22 套真实 Chrome 回归通过,桌面/390px 移动端零文档级溢出、零 offender、零运行时异常;首页首发卡与陈旧跨页断言同步到 Round 06。
|
||||
- [x] K3 Round 06 以运行源提交 `b393780`、不可变镜像 `20260730T042651Z-b393780` 发布;OCI index digest `sha256:4e4cb2e065526f50c10cf759bab80a5a871177812ca9fa5c9da77c662c81a63a`,NAS healthy / 0 次重启、Compose Manager 标签、VPS→NAS、NPM host 31 / cert 41、DNS、HTTPS/2、首页/K3 公网内容、门户 `LLM ATLAS / projects / 180` 与全站 22 套生产 Chrome 回归全部通过;保留 Round 05 `20260730T022610Z-f7670ef` 与中间 `20260730T042110Z-a1a52d4` 回滚点。
|
||||
- [x] K3 Round 07 在正式输出前冻结 `llm-atlas-k3-attnres-local-path-v1`:固定 Round 06 三个 depth-32 Block 正式格、layer 21–25 spike set、groups 6+7、14 个 exact selector masks、sufficiency / restoration 双向 log-gap score、主门 50%、单组/分支/输出控制与完整 seed-1 replay;Grok 只读对抗审阅提出的六个 blocking protocol 问题全部在冻结前修正。
|
||||
- [x] 三个正式格与 replay 共消费 262,144,000 target bytes;Round 06 parent equivalence、forward identity、65-node census、selector identity / order / uniqueness、负对照与 loss-scale 闸门全部通过。正式最终 BPC 为 `1.7123525941 / 1.7093240656 / 1.7030966813`。
|
||||
- [x] global gap 在两个指标 × 三 seed 的 6 / 6 格通过。groups 6+7 sufficiency 的 mean score 为 contrast `.677`、peak `1.700`,6 / 6 ≥ `.50`;restoration mean 为 contrast `.650`、peak `.380`,contrast 3 / 3 通过而 peak 0 / 3 通过,正式状态固定为 `one_sided_evidence_localization_not_established`。
|
||||
- [x] 次级控制显示 group 7 MLP-only branch 在 sufficiency 方向 6 / 6 通过 material + margin 门;group 6 MLP-only 为 5 / 6,不能宣布 dominance。output-only mean sufficiency 仅 `.131 / .187`、0 / 6 过 50%;all-depth 为 `.949 / 1.145`。所有 score 都是非加性 log-gap 诊断,不写成贡献率。
|
||||
- [x] K3 Round 07 五视图实验室完成:65-node 路径图、14-mask 全矩阵、双向主门、branch/output controls 与 32 层原始谱/replay 审计可交互;protocol、scoping、Grok 结果前审阅、audit、runner、analyzer、packager、四个 raw JSON、aggregate、compact 与 reproduction 全部进入公开树。
|
||||
- [x] Round 07 本地闸门通过:100 个 Astro 文件零诊断,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;Round 04/05/06/07 四套冻结数据、四套 K3 专项、K3 全量与全站 23 套真实 Chrome 回归通过。动态生成矩阵的 scoped CSS 退化由截图复查发现并修复;桌面/390px 移动端零文档级溢出、零 offender、零运行时异常。
|
||||
|
||||
## 正在进行
|
||||
|
||||
- [ ] K3 六轮下一闸门:把全局 value-route sensitivity 收缩成 group 6 / group 7 的局部 mixer intervention matrix,并设计前向训练变体;真实 K3 forward 继续等待 `A_log [128]→[96]` 官方转换或权重修订。
|
||||
- [ ] K3 七轮下一闸门:设计前向训练变体与非加性局部交互地图;真实 K3 forward 继续等待 `A_log [128]↔[96]` 社区候选的官方裁决或权重修订。
|
||||
- [ ] DeepSeek 八轮下一闸门:推进干预式 mediation、SM90 FlashMLA、FP8 / pipeline traces 与 R1-like RL 小模型复现。
|
||||
- [ ] 表示、位置与残差二轮:真实 hidden-state / norm traces、长上下文位置外推复现与 mHC / AttnRes 深层稳定性消融。
|
||||
- [ ] 评测安全二轮:真实 cross-harness / pass@k 复跑、Judge 元评测、动态污染与过拒案例。
|
||||
@@ -490,7 +496,7 @@
|
||||
| 2026-07-30 | 输出长度揭示强任务域交互 | system-at-period 在 Code 为负、Math 为正,两个选定任务带都不跨零;不从长度外推能力 |
|
||||
| 2026-07-30 | Round 08 十二字段重放过闸 | 64/64 exact;uniform hash、完整 token IDs、文本、stop 与 RNG 一并进入复现合同,评分 gold 提前加载的流程偏差公开保留 |
|
||||
| 2026-07-30 | DeepSeek Round 08 任务 bootstrap / CRN 里程碑发布 | 源提交 `975ed3d`、镜像 `20260729T221654Z-975ed3d`、OCI `sha256:759a8446…21452b`;21/21 公网路由与生产专项/全量 Chrome 通过,保留 Round 07 回滚点 |
|
||||
| 2026-07-30 | K3 `A_log [128]` 冲突继续阻断真实 checkpoint forward | 当前 HF / GitHub / FlashKDA / vLLM / SGLang 均未公开 128→96 转换;不裁剪、不 reshape、不把假设输出冒充 K3 |
|
||||
| 2026-07-30 | K3 `A_log [128]` 冲突仍没有官方裁决 | 官方 main 仍期望 96;社区 #144 改为 128,#150 验证 tail zero 后裁为 96,两案都未合并。不把候选 patch 冒充官方 K3 forward |
|
||||
| 2026-07-30 | AttnRes 缩小实验先冻结再训练 | 三结构共享公共主干、初始化、窗口与优化器;三 seed paired BPC 只按预注册 `3/3 same direction + mean≤−.010` 判为本协议内方向支持 |
|
||||
| 2026-07-30 | 主结果与机制反结果同时发布 | Full / Block BPC 方向支持;核心参数 gradient RMS CV 却高于 Baseline,明确写成未复现论文梯度叙述 |
|
||||
| 2026-07-30 | 独立重放按数值合同而非计时合同验收 | Block / seed-1 的八组冻结字段 2,000 steps exact;wall time 受调度影响,不要求或声称 bit-exact |
|
||||
@@ -506,6 +512,9 @@
|
||||
| 2026-07-30 | value-route 降幅不写成因果贡献百分比 | 全局 uniform value-backward 让 contrast 平均下降 70.2%,只支持预注册阈值下的材料级敏感性;不声称 value 路径“解释了 70.2%” |
|
||||
| 2026-07-30 | reduction 稳健性限定在预注册家族 | element RMS、token RMS mean/median/P95 的 12/12 格通过;探索性 reduction 和其他 batch 不被纳入确认性外推 |
|
||||
| 2026-07-30 | K3 Round 06 尖峰路径里程碑发布 | 运行源 `b393780`、镜像 `20260730T042651Z-b393780`、OCI `sha256:4e4cb2e0…1a63a`;21/21 公网页面链路与全站 22 套生产 Chrome 通过,保留 Round 05 与中间 Round 06 回滚点 |
|
||||
| 2026-07-30 | 局部路径必须同时通过 sufficiency 与 restoration | groups 6+7 的 sufficiency 6/6 过 `.50`,restoration 仅 contrast 3/3 通过、peak 0/3 通过;强单侧证据不升级为 localization |
|
||||
| 2026-07-30 | 局部 score 不写成可加贡献率 | 同一 scope 在 learned 与 uniform 背景的响应不同,`S_peak > 1` 与负 interaction residual 都是非加性诊断,不是 170% 贡献或方差分解 |
|
||||
| 2026-07-30 | branch 与 output 控制保持次级证据身份 | group 7 MLP-only 的 6/6 只属于 sufficiency branch gate;group 6 为 5/6,output-only 为 0/6,都不能补救失败的双向主门 |
|
||||
|
||||
## 未决问题
|
||||
|
||||
|
||||
@@ -19,7 +19,7 @@
|
||||
|
||||
当前里程碑包含 17 专题学习地图、486 篇关键论文索引、Kimi K3 完整导读,
|
||||
语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 技术谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、工具使用与长程 Agent、原生多模态、训练系统、推理服务、数值优化,以及评测与安全深度专题,
|
||||
以及 104 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、
|
||||
以及 109 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、
|
||||
8 个交互实验和 100 个一手/官方节点,完整覆盖架构、预训练、后训练、系统、评测、案例与附录。
|
||||
第三轮已完成开放工件与首个真实 kernel 里程碑:固定官方模型与 FlashKDA revisions,审计 96 个 checkpoint shards、
|
||||
497,220 个 tensor entries、真实 KDA / MLA / MoE / MoonViT shapes 与小范围参数统计,并用 4 个新视图
|
||||
@@ -66,6 +66,18 @@ value-backward coefficients 换成均匀系数后,contrast 平均下降 `70.2%
|
||||
[K3_ATTNRES_SPIKE_PROTOCOL.md](./research/K3_ATTNRES_SPIKE_PROTOCOL.md)、
|
||||
[K3_ATTNRES_SPIKE_AUDIT.md](./research/K3_ATTNRES_SPIKE_AUDIT.md) 与
|
||||
[spike-path experiment](./experiments/k3/attnres_spike/)。
|
||||
第七轮冻结 `llm-atlas-k3-attnres-local-path-v1`,把 Round 06 的全局 value-route
|
||||
sensitivity 收缩为 14 个 exact selector masks,并同时从 learned 背景测 sufficiency、
|
||||
从 all-uniform 背景测 restoration。groups 6+7 的 16 个 depth mixers 在充分性方向
|
||||
两个指标 × 三 seed 的 `6 / 6` 格全部超过 `.50` global log gap;恢复性却只有 contrast
|
||||
`3 / 3` 通过,peak `0 / 3` 通过,mean peak restoration 仅 `.380`。因此冻结结论是
|
||||
`one_sided_evidence_localization_not_established`,不是“定位成功”。group 7 的
|
||||
MLP-only 次级门 6 / 6 通过,但只属于 sufficiency 探索;output-only 为 0 / 6。
|
||||
三个正式格与完整 replay 共处理 262,144,000 target bytes,Round 06 equivalence、
|
||||
14 / 14 forward identity、65-node census、selector 与 canonical replay 全部 exact。
|
||||
详见 [K3_ATTNRES_LOCAL_PATH_PROTOCOL.md](./research/K3_ATTNRES_LOCAL_PATH_PROTOCOL.md)、
|
||||
[K3_ATTNRES_LOCAL_PATH_AUDIT.md](./research/K3_ATTNRES_LOCAL_PATH_AUDIT.md) 与
|
||||
[local-path experiment](./experiments/k3/attnres_local_path/)。
|
||||
DeepSeek 八轮专题以 24 张问题账、10 次技术转向、
|
||||
22 个交互实验和 60 个一手/官方节点,串起 Dense、MoE、MLA、V3 协同、R1 与 V4;
|
||||
并固定官方 V2-Lite revision,在 RTX 5090 上连续执行 7/27 层,记录 3,240 次真实专家选择、
|
||||
|
||||
@@ -0,0 +1,50 @@
|
||||
# Attention Residuals local mixer-path diagnostics
|
||||
|
||||
This directory implements preregistered protocol
|
||||
`llm-atlas-k3-attnres-local-path-v1`.
|
||||
|
||||
It is a targeted follow-up to Round 06. It exact-replays the same depth-32
|
||||
Block training and keeps the learned forward unchanged while switching source
|
||||
value-gradient coefficients only at frozen mixer scopes. It is not a Kimi K3
|
||||
checkpoint run, a trainable variant, an additive attribution, or a reproduction
|
||||
of unpublished Figure 5 telemetry.
|
||||
|
||||
## Frozen environment
|
||||
|
||||
```text
|
||||
Python /home/wuyang/.pyenv/versions/3.10.14/envs/navi-router-cu128/bin/python
|
||||
PyTorch 2.11.0+cu128
|
||||
GPU NVIDIA GeForce RTX 5090
|
||||
CUBLAS_WORKSPACE_CONFIG=:4096:8
|
||||
```
|
||||
|
||||
The preregistration was committed as `6911efc` before the runner or any result
|
||||
file existed.
|
||||
|
||||
## Step-0 smoke
|
||||
|
||||
```bash
|
||||
CUBLAS_WORKSPACE_CONFIG=:4096:8 \
|
||||
/home/wuyang/.pyenv/versions/3.10.14/envs/navi-router-cu128/bin/python \
|
||||
experiments/k3/attnres_local_path/train.py \
|
||||
--run-kind smoke \
|
||||
--seed 2026073001 \
|
||||
--cache-dir /home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1 \
|
||||
--data-manifest experiments/k3/attnres_gradient/manifest.json \
|
||||
--parent-manifest experiments/k3/attnres_spike/manifest.json \
|
||||
--manifest experiments/k3/attnres_local_path/manifest.json \
|
||||
--output /home/wuyang/.cache/llm-atlas/k3-attnres-local-path-v1/smoke/seed-2026073001.json
|
||||
```
|
||||
|
||||
The first smoke passed all 14-mode forward-identity, selector, Round 06 endpoint,
|
||||
initialization-negative-control, parent-learned, and loss-scale gates. Its
|
||||
canonical content hash is
|
||||
`f708200f4fb122f61f30b97393839382a2094a71a8ea5d48cf47fb7fa094e69b`.
|
||||
|
||||
Formal cells use the same command with `--run-kind formal`, one of the three
|
||||
manifest seeds, and a new output path. The independent seed-2026073001 run uses
|
||||
`--run-kind replay`.
|
||||
|
||||
Only `analyze.py` may calculate the global log gap, sufficiency/restoration
|
||||
scores, and preregistered gates. The site consumes its frozen aggregate rather
|
||||
than reimplementing thresholds in TypeScript.
|
||||
@@ -0,0 +1,498 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Aggregate and gate preregistered Round 07 local-path results."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import copy
|
||||
import hashlib
|
||||
import json
|
||||
import math
|
||||
import statistics
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
|
||||
PROTOCOL_ID = "llm-atlas-k3-attnres-local-path-v1"
|
||||
METRICS = ("spike_contrast", "peak_normalized")
|
||||
SUFFICIENCY_MODES = (
|
||||
"uniform_group_6_only",
|
||||
"uniform_group_7_only",
|
||||
"uniform_groups_6_7_only",
|
||||
"uniform_group_6_attention_only",
|
||||
"uniform_group_6_mlp_only",
|
||||
"uniform_group_7_attention_only",
|
||||
"uniform_group_7_mlp_only",
|
||||
"uniform_output_only",
|
||||
"uniform_depth_all",
|
||||
"uniform_all",
|
||||
)
|
||||
RESTORATION_MODES = (
|
||||
"uniform_except_group_6",
|
||||
"uniform_except_group_7",
|
||||
"uniform_except_groups_6_7",
|
||||
)
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--manifest", type=Path, required=True)
|
||||
parser.add_argument(
|
||||
"--formal", type=Path, action="append", required=True
|
||||
)
|
||||
parser.add_argument("--replay", type=Path, required=True)
|
||||
parser.add_argument("--output", type=Path, required=True)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def file_sha256(path: Path) -> str:
|
||||
digest = hashlib.sha256()
|
||||
with path.open("rb") as handle:
|
||||
for chunk in iter(lambda: handle.read(1024 * 1024), b""):
|
||||
digest.update(chunk)
|
||||
return digest.hexdigest()
|
||||
|
||||
|
||||
def canonical_sha256(value: Any) -> str:
|
||||
payload = json.dumps(
|
||||
value, ensure_ascii=False, sort_keys=True, separators=(",", ":")
|
||||
).encode()
|
||||
return hashlib.sha256(payload).hexdigest()
|
||||
|
||||
|
||||
def read_result(path: Path) -> dict[str, Any]:
|
||||
value = json.loads(path.read_text())
|
||||
if value["protocol_id"] != PROTOCOL_ID:
|
||||
raise RuntimeError(f"protocol mismatch: {path}")
|
||||
expected = value["canonical_sha256_without_self"]
|
||||
without_self = {
|
||||
key: item
|
||||
for key, item in value.items()
|
||||
if key != "canonical_sha256_without_self"
|
||||
}
|
||||
if canonical_sha256(without_self) != expected:
|
||||
raise RuntimeError(f"canonical self-hash mismatch: {path}")
|
||||
return value
|
||||
|
||||
|
||||
def replay_payload(value: dict[str, Any]) -> dict[str, Any]:
|
||||
cleaned = copy.deepcopy(value)
|
||||
for key in ("run_kind", "timing", "canonical_sha256_without_self"):
|
||||
cleaned.pop(key)
|
||||
return cleaned
|
||||
|
||||
|
||||
def primary_metrics(mode: dict[str, Any]) -> dict[str, float]:
|
||||
stats = mode["positions"]["post_mlp_state"]["reductions"][
|
||||
"element_rms"
|
||||
]["statistics"]
|
||||
return {name: float(stats[name]) for name in METRICS}
|
||||
|
||||
|
||||
def all_true(values: list[bool]) -> bool:
|
||||
return len(values) > 0 and all(values)
|
||||
|
||||
|
||||
def mixed_status(
|
||||
by_seed_metric: dict[str, dict[str, dict[str, Any]]],
|
||||
threshold: float,
|
||||
) -> dict[str, Any]:
|
||||
passed = []
|
||||
signs = []
|
||||
metric_passes = {metric: [] for metric in METRICS}
|
||||
for seed_values in by_seed_metric.values():
|
||||
for metric in METRICS:
|
||||
score = seed_values[metric]["score"]
|
||||
cell_passed = score is not None and score >= threshold
|
||||
passed.append(cell_passed)
|
||||
metric_passes[metric].append(cell_passed)
|
||||
if score is not None:
|
||||
signs.append(1 if score >= 0 else -1)
|
||||
reasons = []
|
||||
if any(passed) and not all(passed):
|
||||
reasons.append("seed_or_metric_pass_split")
|
||||
if metric_passes[METRICS[0]] != metric_passes[METRICS[1]]:
|
||||
reasons.append("metric_direction_split")
|
||||
if len(set(signs)) > 1:
|
||||
reasons.append("score_sign_split")
|
||||
return {
|
||||
"passed": all_true(passed),
|
||||
"threshold": threshold,
|
||||
"required_cells": len(passed),
|
||||
"passed_cells": sum(passed),
|
||||
"mixed": bool(reasons),
|
||||
"mixed_reasons": reasons,
|
||||
}
|
||||
|
||||
|
||||
def main() -> None:
|
||||
args = parse_args()
|
||||
manifest = json.loads(args.manifest.read_text())
|
||||
if manifest["protocol_id"] != PROTOCOL_ID:
|
||||
raise RuntimeError("manifest protocol mismatch")
|
||||
if len(args.formal) != 3:
|
||||
raise RuntimeError("exactly three formal result paths are required")
|
||||
|
||||
formal_pairs = [(path, read_result(path)) for path in args.formal]
|
||||
formal_pairs.sort(key=lambda item: item[1]["seed"])
|
||||
expected_seeds = manifest["formal_seeds"]
|
||||
if [value["seed"] for _, value in formal_pairs] != expected_seeds:
|
||||
raise RuntimeError("formal seeds do not match manifest")
|
||||
for path, value in formal_pairs:
|
||||
if value["run_kind"] != "formal" or value["steps"] != 8000:
|
||||
raise RuntimeError(f"invalid formal cell: {path}")
|
||||
if not value["round06_equivalence"]["passed"]:
|
||||
raise RuntimeError(f"parent equivalence failed: {path}")
|
||||
for diagnostic in value["diagnostics"]:
|
||||
if not diagnostic["parent_learned_round06_exact"]:
|
||||
raise RuntimeError(f"parent diagnostic mismatch: {path}")
|
||||
if diagnostic["local_matrix"] is not None:
|
||||
for gate in ("forward_identity_gate", "endpoint_exactness"):
|
||||
if not diagnostic[gate]["passed"]:
|
||||
raise RuntimeError(f"{gate} failed: {path}")
|
||||
for mode in manifest["matrix_modes"]:
|
||||
if not diagnostic["local_matrix"][mode]["selector"][
|
||||
"passed"
|
||||
]:
|
||||
raise RuntimeError(f"selector failed: {path}:{mode}")
|
||||
step0 = value["diagnostics"][0]
|
||||
if (
|
||||
not step0["initialization_negative_control"]["passed"]
|
||||
or not step0["loss_scale_gate"]["passed"]
|
||||
):
|
||||
raise RuntimeError(f"step-0 control failed: {path}")
|
||||
|
||||
replay = read_result(args.replay)
|
||||
if (
|
||||
replay["run_kind"] != "replay"
|
||||
or replay["seed"] != expected_seeds[0]
|
||||
or replay["steps"] != 8000
|
||||
):
|
||||
raise RuntimeError("invalid replay cell")
|
||||
replay_exact = (
|
||||
replay_payload(formal_pairs[0][1]) == replay_payload(replay)
|
||||
)
|
||||
if not replay_exact:
|
||||
raise RuntimeError("formal seed1 and replay are not canonical exact")
|
||||
|
||||
thresholds = manifest["thresholds"]
|
||||
cells = []
|
||||
sufficiency_by_mode: dict[str, dict[str, dict[str, Any]]] = {
|
||||
mode: {} for mode in SUFFICIENCY_MODES
|
||||
}
|
||||
restoration_by_mode: dict[str, dict[str, dict[str, Any]]] = {
|
||||
mode: {} for mode in RESTORATION_MODES
|
||||
}
|
||||
|
||||
for path, value in formal_pairs:
|
||||
seed_key = str(value["seed"])
|
||||
final = value["diagnostics"][-1]["local_matrix"]
|
||||
mode_metrics = {
|
||||
mode: primary_metrics(final[mode])
|
||||
for mode in manifest["matrix_modes"]
|
||||
}
|
||||
global_metrics = {}
|
||||
for metric in METRICS:
|
||||
reference = mode_metrics["detached_learned"][metric]
|
||||
uniform_all = mode_metrics["uniform_all"][metric]
|
||||
log_gap = math.log(reference / uniform_all)
|
||||
relative_drop = (reference - uniform_all) / reference
|
||||
established = (
|
||||
reference
|
||||
> thresholds["positive_denominator_epsilon"]
|
||||
and uniform_all
|
||||
> thresholds["positive_denominator_epsilon"]
|
||||
and log_gap > 0
|
||||
and relative_drop
|
||||
>= thresholds["global_relative_drop_minimum"]
|
||||
)
|
||||
global_metrics[metric] = {
|
||||
"reference": reference,
|
||||
"uniform_all": uniform_all,
|
||||
"log_gap": log_gap,
|
||||
"relative_drop": relative_drop,
|
||||
"established": established,
|
||||
}
|
||||
for mode in SUFFICIENCY_MODES:
|
||||
sufficiency_by_mode[mode][seed_key] = {}
|
||||
for metric in METRICS:
|
||||
established = global_metrics[metric]["established"]
|
||||
score = (
|
||||
math.log(
|
||||
mode_metrics["detached_learned"][metric]
|
||||
/ mode_metrics[mode][metric]
|
||||
)
|
||||
/ global_metrics[metric]["log_gap"]
|
||||
if established
|
||||
else None
|
||||
)
|
||||
sufficiency_by_mode[mode][seed_key][metric] = {
|
||||
"score": score,
|
||||
"metric_value": mode_metrics[mode][metric],
|
||||
"global_gap_established": established,
|
||||
}
|
||||
for mode in RESTORATION_MODES:
|
||||
restoration_by_mode[mode][seed_key] = {}
|
||||
for metric in METRICS:
|
||||
established = global_metrics[metric]["established"]
|
||||
score = (
|
||||
math.log(
|
||||
mode_metrics[mode][metric]
|
||||
/ mode_metrics["uniform_all"][metric]
|
||||
)
|
||||
/ global_metrics[metric]["log_gap"]
|
||||
if established
|
||||
else None
|
||||
)
|
||||
restoration_by_mode[mode][seed_key][metric] = {
|
||||
"score": score,
|
||||
"metric_value": mode_metrics[mode][metric],
|
||||
"global_gap_established": established,
|
||||
}
|
||||
cells.append(
|
||||
{
|
||||
"seed": value["seed"],
|
||||
"path": str(path),
|
||||
"file_sha256": file_sha256(path),
|
||||
"canonical_sha256": value[
|
||||
"canonical_sha256_without_self"
|
||||
],
|
||||
"mode_metrics": mode_metrics,
|
||||
"global": global_metrics,
|
||||
"interaction_residual": {
|
||||
metric: (
|
||||
1.0
|
||||
- sufficiency_by_mode["uniform_output_only"][
|
||||
seed_key
|
||||
][metric]["score"]
|
||||
- sufficiency_by_mode["uniform_depth_all"][
|
||||
seed_key
|
||||
][metric]["score"]
|
||||
)
|
||||
for metric in METRICS
|
||||
},
|
||||
}
|
||||
)
|
||||
|
||||
global_gap_passed = all_true(
|
||||
[
|
||||
cell["global"][metric]["established"]
|
||||
for cell in cells
|
||||
for metric in METRICS
|
||||
]
|
||||
)
|
||||
sufficiency_gates = {
|
||||
"groups_6_7": mixed_status(
|
||||
sufficiency_by_mode["uniform_groups_6_7_only"],
|
||||
thresholds["groups_6_7_sufficiency_minimum"],
|
||||
),
|
||||
"group_6": mixed_status(
|
||||
sufficiency_by_mode["uniform_group_6_only"],
|
||||
thresholds["single_group_material_minimum"],
|
||||
),
|
||||
"group_7": mixed_status(
|
||||
sufficiency_by_mode["uniform_group_7_only"],
|
||||
thresholds["single_group_material_minimum"],
|
||||
),
|
||||
"output_half_gap": mixed_status(
|
||||
sufficiency_by_mode["uniform_output_only"],
|
||||
thresholds["output_half_gap_minimum"],
|
||||
),
|
||||
}
|
||||
restoration_gates = {
|
||||
"groups_6_7": mixed_status(
|
||||
restoration_by_mode["uniform_except_groups_6_7"],
|
||||
thresholds["groups_6_7_restoration_minimum"],
|
||||
),
|
||||
"group_6": mixed_status(
|
||||
restoration_by_mode["uniform_except_group_6"],
|
||||
thresholds["single_group_material_minimum"],
|
||||
),
|
||||
"group_7": mixed_status(
|
||||
restoration_by_mode["uniform_except_group_7"],
|
||||
thresholds["single_group_material_minimum"],
|
||||
),
|
||||
}
|
||||
|
||||
branch_gates = {}
|
||||
for group in (6, 7):
|
||||
group_passed = sufficiency_gates[f"group_{group}"]["passed"]
|
||||
candidates = {}
|
||||
for branch, sibling in (("attention", "mlp"), ("mlp", "attention")):
|
||||
branch_mode = f"uniform_group_{group}_{branch}_only"
|
||||
sibling_mode = f"uniform_group_{group}_{sibling}_only"
|
||||
checks = []
|
||||
margins = []
|
||||
for seed in expected_seeds:
|
||||
seed_key = str(seed)
|
||||
for metric in METRICS:
|
||||
left = sufficiency_by_mode[branch_mode][seed_key][
|
||||
metric
|
||||
]["score"]
|
||||
right = sufficiency_by_mode[sibling_mode][seed_key][
|
||||
metric
|
||||
]["score"]
|
||||
margin = (
|
||||
left - right
|
||||
if left is not None and right is not None
|
||||
else None
|
||||
)
|
||||
margins.append(margin)
|
||||
checks.append(
|
||||
left is not None
|
||||
and left >= thresholds["branch_material_minimum"]
|
||||
and margin is not None
|
||||
and margin
|
||||
>= thresholds["branch_dominance_margin"]
|
||||
)
|
||||
candidates[branch] = {
|
||||
"passed": group_passed and all_true(checks),
|
||||
"group_gate_passed": group_passed,
|
||||
"passed_cells": sum(checks),
|
||||
"required_cells": len(checks),
|
||||
"margins": margins,
|
||||
}
|
||||
dominant = [
|
||||
branch
|
||||
for branch, gate in candidates.items()
|
||||
if gate["passed"]
|
||||
]
|
||||
branch_gates[f"group_{group}"] = {
|
||||
"passed": len(dominant) == 1,
|
||||
"dominant_branch": dominant[0] if len(dominant) == 1 else None,
|
||||
"exploratory_sufficiency_only": True,
|
||||
"candidates": candidates,
|
||||
}
|
||||
|
||||
localization_passed = (
|
||||
global_gap_passed
|
||||
and sufficiency_gates["groups_6_7"]["passed"]
|
||||
and restoration_gates["groups_6_7"]["passed"]
|
||||
)
|
||||
localization_status = (
|
||||
"established_at_preregistered_bidirectional_50pct_threshold"
|
||||
if localization_passed
|
||||
else (
|
||||
"one_sided_evidence_localization_not_established"
|
||||
if (
|
||||
sufficiency_gates["groups_6_7"]["passed"]
|
||||
!= restoration_gates["groups_6_7"]["passed"]
|
||||
)
|
||||
else "not_established_at_preregistered_threshold"
|
||||
)
|
||||
)
|
||||
|
||||
mode_means = {}
|
||||
for mode in manifest["matrix_modes"]:
|
||||
mode_means[mode] = {
|
||||
metric: statistics.fmean(
|
||||
cell["mode_metrics"][mode][metric] for cell in cells
|
||||
)
|
||||
for metric in METRICS
|
||||
}
|
||||
sufficiency_means = {
|
||||
mode: {
|
||||
metric: statistics.fmean(
|
||||
sufficiency_by_mode[mode][str(seed)][metric]["score"]
|
||||
for seed in expected_seeds
|
||||
)
|
||||
for metric in METRICS
|
||||
}
|
||||
for mode in SUFFICIENCY_MODES
|
||||
}
|
||||
restoration_means = {
|
||||
mode: {
|
||||
metric: statistics.fmean(
|
||||
restoration_by_mode[mode][str(seed)][metric]["score"]
|
||||
for seed in expected_seeds
|
||||
)
|
||||
for metric in METRICS
|
||||
}
|
||||
for mode in RESTORATION_MODES
|
||||
}
|
||||
|
||||
result = {
|
||||
"schema_version": 1,
|
||||
"protocol_id": PROTOCOL_ID,
|
||||
"study_identity": manifest["study_identity"],
|
||||
"manifest": {
|
||||
"path": str(args.manifest),
|
||||
"file_sha256": file_sha256(args.manifest),
|
||||
},
|
||||
"formal_cells": cells,
|
||||
"replay": {
|
||||
"path": str(args.replay),
|
||||
"file_sha256": file_sha256(args.replay),
|
||||
"canonical_sha256": replay[
|
||||
"canonical_sha256_without_self"
|
||||
],
|
||||
"formal_seed1_exact_excluding_run_kind_and_timing": replay_exact,
|
||||
},
|
||||
"scores": {
|
||||
"sufficiency": sufficiency_by_mode,
|
||||
"restoration": restoration_by_mode,
|
||||
},
|
||||
"means": {
|
||||
"mode_metrics": mode_means,
|
||||
"sufficiency": sufficiency_means,
|
||||
"restoration": restoration_means,
|
||||
},
|
||||
"gates": {
|
||||
"all_input_and_parent_gates_passed": True,
|
||||
"global_gap": {
|
||||
"passed": global_gap_passed,
|
||||
"required_cells": 6,
|
||||
"passed_cells": sum(
|
||||
cell["global"][metric]["established"]
|
||||
for cell in cells
|
||||
for metric in METRICS
|
||||
),
|
||||
},
|
||||
"sufficiency": sufficiency_gates,
|
||||
"restoration": restoration_gates,
|
||||
"localization": {
|
||||
"passed": localization_passed,
|
||||
"status": localization_status,
|
||||
"requires": (
|
||||
"groups 6+7 sufficiency and restoration both >=0.50 "
|
||||
"for 3/3 seeds and both metrics"
|
||||
),
|
||||
},
|
||||
"branch_dominance": branch_gates,
|
||||
},
|
||||
"limitations": [
|
||||
"same-forward diagnostic backward-rule sensitivity only",
|
||||
"reduced byte-level language model, not the Kimi K3 checkpoint",
|
||||
"effects are non-additive and are not contribution percentages",
|
||||
"group 7 includes layers 26-28 outside fixed spike set 21-25",
|
||||
"three-seed threshold gates are not population inference",
|
||||
],
|
||||
}
|
||||
result["canonical_sha256_without_self"] = canonical_sha256(result)
|
||||
args.output.parent.mkdir(parents=True, exist_ok=True)
|
||||
temporary = args.output.with_suffix(args.output.suffix + ".tmp")
|
||||
temporary.write_text(
|
||||
json.dumps(result, ensure_ascii=False, indent=2, sort_keys=True)
|
||||
+ "\n"
|
||||
)
|
||||
temporary.replace(args.output)
|
||||
print(
|
||||
json.dumps(
|
||||
{
|
||||
"output": str(args.output),
|
||||
"formal_cells": len(cells),
|
||||
"replay_exact": replay_exact,
|
||||
"global_gap": global_gap_passed,
|
||||
"localization": localization_status,
|
||||
"canonical_sha256": result[
|
||||
"canonical_sha256_without_self"
|
||||
],
|
||||
},
|
||||
ensure_ascii=False,
|
||||
indent=2,
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,540 @@
|
||||
{
|
||||
"schema_version": 1,
|
||||
"protocol_id": "llm-atlas-k3-attnres-local-path-v1",
|
||||
"parent_protocol_id": "llm-atlas-k3-attnres-spike-path-v1",
|
||||
"study_identity": "targeted local-path follow-up informed by Round 05 and Round 06; not blind discovery",
|
||||
"architecture": "block",
|
||||
"depth": 32,
|
||||
"aggregation_groups": 8,
|
||||
"blocks_per_group": 4,
|
||||
"depth_mixers": 64,
|
||||
"output_mixers": 1,
|
||||
"formal_seeds": [
|
||||
2026073001,
|
||||
2026073002,
|
||||
2026073003
|
||||
],
|
||||
"replay": {
|
||||
"architecture": "block",
|
||||
"depth": 32,
|
||||
"seed": 2026073001,
|
||||
"environment_scope": "same host, GPU, Python, PyTorch, CUDA and CUBLAS_WORKSPACE_CONFIG"
|
||||
},
|
||||
"training": {
|
||||
"steps": 8000,
|
||||
"batch_size": 32,
|
||||
"context": 256,
|
||||
"target_bytes_per_cell": 65536000,
|
||||
"parent_diagnostic_steps": [
|
||||
0,
|
||||
100,
|
||||
500,
|
||||
2000,
|
||||
4000,
|
||||
8000
|
||||
],
|
||||
"local_matrix_steps": [
|
||||
0,
|
||||
8000
|
||||
],
|
||||
"matrix_used_during_training": false
|
||||
},
|
||||
"primary_object": {
|
||||
"position": "post_mlp_state",
|
||||
"reduction": "element_rms",
|
||||
"spike_layers_one_based": [
|
||||
21,
|
||||
22,
|
||||
23,
|
||||
24,
|
||||
25
|
||||
],
|
||||
"metrics": [
|
||||
"spike_contrast",
|
||||
"peak_normalized"
|
||||
]
|
||||
},
|
||||
"matrix_modes": [
|
||||
"detached_learned",
|
||||
"uniform_group_6_only",
|
||||
"uniform_group_7_only",
|
||||
"uniform_groups_6_7_only",
|
||||
"uniform_group_6_attention_only",
|
||||
"uniform_group_6_mlp_only",
|
||||
"uniform_group_7_attention_only",
|
||||
"uniform_group_7_mlp_only",
|
||||
"uniform_output_only",
|
||||
"uniform_depth_all",
|
||||
"uniform_all",
|
||||
"uniform_except_group_6",
|
||||
"uniform_except_group_7",
|
||||
"uniform_except_groups_6_7"
|
||||
],
|
||||
"selector": {
|
||||
"coefficient_choices": [
|
||||
"detached_learned",
|
||||
"uniform"
|
||||
],
|
||||
"full_autograd_learned_in_matrix": false,
|
||||
"depth_identity": "kind=depth,index=0..63; layer=floor(index/2)+1; attention iff index even",
|
||||
"output_identity": "kind=output,index=64; parallel-schema alias only; parent mixer_index remains null",
|
||||
"uniform_depth_indices": {
|
||||
"detached_learned": [],
|
||||
"uniform_group_6_only": [
|
||||
40,
|
||||
41,
|
||||
42,
|
||||
43,
|
||||
44,
|
||||
45,
|
||||
46,
|
||||
47
|
||||
],
|
||||
"uniform_group_7_only": [
|
||||
48,
|
||||
49,
|
||||
50,
|
||||
51,
|
||||
52,
|
||||
53,
|
||||
54,
|
||||
55
|
||||
],
|
||||
"uniform_groups_6_7_only": [
|
||||
40,
|
||||
41,
|
||||
42,
|
||||
43,
|
||||
44,
|
||||
45,
|
||||
46,
|
||||
47,
|
||||
48,
|
||||
49,
|
||||
50,
|
||||
51,
|
||||
52,
|
||||
53,
|
||||
54,
|
||||
55
|
||||
],
|
||||
"uniform_group_6_attention_only": [
|
||||
40,
|
||||
42,
|
||||
44,
|
||||
46
|
||||
],
|
||||
"uniform_group_6_mlp_only": [
|
||||
41,
|
||||
43,
|
||||
45,
|
||||
47
|
||||
],
|
||||
"uniform_group_7_attention_only": [
|
||||
48,
|
||||
50,
|
||||
52,
|
||||
54
|
||||
],
|
||||
"uniform_group_7_mlp_only": [
|
||||
49,
|
||||
51,
|
||||
53,
|
||||
55
|
||||
],
|
||||
"uniform_output_only": [],
|
||||
"uniform_depth_all": [
|
||||
0,
|
||||
1,
|
||||
2,
|
||||
3,
|
||||
4,
|
||||
5,
|
||||
6,
|
||||
7,
|
||||
8,
|
||||
9,
|
||||
10,
|
||||
11,
|
||||
12,
|
||||
13,
|
||||
14,
|
||||
15,
|
||||
16,
|
||||
17,
|
||||
18,
|
||||
19,
|
||||
20,
|
||||
21,
|
||||
22,
|
||||
23,
|
||||
24,
|
||||
25,
|
||||
26,
|
||||
27,
|
||||
28,
|
||||
29,
|
||||
30,
|
||||
31,
|
||||
32,
|
||||
33,
|
||||
34,
|
||||
35,
|
||||
36,
|
||||
37,
|
||||
38,
|
||||
39,
|
||||
40,
|
||||
41,
|
||||
42,
|
||||
43,
|
||||
44,
|
||||
45,
|
||||
46,
|
||||
47,
|
||||
48,
|
||||
49,
|
||||
50,
|
||||
51,
|
||||
52,
|
||||
53,
|
||||
54,
|
||||
55,
|
||||
56,
|
||||
57,
|
||||
58,
|
||||
59,
|
||||
60,
|
||||
61,
|
||||
62,
|
||||
63
|
||||
],
|
||||
"uniform_all": [
|
||||
0,
|
||||
1,
|
||||
2,
|
||||
3,
|
||||
4,
|
||||
5,
|
||||
6,
|
||||
7,
|
||||
8,
|
||||
9,
|
||||
10,
|
||||
11,
|
||||
12,
|
||||
13,
|
||||
14,
|
||||
15,
|
||||
16,
|
||||
17,
|
||||
18,
|
||||
19,
|
||||
20,
|
||||
21,
|
||||
22,
|
||||
23,
|
||||
24,
|
||||
25,
|
||||
26,
|
||||
27,
|
||||
28,
|
||||
29,
|
||||
30,
|
||||
31,
|
||||
32,
|
||||
33,
|
||||
34,
|
||||
35,
|
||||
36,
|
||||
37,
|
||||
38,
|
||||
39,
|
||||
40,
|
||||
41,
|
||||
42,
|
||||
43,
|
||||
44,
|
||||
45,
|
||||
46,
|
||||
47,
|
||||
48,
|
||||
49,
|
||||
50,
|
||||
51,
|
||||
52,
|
||||
53,
|
||||
54,
|
||||
55,
|
||||
56,
|
||||
57,
|
||||
58,
|
||||
59,
|
||||
60,
|
||||
61,
|
||||
62,
|
||||
63
|
||||
],
|
||||
"uniform_except_group_6": [
|
||||
0,
|
||||
1,
|
||||
2,
|
||||
3,
|
||||
4,
|
||||
5,
|
||||
6,
|
||||
7,
|
||||
8,
|
||||
9,
|
||||
10,
|
||||
11,
|
||||
12,
|
||||
13,
|
||||
14,
|
||||
15,
|
||||
16,
|
||||
17,
|
||||
18,
|
||||
19,
|
||||
20,
|
||||
21,
|
||||
22,
|
||||
23,
|
||||
24,
|
||||
25,
|
||||
26,
|
||||
27,
|
||||
28,
|
||||
29,
|
||||
30,
|
||||
31,
|
||||
32,
|
||||
33,
|
||||
34,
|
||||
35,
|
||||
36,
|
||||
37,
|
||||
38,
|
||||
39,
|
||||
48,
|
||||
49,
|
||||
50,
|
||||
51,
|
||||
52,
|
||||
53,
|
||||
54,
|
||||
55,
|
||||
56,
|
||||
57,
|
||||
58,
|
||||
59,
|
||||
60,
|
||||
61,
|
||||
62,
|
||||
63
|
||||
],
|
||||
"uniform_except_group_7": [
|
||||
0,
|
||||
1,
|
||||
2,
|
||||
3,
|
||||
4,
|
||||
5,
|
||||
6,
|
||||
7,
|
||||
8,
|
||||
9,
|
||||
10,
|
||||
11,
|
||||
12,
|
||||
13,
|
||||
14,
|
||||
15,
|
||||
16,
|
||||
17,
|
||||
18,
|
||||
19,
|
||||
20,
|
||||
21,
|
||||
22,
|
||||
23,
|
||||
24,
|
||||
25,
|
||||
26,
|
||||
27,
|
||||
28,
|
||||
29,
|
||||
30,
|
||||
31,
|
||||
32,
|
||||
33,
|
||||
34,
|
||||
35,
|
||||
36,
|
||||
37,
|
||||
38,
|
||||
39,
|
||||
40,
|
||||
41,
|
||||
42,
|
||||
43,
|
||||
44,
|
||||
45,
|
||||
46,
|
||||
47,
|
||||
56,
|
||||
57,
|
||||
58,
|
||||
59,
|
||||
60,
|
||||
61,
|
||||
62,
|
||||
63
|
||||
],
|
||||
"uniform_except_groups_6_7": [
|
||||
0,
|
||||
1,
|
||||
2,
|
||||
3,
|
||||
4,
|
||||
5,
|
||||
6,
|
||||
7,
|
||||
8,
|
||||
9,
|
||||
10,
|
||||
11,
|
||||
12,
|
||||
13,
|
||||
14,
|
||||
15,
|
||||
16,
|
||||
17,
|
||||
18,
|
||||
19,
|
||||
20,
|
||||
21,
|
||||
22,
|
||||
23,
|
||||
24,
|
||||
25,
|
||||
26,
|
||||
27,
|
||||
28,
|
||||
29,
|
||||
30,
|
||||
31,
|
||||
32,
|
||||
33,
|
||||
34,
|
||||
35,
|
||||
36,
|
||||
37,
|
||||
38,
|
||||
39,
|
||||
56,
|
||||
57,
|
||||
58,
|
||||
59,
|
||||
60,
|
||||
61,
|
||||
62,
|
||||
63
|
||||
]
|
||||
},
|
||||
"uniform_output": {
|
||||
"detached_learned": false,
|
||||
"uniform_group_6_only": false,
|
||||
"uniform_group_7_only": false,
|
||||
"uniform_groups_6_7_only": false,
|
||||
"uniform_group_6_attention_only": false,
|
||||
"uniform_group_6_mlp_only": false,
|
||||
"uniform_group_7_attention_only": false,
|
||||
"uniform_group_7_mlp_only": false,
|
||||
"uniform_output_only": true,
|
||||
"uniform_depth_all": false,
|
||||
"uniform_all": true,
|
||||
"uniform_except_group_6": true,
|
||||
"uniform_except_group_7": true,
|
||||
"uniform_except_groups_6_7": true
|
||||
},
|
||||
"expected_uniform_counts": {
|
||||
"detached_learned": 0,
|
||||
"uniform_group_6_only": 8,
|
||||
"uniform_group_7_only": 8,
|
||||
"uniform_groups_6_7_only": 16,
|
||||
"uniform_group_6_attention_only": 4,
|
||||
"uniform_group_6_mlp_only": 4,
|
||||
"uniform_group_7_attention_only": 4,
|
||||
"uniform_group_7_mlp_only": 4,
|
||||
"uniform_output_only": 1,
|
||||
"uniform_depth_all": 64,
|
||||
"uniform_all": 65,
|
||||
"uniform_except_group_6": 57,
|
||||
"uniform_except_group_7": 57,
|
||||
"uniform_except_groups_6_7": 49
|
||||
}
|
||||
},
|
||||
"thresholds": {
|
||||
"positive_denominator_epsilon": 1e-30,
|
||||
"step0_spectrum_tolerance": 1e-6,
|
||||
"loss_scale_tolerance": 1e-5,
|
||||
"global_relative_drop_minimum": 0.2,
|
||||
"groups_6_7_sufficiency_minimum": 0.5,
|
||||
"groups_6_7_restoration_minimum": 0.5,
|
||||
"single_group_material_minimum": 0.2,
|
||||
"branch_material_minimum": 0.2,
|
||||
"branch_dominance_margin": 0.15,
|
||||
"output_half_gap_minimum": 0.5,
|
||||
"formal_seed_gate": "3/3 independently for both metrics; means are display-only"
|
||||
},
|
||||
"formulas": {
|
||||
"global_log_gap": "G_X = ln(X_ref / X_uniform_all)",
|
||||
"global_relative_drop": "(X_ref - X_uniform_all) / X_ref",
|
||||
"sufficiency": "S_X(m) = ln(X_ref / X_m) / G_X",
|
||||
"restoration": "R_X(r) = ln(X_r / X_uniform_all) / G_X",
|
||||
"score_clipping": false
|
||||
},
|
||||
"parent_artifacts": {
|
||||
"manifest_path": "experiments/k3/attnres_spike/manifest.json",
|
||||
"manifest_sha256": "d5302a249249a07d362819134763d14e7d32307f22cff416c665ed9606142fef",
|
||||
"runner_path": "experiments/k3/attnres_spike/train.py",
|
||||
"runner_sha256": "77298081d3c491d2e88e4705995174b9879ef377f520eb5fe5ea107e7a1da084",
|
||||
"protocol_path": "research/K3_ATTNRES_SPIKE_PROTOCOL.md",
|
||||
"protocol_sha256": "6cb101b8760d9f1c81caeb2f16880b16152da103867224a06761a75a12984a16",
|
||||
"scoping_path": "research/K3_ATTNRES_SPIKE_SCOPING.md",
|
||||
"scoping_sha256": "590166bd62580bb8238293823cfcc39bc0a465fec4c697025343f3f1138abd27",
|
||||
"formal_schedule_sha256": "5041e09b167f229248d2462324e8c254b8f5938975f135dcd8192b00a54a4f4e",
|
||||
"validation_tensor_sha256": "f459316f13078a163b47c133511bb7181e05170ab89516e196490113893ce338",
|
||||
"diagnostic_tensor_sha256": "21117e31db302b10d67b63f035665dc8f220b879d216ccd12b7d2ba86e7b1716"
|
||||
},
|
||||
"current_artifacts": {
|
||||
"protocol_path": "research/K3_ATTNRES_LOCAL_PATH_PROTOCOL.md",
|
||||
"protocol_sha256": "5ecc7ca92314ddb50aecf0cb50e115814c8983aa8bffb30e3634f7b3ce6dca1d",
|
||||
"scoping_path": "research/K3_ATTNRES_LOCAL_PATH_SCOPING.md",
|
||||
"scoping_sha256": "670ca4edf31a4be1f54937d9c7a760dba7a96e1e820c38c6b10405e22b078fc8",
|
||||
"grok_review_path": "research/K3_ATTNRES_LOCAL_PATH_GROK_REVIEW.md",
|
||||
"grok_review_sha256": "2da1b6bf1f455c4121a7a2c5cfe40e102327dabafc7e24dccc23ed0d00ac6d71",
|
||||
"grok_session": "019fb151-9627-76c1-b7d7-53012874f85c"
|
||||
},
|
||||
"round06_expected": {
|
||||
"2026073001": {
|
||||
"raw_file_sha256": "e39e93b7a7fce3c56f5f14f95cfdc04afdce53628affee1202fe62bd1bdb7f71",
|
||||
"canonical_sha256": "76b0ccfb55c38baef50c395788ac4b351cbe0d58de70064b50702acb5c93f515",
|
||||
"final_model_state": "3f0b97ece3a15571ba3d656f589f512ca0bb9e20083c9f58a42ccaee14892f59",
|
||||
"final_optimizer_state": "ed03e6fbd4a12d8b063dcb22e0437754285f54d585374cd52fbd534f05d24637"
|
||||
},
|
||||
"2026073002": {
|
||||
"raw_file_sha256": "1c6f6c731030ec0adb2a8e7a4d586e0c4005cc3319568a7ac83c08c2a4b8eaf8",
|
||||
"canonical_sha256": "5352c74eca853b375c0e85933dafd7c5916c39fc59052e742ca14ffd6d68bc78",
|
||||
"final_model_state": "bd2556388aeaa211b798c283c7cbd8ccd29edf166a2922fa13d172e8dfdc38d1",
|
||||
"final_optimizer_state": "0b101eab3bc7d8d654be2ea335c86fc25563ce19912d721844ee4e639c569e77"
|
||||
},
|
||||
"2026073003": {
|
||||
"raw_file_sha256": "115f8245577ece6dfaaa8ada68445c186e6523a7f3b26efcc3eb4c0c4ce82406",
|
||||
"canonical_sha256": "7e764c07e90b78c4cd0acc2e99600225f16428cbb25d5188a0a5a8fe797f8766",
|
||||
"final_model_state": "638568aede21890773b6932a19ec4e112f5ac0a4770ba3402fcd82980a9ecf76",
|
||||
"final_optimizer_state": "83947fd743ec8e3e31ca7788fd201981846f0f1c7e9e38c88afcef95cfc6ec4e"
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,274 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Package frozen Round 07 outputs without recomputing any result gate."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import copy
|
||||
import hashlib
|
||||
import json
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
|
||||
PROTOCOL_ID = "llm-atlas-k3-attnres-local-path-v1"
|
||||
SEEDS = (2026073001, 2026073002, 2026073003)
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--raw-dir", type=Path, required=True)
|
||||
parser.add_argument("--manifest", type=Path, required=True)
|
||||
parser.add_argument("--aggregate", type=Path, required=True)
|
||||
parser.add_argument("--reproduction-output", type=Path, required=True)
|
||||
parser.add_argument("--compact-output", type=Path, required=True)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def file_sha256(path: Path) -> str:
|
||||
digest = hashlib.sha256()
|
||||
with path.open("rb") as handle:
|
||||
for chunk in iter(lambda: handle.read(1024 * 1024), b""):
|
||||
digest.update(chunk)
|
||||
return digest.hexdigest()
|
||||
|
||||
|
||||
def canonical_sha256(value: Any) -> str:
|
||||
payload = json.dumps(
|
||||
value, ensure_ascii=False, sort_keys=True, separators=(",", ":")
|
||||
).encode()
|
||||
return hashlib.sha256(payload).hexdigest()
|
||||
|
||||
|
||||
def load_canonical(path: Path) -> dict[str, Any]:
|
||||
value = json.loads(path.read_text())
|
||||
expected = value["canonical_sha256_without_self"]
|
||||
payload = {
|
||||
key: item
|
||||
for key, item in value.items()
|
||||
if key != "canonical_sha256_without_self"
|
||||
}
|
||||
if canonical_sha256(payload) != expected:
|
||||
raise RuntimeError(f"canonical hash mismatch: {path}")
|
||||
return value
|
||||
|
||||
|
||||
def write_canonical(path: Path, value: dict[str, Any]) -> None:
|
||||
value["canonical_sha256_without_self"] = canonical_sha256(value)
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
temporary = path.with_suffix(path.suffix + ".tmp")
|
||||
temporary.write_text(
|
||||
json.dumps(value, ensure_ascii=False, indent=2, sort_keys=True)
|
||||
+ "\n"
|
||||
)
|
||||
temporary.replace(path)
|
||||
|
||||
|
||||
def replay_payload(value: dict[str, Any]) -> dict[str, Any]:
|
||||
cleaned = copy.deepcopy(value)
|
||||
for key in ("run_kind", "timing", "canonical_sha256_without_self"):
|
||||
cleaned.pop(key)
|
||||
return cleaned
|
||||
|
||||
|
||||
def artifact_hashes(repo_root: Path) -> dict[str, str]:
|
||||
paths = {
|
||||
"runner": "experiments/k3/attnres_local_path/train.py",
|
||||
"analyzer": "experiments/k3/attnres_local_path/analyze.py",
|
||||
"packager": "experiments/k3/attnres_local_path/package.py",
|
||||
"manifest": "experiments/k3/attnres_local_path/manifest.json",
|
||||
"protocol": "research/K3_ATTNRES_LOCAL_PATH_PROTOCOL.md",
|
||||
"scoping": "research/K3_ATTNRES_LOCAL_PATH_SCOPING.md",
|
||||
"preresult_grok_review": (
|
||||
"research/K3_ATTNRES_LOCAL_PATH_GROK_REVIEW.md"
|
||||
),
|
||||
}
|
||||
return {
|
||||
name: file_sha256(repo_root / path) for name, path in paths.items()
|
||||
}
|
||||
|
||||
|
||||
def main() -> None:
|
||||
args = parse_args()
|
||||
repo_root = Path(__file__).resolve().parents[3]
|
||||
manifest = json.loads(args.manifest.read_text())
|
||||
aggregate = load_canonical(args.aggregate)
|
||||
if (
|
||||
manifest["protocol_id"] != PROTOCOL_ID
|
||||
or aggregate["protocol_id"] != PROTOCOL_ID
|
||||
):
|
||||
raise RuntimeError("protocol mismatch")
|
||||
|
||||
formal = {}
|
||||
raw_files = {}
|
||||
for seed in SEEDS:
|
||||
name = f"formal-seed-{seed}.json"
|
||||
path = args.raw_dir / name
|
||||
run = load_canonical(path)
|
||||
if (
|
||||
run["run_kind"] != "formal"
|
||||
or run["seed"] != seed
|
||||
or not run["round06_equivalence"]["passed"]
|
||||
):
|
||||
raise RuntimeError(f"invalid formal run: {name}")
|
||||
formal[seed] = run
|
||||
raw_files[name] = {
|
||||
"file_sha256": file_sha256(path),
|
||||
"canonical_sha256": run["canonical_sha256_without_self"],
|
||||
}
|
||||
|
||||
replay_name = f"replay-seed-{SEEDS[0]}.json"
|
||||
replay_path = args.raw_dir / replay_name
|
||||
replay = load_canonical(replay_path)
|
||||
if replay["run_kind"] != "replay" or replay["seed"] != SEEDS[0]:
|
||||
raise RuntimeError("invalid replay")
|
||||
raw_files[replay_name] = {
|
||||
"file_sha256": file_sha256(replay_path),
|
||||
"canonical_sha256": replay["canonical_sha256_without_self"],
|
||||
}
|
||||
compare_payload = replay_payload(formal[SEEDS[0]])
|
||||
replay_exact = compare_payload == replay_payload(replay)
|
||||
if not replay_exact or not aggregate["replay"][
|
||||
"formal_seed1_exact_excluding_run_kind_and_timing"
|
||||
]:
|
||||
raise RuntimeError("replay exactness failed")
|
||||
replay_gate = {
|
||||
"passed": True,
|
||||
"excluded_fields": [
|
||||
"run_kind",
|
||||
"timing",
|
||||
"canonical_sha256_without_self",
|
||||
],
|
||||
"frozen_compare_sha256": canonical_sha256(compare_payload),
|
||||
}
|
||||
|
||||
reproduction = {
|
||||
"schema_version": 1,
|
||||
"protocol_id": PROTOCOL_ID,
|
||||
"raw_files": raw_files,
|
||||
"replay_gate": replay_gate,
|
||||
"artifacts": artifact_hashes(repo_root),
|
||||
"aggregate": {
|
||||
"file_sha256": file_sha256(args.aggregate),
|
||||
"canonical_sha256": aggregate[
|
||||
"canonical_sha256_without_self"
|
||||
],
|
||||
},
|
||||
"post_result_grok_review": {
|
||||
"session": "019fb19d-94a3-7231-9a63-3a1ef33a9892",
|
||||
"role": "read-only adversarial implementation audit; not an evidence source",
|
||||
"blocking_errors": 0,
|
||||
"localization_status_confirmed": True,
|
||||
},
|
||||
}
|
||||
write_canonical(args.reproduction_output, reproduction)
|
||||
|
||||
final_spectra = []
|
||||
for seed in SEEDS:
|
||||
final = formal[seed]["diagnostics"][-1]["local_matrix"]
|
||||
final_spectra.append(
|
||||
{
|
||||
"seed": seed,
|
||||
"modes": {
|
||||
mode: {
|
||||
"normalized": final[mode]["positions"][
|
||||
"post_mlp_state"
|
||||
]["reductions"]["element_rms"]["statistics"][
|
||||
"normalized"
|
||||
],
|
||||
"spike_contrast": final[mode]["positions"][
|
||||
"post_mlp_state"
|
||||
]["reductions"]["element_rms"]["statistics"][
|
||||
"spike_contrast"
|
||||
],
|
||||
"peak_normalized": final[mode]["positions"][
|
||||
"post_mlp_state"
|
||||
]["reductions"]["element_rms"]["statistics"][
|
||||
"peak_normalized"
|
||||
],
|
||||
"peak_layer": final[mode]["positions"][
|
||||
"post_mlp_state"
|
||||
]["reductions"]["element_rms"]["statistics"][
|
||||
"peak_layer"
|
||||
],
|
||||
"uniform_count": final[mode]["selector"][
|
||||
"uniform_count"
|
||||
],
|
||||
}
|
||||
for mode in manifest["matrix_modes"]
|
||||
},
|
||||
}
|
||||
)
|
||||
|
||||
compact = {
|
||||
"schema_version": 1,
|
||||
"protocol_id": PROTOCOL_ID,
|
||||
"study": {
|
||||
"identity": manifest["study_identity"],
|
||||
"seeds": list(SEEDS),
|
||||
"steps": manifest["training"]["steps"],
|
||||
"formal_target_bytes": (
|
||||
len(SEEDS)
|
||||
* manifest["training"]["target_bytes_per_cell"]
|
||||
),
|
||||
"total_target_bytes_with_replay": (
|
||||
(len(SEEDS) + 1)
|
||||
* manifest["training"]["target_bytes_per_cell"]
|
||||
),
|
||||
"modes": manifest["matrix_modes"],
|
||||
"spike_layers": manifest["primary_object"][
|
||||
"spike_layers_one_based"
|
||||
],
|
||||
"metrics": manifest["primary_object"]["metrics"],
|
||||
},
|
||||
"thresholds": manifest["thresholds"],
|
||||
"formulas": manifest["formulas"],
|
||||
"formal_cells": aggregate["formal_cells"],
|
||||
"scores": aggregate["scores"],
|
||||
"means": aggregate["means"],
|
||||
"gates": aggregate["gates"],
|
||||
"replay": {
|
||||
**aggregate["replay"],
|
||||
"frozen_compare_sha256": replay_gate[
|
||||
"frozen_compare_sha256"
|
||||
],
|
||||
},
|
||||
"final_spectra": final_spectra,
|
||||
"limitations": aggregate["limitations"],
|
||||
"hashes": {
|
||||
"aggregate_file_sha256": file_sha256(args.aggregate),
|
||||
"aggregate_canonical_sha256": aggregate[
|
||||
"canonical_sha256_without_self"
|
||||
],
|
||||
"reproduction_file_sha256": file_sha256(
|
||||
args.reproduction_output
|
||||
),
|
||||
"reproduction_canonical_sha256": reproduction[
|
||||
"canonical_sha256_without_self"
|
||||
],
|
||||
"manifest_file_sha256": file_sha256(args.manifest),
|
||||
},
|
||||
}
|
||||
write_canonical(args.compact_output, compact)
|
||||
print(
|
||||
json.dumps(
|
||||
{
|
||||
"reproduction": str(args.reproduction_output),
|
||||
"compact": str(args.compact_output),
|
||||
"raw_files": len(raw_files),
|
||||
"replay_exact": replay_exact,
|
||||
"localization": aggregate["gates"]["localization"][
|
||||
"status"
|
||||
],
|
||||
"compact_canonical_sha256": compact[
|
||||
"canonical_sha256_without_self"
|
||||
],
|
||||
},
|
||||
ensure_ascii=False,
|
||||
indent=2,
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,51 @@
|
||||
{
|
||||
"aggregate": {
|
||||
"canonical_sha256": "b86d119cd2f106e2cbee8a35760ed3244336a2fcfeb9178ea1e7dab13fc6f215",
|
||||
"file_sha256": "bb0ec9fce5b30d50ad5c50c4b95af7a892d614f205a2c20cfc2662125e10160e"
|
||||
},
|
||||
"artifacts": {
|
||||
"analyzer": "e0921562463e43d1ba6d47e4d23015107eb2df23579921088550b69afd47d02b",
|
||||
"manifest": "db01e92ef2cf0896212fcd529429bd94a344de0e1195db7f87b9a56dc3449139",
|
||||
"packager": "6a9ada0bc35b40475f45d7aca82877a93aeca667ed117c8e4417d924416c9196",
|
||||
"preresult_grok_review": "2da1b6bf1f455c4121a7a2c5cfe40e102327dabafc7e24dccc23ed0d00ac6d71",
|
||||
"protocol": "5ecc7ca92314ddb50aecf0cb50e115814c8983aa8bffb30e3634f7b3ce6dca1d",
|
||||
"runner": "b42879e242a2f2d54aa6a87a718aeac4cf4509eae42da2b14403656675a8b03d",
|
||||
"scoping": "670ca4edf31a4be1f54937d9c7a760dba7a96e1e820c38c6b10405e22b078fc8"
|
||||
},
|
||||
"canonical_sha256_without_self": "6f5d98fce6446fecc966dd2675f272f2c4f0c9a39a5741fabc4ffad6852ca7f4",
|
||||
"post_result_grok_review": {
|
||||
"blocking_errors": 0,
|
||||
"localization_status_confirmed": true,
|
||||
"role": "read-only adversarial implementation audit; not an evidence source",
|
||||
"session": "019fb19d-94a3-7231-9a63-3a1ef33a9892"
|
||||
},
|
||||
"protocol_id": "llm-atlas-k3-attnres-local-path-v1",
|
||||
"raw_files": {
|
||||
"formal-seed-2026073001.json": {
|
||||
"canonical_sha256": "f0a44f119836ed632c15880c3c2bb225173c0a05c50ea07abbe0e464ff407592",
|
||||
"file_sha256": "73d46ae443d3e5ae3fe839c1656cda758f5f41aaee5f220c971c3c39b8a8cc3f"
|
||||
},
|
||||
"formal-seed-2026073002.json": {
|
||||
"canonical_sha256": "b4629672b7d3b88a6be5525d2839e63e34fc9e4603ba1e7b98e957558da6da05",
|
||||
"file_sha256": "bca4674c746e35a035acde7d2094a9c3bd59a988b052cd3feb30eb66eb0ca60a"
|
||||
},
|
||||
"formal-seed-2026073003.json": {
|
||||
"canonical_sha256": "190b3deb06ae06caba287fce047b55cee613af6f1ebeb1661fcb53dd245abab0",
|
||||
"file_sha256": "712f349e7715fee71f8e4678dcde0619d01b8d6c3b5c34825c88b9af0bf2abbb"
|
||||
},
|
||||
"replay-seed-2026073001.json": {
|
||||
"canonical_sha256": "378df53ed9c89b2a4e0f3045b4c1e72754a7108d87fa9436e1db7aefa442e5eb",
|
||||
"file_sha256": "872aabd9285ac346b4016c23de10769ab83c4dc29e62d8bc8b3156af98228d4e"
|
||||
}
|
||||
},
|
||||
"replay_gate": {
|
||||
"excluded_fields": [
|
||||
"run_kind",
|
||||
"timing",
|
||||
"canonical_sha256_without_self"
|
||||
],
|
||||
"frozen_compare_sha256": "7dbd15ad03fbd357c5d91e159706d63b24703722f76c492ed1dc733535d6b9cf",
|
||||
"passed": true
|
||||
},
|
||||
"schema_version": 1
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,933 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Exact-replay Round 06 with preregistered local mixer-path interventions."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import importlib.util
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import platform
|
||||
import statistics
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
|
||||
|
||||
PROTOCOL_ID = "llm-atlas-k3-attnres-local-path-v1"
|
||||
PARENT_PROTOCOL_ID = "llm-atlas-k3-attnres-spike-path-v1"
|
||||
DATA_PROTOCOL_ID = "llm-atlas-k3-attnres-gradient-scale-v1"
|
||||
DEPTH = 32
|
||||
SEEDS = (2026073001, 2026073002, 2026073003)
|
||||
FORMAL_STEPS = 8000
|
||||
PARENT_DIAGNOSTIC_STEPS = (0, 100, 500, 2000, 4000, 8000)
|
||||
MATRIX_STEPS = (0, 8000)
|
||||
MATRIX_MODES = (
|
||||
"detached_learned",
|
||||
"uniform_group_6_only",
|
||||
"uniform_group_7_only",
|
||||
"uniform_groups_6_7_only",
|
||||
"uniform_group_6_attention_only",
|
||||
"uniform_group_6_mlp_only",
|
||||
"uniform_group_7_attention_only",
|
||||
"uniform_group_7_mlp_only",
|
||||
"uniform_output_only",
|
||||
"uniform_depth_all",
|
||||
"uniform_all",
|
||||
"uniform_except_group_6",
|
||||
"uniform_except_group_7",
|
||||
"uniform_except_groups_6_7",
|
||||
)
|
||||
TRAIN_BATCH_SIZE = 32
|
||||
VALIDATION_WINDOWS = 64
|
||||
EVAL_BATCH_SIZE = 8
|
||||
TIMING_WARMUP = 20
|
||||
|
||||
|
||||
def load_parent() -> Any:
|
||||
path = Path(__file__).resolve().parents[1] / "attnres_spike" / "train.py"
|
||||
spec = importlib.util.spec_from_file_location(
|
||||
"k3_attnres_spike_parent", path
|
||||
)
|
||||
if spec is None or spec.loader is None:
|
||||
raise RuntimeError(f"cannot import parent runner from {path}")
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
sys.modules[spec.name] = module
|
||||
spec.loader.exec_module(module)
|
||||
return module
|
||||
|
||||
|
||||
parent = load_parent()
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument(
|
||||
"--run-kind", choices=("smoke", "formal", "replay"), required=True
|
||||
)
|
||||
parser.add_argument("--seed", type=int, required=True)
|
||||
parser.add_argument("--cache-dir", type=Path, required=True)
|
||||
parser.add_argument("--data-manifest", type=Path, required=True)
|
||||
parser.add_argument("--parent-manifest", type=Path, required=True)
|
||||
parser.add_argument("--manifest", type=Path, required=True)
|
||||
parser.add_argument("--output", type=Path, required=True)
|
||||
args = parser.parse_args()
|
||||
if args.seed not in SEEDS:
|
||||
parser.error(f"seed must be one of {SEEDS}")
|
||||
if args.run_kind == "replay" and args.seed != SEEDS[0]:
|
||||
parser.error(f"replay seed must be {SEEDS[0]}")
|
||||
return args
|
||||
|
||||
|
||||
class LocalPathLanguageModel(parent.SpikeLanguageModel):
|
||||
"""Parent model with a parallel, audited per-mixer coefficient selector."""
|
||||
|
||||
def __init__(self, architecture: str, manifest: dict[str, Any]):
|
||||
super().__init__(architecture)
|
||||
selector = manifest["selector"]
|
||||
self.uniform_depth = {
|
||||
mode: frozenset(indices)
|
||||
for mode, indices in selector["uniform_depth_indices"].items()
|
||||
}
|
||||
self.uniform_output = selector["uniform_output"]
|
||||
self.expected_uniform_counts = selector["expected_uniform_counts"]
|
||||
self._active_selector_visits: list[dict[str, Any]] | None = None
|
||||
self.last_selector_visits: list[dict[str, Any]] | None = None
|
||||
|
||||
def forward(
|
||||
self,
|
||||
input_ids: torch.Tensor,
|
||||
capture: bool = False,
|
||||
mixer_backward_mode: str = "learned",
|
||||
) -> tuple[torch.Tensor, parent.SpikeTrace | None]:
|
||||
if not capture:
|
||||
if mixer_backward_mode != "learned":
|
||||
raise RuntimeError("training/evaluation cannot use an intervention")
|
||||
return super().forward(
|
||||
input_ids, capture=False, mixer_backward_mode="learned"
|
||||
)
|
||||
if mixer_backward_mode == "learned":
|
||||
self.last_selector_visits = None
|
||||
return super().forward(
|
||||
input_ids, capture=True, mixer_backward_mode="learned"
|
||||
)
|
||||
if mixer_backward_mode not in MATRIX_MODES:
|
||||
raise ValueError(f"unknown local matrix mode: {mixer_backward_mode}")
|
||||
self._active_selector_visits = []
|
||||
logits, trace = self._diagnostic_forward(
|
||||
input_ids, mixer_backward_mode
|
||||
)
|
||||
self.last_selector_visits = self._active_selector_visits
|
||||
self._active_selector_visits = None
|
||||
return logits, trace
|
||||
|
||||
def _mix(
|
||||
self,
|
||||
mixer: nn.Module,
|
||||
sources: list[torch.Tensor],
|
||||
labels: list[str],
|
||||
*,
|
||||
mode: str,
|
||||
mixer_index: int | None,
|
||||
layer: int | None,
|
||||
branch: str,
|
||||
group: int | None,
|
||||
offset: int | None,
|
||||
) -> tuple[torch.Tensor, dict[str, Any]]:
|
||||
if mode == "learned":
|
||||
return super()._mix(
|
||||
mixer,
|
||||
sources,
|
||||
labels,
|
||||
mode=mode,
|
||||
mixer_index=mixer_index,
|
||||
layer=layer,
|
||||
branch=branch,
|
||||
group=group,
|
||||
offset=offset,
|
||||
)
|
||||
if self._active_selector_visits is None:
|
||||
raise RuntimeError("local selector visit log is not active")
|
||||
|
||||
if mixer_index is None:
|
||||
if (
|
||||
layer is not None
|
||||
or group is not None
|
||||
or branch != "output"
|
||||
or offset is not None
|
||||
):
|
||||
raise RuntimeError("invalid output mixer identity")
|
||||
identity = {
|
||||
"kind": "output",
|
||||
"index": 64,
|
||||
"layer": None,
|
||||
"group": None,
|
||||
"branch": "output",
|
||||
"offset": None,
|
||||
}
|
||||
use_uniform = bool(self.uniform_output[mode])
|
||||
else:
|
||||
expected_layer = mixer_index // 2 + 1
|
||||
expected_branch = "attention" if mixer_index % 2 == 0 else "mlp"
|
||||
expected_group = (expected_layer - 1) // 4 + 1
|
||||
expected_offset = (expected_layer - 1) % 4 + 1
|
||||
if (
|
||||
not 0 <= mixer_index < 64
|
||||
or layer != expected_layer
|
||||
or branch != expected_branch
|
||||
or group != expected_group
|
||||
or offset != expected_offset
|
||||
):
|
||||
raise RuntimeError("invalid depth mixer identity")
|
||||
identity = {
|
||||
"kind": "depth",
|
||||
"index": mixer_index,
|
||||
"layer": layer,
|
||||
"group": group,
|
||||
"branch": branch,
|
||||
"offset": offset,
|
||||
}
|
||||
use_uniform = mixer_index in self.uniform_depth[mode]
|
||||
|
||||
parent_output, _ = mixer(sources, False)
|
||||
weights = parent.recompute_weights(mixer, sources)
|
||||
summary = parent.weight_summary(
|
||||
weights,
|
||||
labels,
|
||||
mixer_index=mixer_index,
|
||||
layer=layer,
|
||||
branch=branch,
|
||||
group=group,
|
||||
offset=offset,
|
||||
)
|
||||
backward_weights = (
|
||||
torch.full_like(weights, 1.0 / len(sources))
|
||||
if use_uniform
|
||||
else weights
|
||||
)
|
||||
values = torch.stack(sources, dim=0)
|
||||
routed = parent.RoutedSourceBackward.apply(
|
||||
values, parent_output, backward_weights
|
||||
)
|
||||
self._active_selector_visits.append(
|
||||
{
|
||||
**identity,
|
||||
"sources": len(sources),
|
||||
"use_uniform": use_uniform,
|
||||
"coefficient": (
|
||||
"uniform" if use_uniform else "detached_learned"
|
||||
),
|
||||
}
|
||||
)
|
||||
return routed, summary
|
||||
|
||||
|
||||
def expected_visit_identities() -> list[dict[str, Any]]:
|
||||
result = []
|
||||
for index in range(64):
|
||||
layer = index // 2 + 1
|
||||
result.append(
|
||||
{
|
||||
"kind": "depth",
|
||||
"index": index,
|
||||
"layer": layer,
|
||||
"group": (layer - 1) // 4 + 1,
|
||||
"branch": "attention" if index % 2 == 0 else "mlp",
|
||||
"offset": (layer - 1) % 4 + 1,
|
||||
}
|
||||
)
|
||||
result.append(
|
||||
{
|
||||
"kind": "output",
|
||||
"index": 64,
|
||||
"layer": None,
|
||||
"group": None,
|
||||
"branch": "output",
|
||||
"offset": None,
|
||||
}
|
||||
)
|
||||
return result
|
||||
|
||||
|
||||
def validate_selector_visits(
|
||||
mode: str,
|
||||
visits: list[dict[str, Any]] | None,
|
||||
manifest: dict[str, Any],
|
||||
) -> dict[str, Any]:
|
||||
if visits is None or len(visits) != 65:
|
||||
raise RuntimeError("selector visit count mismatch")
|
||||
identity_keys = ("kind", "index", "layer", "group", "branch", "offset")
|
||||
actual_identities = [
|
||||
{key: visit[key] for key in identity_keys} for visit in visits
|
||||
]
|
||||
expected_identities = expected_visit_identities()
|
||||
if actual_identities != expected_identities:
|
||||
raise RuntimeError("selector identity order mismatch")
|
||||
if len({(item["kind"], item["index"]) for item in visits}) != 65:
|
||||
raise RuntimeError("selector identities are not unique")
|
||||
|
||||
expected_depth = set(
|
||||
manifest["selector"]["uniform_depth_indices"][mode]
|
||||
)
|
||||
expected_output = manifest["selector"]["uniform_output"][mode]
|
||||
actual_depth = {
|
||||
item["index"]
|
||||
for item in visits
|
||||
if item["kind"] == "depth" and item["use_uniform"]
|
||||
}
|
||||
actual_output = visits[-1]["use_uniform"]
|
||||
if actual_depth != expected_depth or actual_output != expected_output:
|
||||
raise RuntimeError("selector exact-set mismatch")
|
||||
uniform_count = sum(int(item["use_uniform"]) for item in visits)
|
||||
expected_count = manifest["selector"]["expected_uniform_counts"][mode]
|
||||
if uniform_count != expected_count:
|
||||
raise RuntimeError("selector uniform census mismatch")
|
||||
return {
|
||||
"passed": True,
|
||||
"visit_count": len(visits),
|
||||
"identities_unique": True,
|
||||
"identity_order_sha256": parent.canonical_sha256(
|
||||
actual_identities
|
||||
),
|
||||
"uniform_indices": [
|
||||
item["index"] for item in visits if item["use_uniform"]
|
||||
],
|
||||
"uniform_count": uniform_count,
|
||||
"expected_uniform_count": expected_count,
|
||||
"visits_sha256": parent.canonical_sha256(visits),
|
||||
"visits": visits,
|
||||
}
|
||||
|
||||
|
||||
def run_local_diagnostic(
|
||||
model: LocalPathLanguageModel,
|
||||
corpus: Any,
|
||||
optimizer: torch.optim.Optimizer,
|
||||
manifest: dict[str, Any],
|
||||
*,
|
||||
mode: str,
|
||||
loss_scale: float = 1.0,
|
||||
) -> dict[str, Any]:
|
||||
result = parent.run_diagnostic(
|
||||
model, corpus, optimizer, mode=mode, loss_scale=loss_scale
|
||||
)
|
||||
result["selector"] = validate_selector_visits(
|
||||
mode, model.last_selector_visits, manifest
|
||||
)
|
||||
return result
|
||||
|
||||
|
||||
def forward_identity_gate(matrix: dict[str, Any]) -> dict[str, Any]:
|
||||
reference = matrix["detached_learned"]["forward"]
|
||||
comparisons = {}
|
||||
for mode in MATRIX_MODES[1:]:
|
||||
other = matrix[mode]["forward"]
|
||||
comparisons[mode] = {
|
||||
"logits_exact": (
|
||||
other["logits_sha256"] == reference["logits_sha256"]
|
||||
),
|
||||
"loss_exact": other["loss_nats"] == reference["loss_nats"],
|
||||
"activations_exact": (
|
||||
other["activation_sha256"]
|
||||
== reference["activation_sha256"]
|
||||
),
|
||||
"mixer_summaries_exact": (
|
||||
other["mixer_summary_sha256"]
|
||||
== reference["mixer_summary_sha256"]
|
||||
),
|
||||
}
|
||||
if not all(all(checks.values()) for checks in comparisons.values()):
|
||||
raise RuntimeError("local matrix forward identity failed")
|
||||
return {"passed": True, "comparisons": comparisons}
|
||||
|
||||
|
||||
def spectrum_agreement(
|
||||
left: dict[str, Any], right: dict[str, Any]
|
||||
) -> dict[str, Any]:
|
||||
checks = {}
|
||||
passed = True
|
||||
for position in parent.POSITIONS:
|
||||
left_metric = left["positions"][position]["reductions"][
|
||||
"element_rms"
|
||||
]
|
||||
right_metric = right["positions"][position]["reductions"][
|
||||
"element_rms"
|
||||
]
|
||||
raw_errors = [
|
||||
abs(a - b) / a
|
||||
for a, b in zip(
|
||||
left_metric["values"], right_metric["values"]
|
||||
)
|
||||
]
|
||||
normalized_errors = [
|
||||
abs(a - b)
|
||||
for a, b in zip(
|
||||
left_metric["statistics"]["normalized"],
|
||||
right_metric["statistics"]["normalized"],
|
||||
)
|
||||
]
|
||||
item_passed = (
|
||||
all(
|
||||
math.isfinite(value) and value > 0
|
||||
for value in left_metric["values"]
|
||||
)
|
||||
and max(raw_errors) <= parent.SPECTRUM_TOLERANCE
|
||||
and max(normalized_errors) <= parent.SPECTRUM_TOLERANCE
|
||||
)
|
||||
passed = passed and item_passed
|
||||
checks[position] = {
|
||||
"passed": item_passed,
|
||||
"max_raw_relative_error": max(raw_errors),
|
||||
"max_normalized_absolute_error": max(normalized_errors),
|
||||
}
|
||||
return {"passed": passed, "checks": checks}
|
||||
|
||||
|
||||
def initialization_negative_control(
|
||||
parent_learned: dict[str, Any], matrix: dict[str, Any]
|
||||
) -> dict[str, Any]:
|
||||
reference = matrix["detached_learned"]
|
||||
comparisons = {
|
||||
"parent_learned_vs_detached": spectrum_agreement(
|
||||
parent_learned, reference
|
||||
)
|
||||
}
|
||||
for mode in MATRIX_MODES[1:]:
|
||||
comparisons[mode] = spectrum_agreement(reference, matrix[mode])
|
||||
passed = all(item["passed"] for item in comparisons.values())
|
||||
if not passed:
|
||||
raise RuntimeError("initialization negative control failed")
|
||||
return {"passed": True, "comparisons": comparisons}
|
||||
|
||||
|
||||
def without_selector(result: dict[str, Any], rename: str | None = None) -> dict[str, Any]:
|
||||
cleaned = {key: value for key, value in result.items() if key != "selector"}
|
||||
if rename is not None:
|
||||
cleaned["mode"] = rename
|
||||
return cleaned
|
||||
|
||||
|
||||
def endpoint_exactness(
|
||||
matrix: dict[str, Any], parent_diagnostic: dict[str, Any]
|
||||
) -> dict[str, Any]:
|
||||
reference_exact = (
|
||||
without_selector(matrix["detached_learned"])
|
||||
== parent_diagnostic["modes"]["detached_learned"]
|
||||
)
|
||||
uniform_exact = (
|
||||
without_selector(
|
||||
matrix["uniform_all"], rename="uniform_value_backward"
|
||||
)
|
||||
== parent_diagnostic["modes"]["uniform_value_backward"]
|
||||
)
|
||||
checks = {
|
||||
"detached_learned_round06_exact": reference_exact,
|
||||
"uniform_all_round06_exact": uniform_exact,
|
||||
}
|
||||
if not all(checks.values()):
|
||||
raise RuntimeError(f"Round 06 endpoint exactness failed: {checks}")
|
||||
return {"passed": True, "checks": checks}
|
||||
|
||||
|
||||
def run_diagnostic_bundle(
|
||||
model: LocalPathLanguageModel,
|
||||
corpus: Any,
|
||||
optimizer: torch.optim.Optimizer,
|
||||
manifest: dict[str, Any],
|
||||
parent_diagnostic: dict[str, Any],
|
||||
step: int,
|
||||
) -> dict[str, Any]:
|
||||
parent_learned = parent.run_diagnostic(
|
||||
model, corpus, optimizer, mode="learned"
|
||||
)
|
||||
if parent_learned != parent_diagnostic["modes"]["learned"]:
|
||||
raise RuntimeError("parent learned diagnostic is not Round 06 exact")
|
||||
result: dict[str, Any] = {
|
||||
"step": step,
|
||||
"parent_learned": parent_learned,
|
||||
"parent_learned_round06_exact": True,
|
||||
"local_matrix": None,
|
||||
}
|
||||
if step not in MATRIX_STEPS:
|
||||
return result
|
||||
|
||||
matrix = {
|
||||
mode: run_local_diagnostic(
|
||||
model, corpus, optimizer, manifest, mode=mode
|
||||
)
|
||||
for mode in MATRIX_MODES
|
||||
}
|
||||
result["local_matrix"] = matrix
|
||||
result["forward_identity_gate"] = forward_identity_gate(matrix)
|
||||
result["endpoint_exactness"] = endpoint_exactness(
|
||||
matrix, parent_diagnostic
|
||||
)
|
||||
if step == 0:
|
||||
result["initialization_negative_control"] = (
|
||||
initialization_negative_control(parent_learned, matrix)
|
||||
)
|
||||
doubled = run_local_diagnostic(
|
||||
model,
|
||||
corpus,
|
||||
optimizer,
|
||||
manifest,
|
||||
mode="detached_learned",
|
||||
loss_scale=2.0,
|
||||
)
|
||||
result["loss_scale_gate"] = parent.loss_scale_gate(
|
||||
matrix["detached_learned"], doubled
|
||||
)
|
||||
model.zero_grad(set_to_none=True)
|
||||
return result
|
||||
|
||||
|
||||
def load_and_verify_inputs(
|
||||
args: argparse.Namespace,
|
||||
) -> tuple[dict[str, Any], dict[str, Any], dict[str, Any], Path]:
|
||||
manifest = json.loads(args.manifest.read_text())
|
||||
parent_manifest = json.loads(args.parent_manifest.read_text())
|
||||
data_manifest = json.loads(args.data_manifest.read_text())
|
||||
repo_root = Path(__file__).resolve().parents[3]
|
||||
|
||||
if manifest["protocol_id"] != PROTOCOL_ID:
|
||||
raise RuntimeError("Round 07 manifest protocol mismatch")
|
||||
if parent_manifest["protocol_id"] != PARENT_PROTOCOL_ID:
|
||||
raise RuntimeError("Round 06 parent manifest protocol mismatch")
|
||||
if data_manifest["protocol_id"] != DATA_PROTOCOL_ID:
|
||||
raise RuntimeError("data manifest protocol mismatch")
|
||||
if manifest["formal_seeds"] != list(SEEDS):
|
||||
raise RuntimeError("formal seed mismatch")
|
||||
if manifest["matrix_modes"] != list(MATRIX_MODES):
|
||||
raise RuntimeError("local matrix mode mismatch")
|
||||
if (
|
||||
manifest["training"]["parent_diagnostic_steps"]
|
||||
!= list(PARENT_DIAGNOSTIC_STEPS)
|
||||
or manifest["training"]["local_matrix_steps"] != list(MATRIX_STEPS)
|
||||
or manifest["training"]["steps"] != FORMAL_STEPS
|
||||
):
|
||||
raise RuntimeError("diagnostic/training schedule mismatch")
|
||||
|
||||
parent_artifacts = manifest["parent_artifacts"]
|
||||
if parent.file_sha256(args.parent_manifest) != parent_artifacts[
|
||||
"manifest_sha256"
|
||||
]:
|
||||
raise RuntimeError("Round 06 manifest physical hash mismatch")
|
||||
if parent.file_sha256(Path(parent.__file__)) != parent_artifacts[
|
||||
"runner_sha256"
|
||||
]:
|
||||
raise RuntimeError("Round 06 runner physical hash mismatch")
|
||||
for name in ("protocol", "scoping"):
|
||||
path = repo_root / parent_artifacts[f"{name}_path"]
|
||||
if parent.file_sha256(path) != parent_artifacts[f"{name}_sha256"]:
|
||||
raise RuntimeError(f"Round 06 {name} physical hash mismatch")
|
||||
for name in ("protocol", "scoping", "grok_review"):
|
||||
path = repo_root / manifest["current_artifacts"][f"{name}_path"]
|
||||
if parent.file_sha256(path) != manifest["current_artifacts"][
|
||||
f"{name}_sha256"
|
||||
]:
|
||||
raise RuntimeError(f"Round 07 {name} physical hash mismatch")
|
||||
for key in (
|
||||
"formal_schedule_sha256",
|
||||
"validation_tensor_sha256",
|
||||
"diagnostic_tensor_sha256",
|
||||
):
|
||||
if (
|
||||
data_manifest["windows"][key]
|
||||
!= parent_artifacts[key]
|
||||
or parent_manifest["parent_artifacts"][key]
|
||||
!= parent_artifacts[key]
|
||||
):
|
||||
raise RuntimeError(f"frozen data hash mismatch: {key}")
|
||||
|
||||
parent_raw_path = (
|
||||
repo_root
|
||||
/ "experiments"
|
||||
/ "k3"
|
||||
/ "attnres_spike"
|
||||
/ "results"
|
||||
/ "raw"
|
||||
/ f"formal-seed-{args.seed}.json"
|
||||
)
|
||||
expected = manifest["round06_expected"][str(args.seed)]
|
||||
if parent.file_sha256(parent_raw_path) != expected["raw_file_sha256"]:
|
||||
raise RuntimeError("Round 06 raw physical hash mismatch")
|
||||
parent_raw = json.loads(parent_raw_path.read_text())
|
||||
if (
|
||||
parent_raw["canonical_sha256_without_self"]
|
||||
!= expected["canonical_sha256"]
|
||||
or parent_raw["hashes"]["final_model_state"]
|
||||
!= expected["final_model_state"]
|
||||
or parent_raw["hashes"]["final_optimizer_state"]
|
||||
!= expected["final_optimizer_state"]
|
||||
):
|
||||
raise RuntimeError("Round 06 raw expected-state mismatch")
|
||||
return manifest, data_manifest, parent_raw, repo_root
|
||||
|
||||
|
||||
def frozen_training_compare(
|
||||
result: dict[str, Any], parent_raw: dict[str, Any]
|
||||
) -> dict[str, Any]:
|
||||
checks = {
|
||||
"final_model_state": (
|
||||
result["hashes"]["final_model_state"]
|
||||
== parent_raw["hashes"]["final_model_state"]
|
||||
),
|
||||
"final_optimizer_state": (
|
||||
result["hashes"]["final_optimizer_state"]
|
||||
== parent_raw["hashes"]["final_optimizer_state"]
|
||||
),
|
||||
"evaluations": result["evaluations"] == parent_raw["evaluations"],
|
||||
"training_history": (
|
||||
result["training_history"] == parent_raw["training_history"]
|
||||
),
|
||||
}
|
||||
parent_diagnostics_exact = []
|
||||
endpoint_exact = []
|
||||
for new, old in zip(result["diagnostics"], parent_raw["diagnostics"]):
|
||||
parent_diagnostics_exact.append(
|
||||
new["step"] == old["step"]
|
||||
and new["parent_learned"] == old["modes"]["learned"]
|
||||
)
|
||||
if new["step"] in MATRIX_STEPS:
|
||||
endpoint_exact.append(new["endpoint_exactness"]["passed"])
|
||||
checks["parent_learned_diagnostics"] = all(parent_diagnostics_exact)
|
||||
checks["round06_endpoints"] = len(endpoint_exact) == 2 and all(
|
||||
endpoint_exact
|
||||
)
|
||||
if not all(checks.values()):
|
||||
raise RuntimeError(f"Round 06 training equivalence failed: {checks}")
|
||||
return {"passed": True, "checks": checks}
|
||||
|
||||
|
||||
def main() -> None:
|
||||
args = parse_args()
|
||||
if not torch.cuda.is_available():
|
||||
raise RuntimeError("CUDA is required")
|
||||
if os.environ.get("CUBLAS_WORKSPACE_CONFIG") != ":4096:8":
|
||||
raise RuntimeError("CUBLAS_WORKSPACE_CONFIG must be :4096:8")
|
||||
manifest, data_manifest, parent_raw, repo_root = load_and_verify_inputs(args)
|
||||
parent.parent.configure_round04_globals(DEPTH)
|
||||
parent.parent.configure_determinism(args.seed)
|
||||
corpus = parent.parent.round04.ByteCorpus(
|
||||
args.cache_dir, data_manifest, torch.device("cuda")
|
||||
)
|
||||
model = LocalPathLanguageModel("block", manifest).to(
|
||||
torch.device("cuda")
|
||||
)
|
||||
|
||||
initial_public_hash = parent.parent.named_state_hash(
|
||||
model, include_mixers=False
|
||||
)
|
||||
initial_mixer_hash = parent.parent.named_state_hash(
|
||||
model, include_mixers=True
|
||||
)
|
||||
public_structure_hash, public_tensors, public_elements = (
|
||||
parent.parent.state_structure_hash(model, include_mixers=False)
|
||||
)
|
||||
input_gate_hashes = parent.parent.model_input_gate_hashes(
|
||||
corpus, data_manifest, args.seed, TRAIN_BATCH_SIZE
|
||||
)
|
||||
|
||||
decay_parameters: list[nn.Parameter] = []
|
||||
no_decay_parameters: list[nn.Parameter] = []
|
||||
for parameter in model.parameters():
|
||||
target = decay_parameters if parameter.ndim >= 2 else no_decay_parameters
|
||||
target.append(parameter)
|
||||
optimizer = torch.optim.AdamW(
|
||||
[
|
||||
{
|
||||
"params": decay_parameters,
|
||||
"weight_decay": parent.parent.WEIGHT_DECAY,
|
||||
},
|
||||
{"params": no_decay_parameters, "weight_decay": 0.0},
|
||||
],
|
||||
lr=parent.parent.PEAK_LR,
|
||||
betas=parent.parent.BETAS,
|
||||
eps=parent.parent.ADAM_EPS,
|
||||
)
|
||||
|
||||
parent_by_step = {
|
||||
item["step"]: item for item in parent_raw["diagnostics"]
|
||||
}
|
||||
evaluations = [
|
||||
{
|
||||
"step": 0,
|
||||
**parent.parent.evaluate(
|
||||
model, corpus, VALIDATION_WINDOWS, EVAL_BATCH_SIZE
|
||||
),
|
||||
}
|
||||
]
|
||||
diagnostics = [
|
||||
run_diagnostic_bundle(
|
||||
model,
|
||||
corpus,
|
||||
optimizer,
|
||||
manifest,
|
||||
parent_by_step[0],
|
||||
0,
|
||||
)
|
||||
]
|
||||
print(
|
||||
json.dumps(
|
||||
{
|
||||
"event": "local_matrix",
|
||||
"step": 0,
|
||||
"seed": args.seed,
|
||||
"modes": len(MATRIX_MODES),
|
||||
"endpoint_exact": diagnostics[0][
|
||||
"endpoint_exactness"
|
||||
]["passed"],
|
||||
},
|
||||
sort_keys=True,
|
||||
),
|
||||
flush=True,
|
||||
)
|
||||
|
||||
steps = 0 if args.run_kind == "smoke" else FORMAL_STEPS
|
||||
training_history: list[dict[str, float | int]] = []
|
||||
step_times: list[float] = []
|
||||
if steps:
|
||||
model.train()
|
||||
for step in range(1, steps + 1):
|
||||
lr = parent.parent.learning_rate(step, steps)
|
||||
for group in optimizer.param_groups:
|
||||
group["lr"] = lr
|
||||
inputs, targets = corpus.training_batch(
|
||||
args.seed, step, TRAIN_BATCH_SIZE
|
||||
)
|
||||
optimizer.zero_grad(set_to_none=True)
|
||||
torch.cuda.synchronize()
|
||||
started = time.perf_counter()
|
||||
with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
|
||||
logits, trace = model(inputs)
|
||||
if trace is not None:
|
||||
raise RuntimeError(
|
||||
"training unexpectedly captured a trace"
|
||||
)
|
||||
loss = parent.parent.cross_entropy(logits, targets)
|
||||
if not torch.isfinite(loss):
|
||||
raise RuntimeError(f"non-finite loss at step {step}")
|
||||
loss.backward()
|
||||
unclipped_norm = torch.nn.utils.clip_grad_norm_(
|
||||
model.parameters(), parent.parent.GRAD_CLIP
|
||||
)
|
||||
optimizer.step()
|
||||
torch.cuda.synchronize()
|
||||
elapsed_ms = (time.perf_counter() - started) * 1000
|
||||
if step == TIMING_WARMUP:
|
||||
torch.cuda.reset_peak_memory_stats()
|
||||
elif step > TIMING_WARMUP:
|
||||
step_times.append(elapsed_ms)
|
||||
if step == 1 or step % 10 == 0 or step == steps:
|
||||
training_history.append(
|
||||
{
|
||||
"step": step,
|
||||
"loss_nats": loss.detach().cpu().item(),
|
||||
"bits_per_byte": (
|
||||
loss.detach().cpu().item() / math.log(2)
|
||||
),
|
||||
"learning_rate": lr,
|
||||
"unclipped_grad_norm": float(
|
||||
unclipped_norm.detach().cpu()
|
||||
),
|
||||
}
|
||||
)
|
||||
if step in PARENT_DIAGNOSTIC_STEPS:
|
||||
evaluations.append(
|
||||
{
|
||||
"step": step,
|
||||
**parent.parent.evaluate(
|
||||
model,
|
||||
corpus,
|
||||
VALIDATION_WINDOWS,
|
||||
EVAL_BATCH_SIZE,
|
||||
),
|
||||
}
|
||||
)
|
||||
diagnostic = run_diagnostic_bundle(
|
||||
model,
|
||||
corpus,
|
||||
optimizer,
|
||||
manifest,
|
||||
parent_by_step[step],
|
||||
step,
|
||||
)
|
||||
diagnostics.append(diagnostic)
|
||||
event = {
|
||||
"event": "diagnostic",
|
||||
"step": step,
|
||||
"seed": args.seed,
|
||||
"validation_bpc": evaluations[-1]["bits_per_byte"],
|
||||
"parent_exact": diagnostic[
|
||||
"parent_learned_round06_exact"
|
||||
],
|
||||
}
|
||||
if diagnostic["local_matrix"] is not None:
|
||||
event["modes"] = len(MATRIX_MODES)
|
||||
event["endpoint_exact"] = diagnostic[
|
||||
"endpoint_exactness"
|
||||
]["passed"]
|
||||
print(json.dumps(event, sort_keys=True), flush=True)
|
||||
model.train()
|
||||
|
||||
timing = {
|
||||
"warmup_steps_excluded": TIMING_WARMUP,
|
||||
"measured_steps": len(step_times),
|
||||
"mean_ms": (
|
||||
statistics.fmean(step_times) if step_times else None
|
||||
),
|
||||
"median_ms": (
|
||||
statistics.median(step_times) if step_times else None
|
||||
),
|
||||
"p95_ms": (
|
||||
parent.type7_quantile(torch.tensor(sorted(step_times)), 0.95)
|
||||
if step_times
|
||||
else None
|
||||
),
|
||||
"peak_allocated_bytes": torch.cuda.max_memory_allocated(),
|
||||
"peak_reserved_bytes": torch.cuda.max_memory_reserved(),
|
||||
}
|
||||
result = {
|
||||
"schema_version": 1,
|
||||
"protocol_id": PROTOCOL_ID,
|
||||
"parent_protocol_id": PARENT_PROTOCOL_ID,
|
||||
"run_kind": args.run_kind,
|
||||
"architecture": "block",
|
||||
"depth": DEPTH,
|
||||
"seed": args.seed,
|
||||
"steps": steps,
|
||||
"batch_size": TRAIN_BATCH_SIZE,
|
||||
"target_bytes_seen": (
|
||||
steps * TRAIN_BATCH_SIZE * parent.CONTEXT
|
||||
),
|
||||
"manifest": {
|
||||
"path": str(args.manifest),
|
||||
"file_sha256": parent.file_sha256(args.manifest),
|
||||
"parent_path": str(args.parent_manifest),
|
||||
"parent_file_sha256": parent.file_sha256(
|
||||
args.parent_manifest
|
||||
),
|
||||
"data_path": str(args.data_manifest),
|
||||
"data_file_sha256": parent.file_sha256(args.data_manifest),
|
||||
"formal_schedule_sha256": data_manifest["windows"][
|
||||
"formal_schedule_sha256"
|
||||
],
|
||||
"validation_tensor_sha256": data_manifest["windows"][
|
||||
"validation_tensor_sha256"
|
||||
],
|
||||
"diagnostic_tensor_sha256": data_manifest["windows"][
|
||||
"diagnostic_tensor_sha256"
|
||||
],
|
||||
"input_gate_tensor_hashes": input_gate_hashes,
|
||||
"selector_contract_sha256": parent.canonical_sha256(
|
||||
manifest["selector"]
|
||||
),
|
||||
},
|
||||
"model": {
|
||||
"layers": DEPTH,
|
||||
"aggregation_groups": 8,
|
||||
"blocks_per_group": 4,
|
||||
"d_model": parent.parent.round04.D_MODEL,
|
||||
"heads": parent.parent.round04.HEADS,
|
||||
"d_ff": parent.parent.round04.D_FF,
|
||||
"parameters": parent.parent.parameter_inventory(model),
|
||||
},
|
||||
"hashes": {
|
||||
"initial_public_parameter_structure": public_structure_hash,
|
||||
"initial_public_parameter_tensors": public_tensors,
|
||||
"initial_public_parameter_elements": public_elements,
|
||||
"initial_public_parameters": initial_public_hash,
|
||||
"initial_mixer_parameters": initial_mixer_hash,
|
||||
"final_public_parameters": parent.parent.named_state_hash(
|
||||
model, include_mixers=False
|
||||
),
|
||||
"final_mixer_parameters": parent.parent.named_state_hash(
|
||||
model, include_mixers=True
|
||||
),
|
||||
"final_model_state": parent.parent.named_state_hash(
|
||||
model, include_mixers=None
|
||||
),
|
||||
"final_optimizer_state": parent.parent.recursive_state_hash(
|
||||
optimizer.state_dict()
|
||||
),
|
||||
},
|
||||
"evaluations": evaluations,
|
||||
"diagnostics": diagnostics,
|
||||
"training_history": training_history,
|
||||
"timing": timing,
|
||||
"environment": {
|
||||
"python": platform.python_version(),
|
||||
"torch": torch.__version__,
|
||||
"cuda": torch.version.cuda,
|
||||
"gpu": torch.cuda.get_device_name(0),
|
||||
"compute_capability": list(
|
||||
torch.cuda.get_device_capability(0)
|
||||
),
|
||||
"cublas_workspace_config": os.environ[
|
||||
"CUBLAS_WORKSPACE_CONFIG"
|
||||
],
|
||||
"deterministic_algorithms": (
|
||||
torch.are_deterministic_algorithms_enabled()
|
||||
),
|
||||
"autocast": "cuda-bfloat16-forward-fp32-cross-entropy",
|
||||
"compile": False,
|
||||
},
|
||||
"artifacts": {
|
||||
"runner_sha256": parent.file_sha256(Path(__file__)),
|
||||
"protocol_sha256": parent.file_sha256(
|
||||
repo_root
|
||||
/ "research"
|
||||
/ "K3_ATTNRES_LOCAL_PATH_PROTOCOL.md"
|
||||
),
|
||||
"scoping_sha256": parent.file_sha256(
|
||||
repo_root
|
||||
/ "research"
|
||||
/ "K3_ATTNRES_LOCAL_PATH_SCOPING.md"
|
||||
),
|
||||
"grok_review_sha256": parent.file_sha256(
|
||||
repo_root
|
||||
/ "research"
|
||||
/ "K3_ATTNRES_LOCAL_PATH_GROK_REVIEW.md"
|
||||
),
|
||||
},
|
||||
}
|
||||
result["round06_equivalence"] = (
|
||||
frozen_training_compare(result, parent_raw) if steps else None
|
||||
)
|
||||
result["canonical_sha256_without_self"] = parent.canonical_sha256(
|
||||
result
|
||||
)
|
||||
args.output.parent.mkdir(parents=True, exist_ok=True)
|
||||
temporary = args.output.with_suffix(args.output.suffix + ".tmp")
|
||||
temporary.write_text(
|
||||
json.dumps(result, ensure_ascii=False, indent=2, sort_keys=True)
|
||||
+ "\n"
|
||||
)
|
||||
os.replace(temporary, args.output)
|
||||
print(
|
||||
json.dumps(
|
||||
{
|
||||
"output": str(args.output),
|
||||
"run_kind": args.run_kind,
|
||||
"seed": args.seed,
|
||||
"steps": steps,
|
||||
"final_bpc": evaluations[-1]["bits_per_byte"],
|
||||
"canonical_sha256": result[
|
||||
"canonical_sha256_without_self"
|
||||
],
|
||||
"timing": timing,
|
||||
},
|
||||
ensure_ascii=False,
|
||||
indent=2,
|
||||
),
|
||||
flush=True,
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -22,6 +22,7 @@
|
||||
"check:data:k3-attnres": "node scripts/check-k3-attnres-data.mjs",
|
||||
"check:data:k3-attnres-gradient": "node scripts/check-k3-attnres-gradient-data.mjs",
|
||||
"check:data:k3-attnres-spike": "node scripts/check-k3-attnres-spike-data.mjs",
|
||||
"check:data:k3-attnres-local-path": "node scripts/check-k3-attnres-local-path-data.mjs",
|
||||
"check:site": "node scripts/check-site.mjs",
|
||||
"check:moe-browser": "node scripts/check-moe-browser.mjs",
|
||||
"check:reasoning-browser": "node scripts/check-reasoning-browser.mjs",
|
||||
@@ -44,6 +45,7 @@
|
||||
"check:k3-attnres-browser": "node scripts/check-k3-attnres-browser.mjs",
|
||||
"check:k3-attnres-gradient-browser": "node scripts/check-k3-attnres-gradient-browser.mjs",
|
||||
"check:k3-attnres-spike-browser": "node scripts/check-k3-attnres-spike-browser.mjs",
|
||||
"check:k3-attnres-local-path-browser": "node scripts/check-k3-attnres-local-path-browser.mjs",
|
||||
"check:k3-browser": "node scripts/check-k3-browser.mjs"
|
||||
},
|
||||
"dependencies": {
|
||||
|
||||
@@ -0,0 +1,302 @@
|
||||
# K3 Attention Residuals 局部 mixer 路径:Round 07 结果审计
|
||||
|
||||
研究日期:2026-07-30
|
||||
协议:`llm-atlas-k3-attnres-local-path-v1`
|
||||
预注册 commit:`6911efc`
|
||||
结果前 runner / analyzer commit:`39a9ad6`
|
||||
研究身份:**受 Round 05 / 06 启发的定向 reduced-model mechanism probe**
|
||||
|
||||
## 0. 一句话结论
|
||||
|
||||
> group 6+7 的 16 个 depth mixers 在 learned 背景上的局部 uniform intervention,
|
||||
> 足以复现全局 log-gap reduction 的至少一半;但从 all-uniform 背景只恢复这 16 个
|
||||
> mixer 时,contrast 恢复超过一半,peak 只恢复约 35.5%–41.3%。因此本轮得到强的
|
||||
> **one-sided evidence**,但没有通过预注册的双向 localization 门。
|
||||
|
||||
这不是一句保守套话,而是协议第 13 节的直接判定:
|
||||
|
||||
```text
|
||||
groups 6+7 sufficiency = PASS 6 / 6
|
||||
groups 6+7 restoration = FAIL 3 / 6
|
||||
localization = NOT ESTABLISHED
|
||||
```
|
||||
|
||||
## 1. 运行与输入闸门
|
||||
|
||||
正式网格:
|
||||
|
||||
| run | seed | steps | target bytes | final validation BPC |
|
||||
|---|---:|---:|---:|---:|
|
||||
| formal | 2026073001 | 8,000 | 65,536,000 | 1.7123525941 |
|
||||
| formal | 2026073002 | 8,000 | 65,536,000 | 1.7093240656 |
|
||||
| formal | 2026073003 | 8,000 | 65,536,000 | 1.7030966813 |
|
||||
| replay | 2026073001 | 8,000 | 65,536,000 | 1.7123525941 |
|
||||
|
||||
正式三格合计 196,608,000 target bytes,含 replay 为 262,144,000。
|
||||
|
||||
全部通过:
|
||||
|
||||
- 三 seed final model-state 与 Round 06 exact;
|
||||
- 三 seed final optimizer-state 与 Round 06 exact;
|
||||
- 六个 validation BPC、training history、六个 parent `learned` diagnostics exact;
|
||||
- step 0 / 8,000 的 `detached_learned` 与 Round 06 同名 endpoint exact;
|
||||
- step 0 / 8,000 的 `uniform_all` 与 Round 06
|
||||
`uniform_value_backward` endpoint exact;
|
||||
- 14 modes 的 logits、loss、六位置 activations、父 mixer summaries exact;
|
||||
- 65 次 selector visits 的 identity、顺序、唯一性、exact mask 与 census 全部通过;
|
||||
- step-0 14-mode negative control、parent learned-vs-detached control、
|
||||
`loss ×1 / ×2` scale gate 全部通过;
|
||||
- seed-1 从初始化完整 replay exact。
|
||||
|
||||
因此后续差异来自同一 forward state 上的预注册 backward coefficient masks,不是不同训练
|
||||
状态、batch、loss、activation 或 selector 漂移。
|
||||
|
||||
## 2. 全局端点先复现
|
||||
|
||||
最终 `post_mlp_state / element_rms`:
|
||||
|
||||
| metric | detached learned(3-seed mean) | uniform all(3-seed mean) | Round 06 per-seed mean relative drop |
|
||||
|---|---:|---:|---:|
|
||||
| spike contrast | 2.8065 | 0.7859 | 70.2% |
|
||||
| peak / layer mean | 3.0651 | 1.8964 | 37.0% |
|
||||
|
||||
每个 seed、两个指标的 `G_X = ln(X_ref / X_uniform_all)` 都严格为正,raw relative
|
||||
drop 6 / 6 超过 20%。global gap gate 完整成立。
|
||||
|
||||
这一步重要,因为局部 score 的分母不是任意“改善空间”,而是同一 seed、同一指标的
|
||||
实测 global log gap。若任一 global gap 不成立,本轮主判定就必须停止;实际没有触发该
|
||||
停止规则。
|
||||
|
||||
## 3. groups 6+7:充分性很强
|
||||
|
||||
只把 mixer indices 40–55 改成 uniform,其他 49 个 mixer 保持 detached learned:
|
||||
|
||||
| seed | `S_contrast` | `S_peak` | 50% × two metrics |
|
||||
|---:|---:|---:|---:|
|
||||
| 2026073001 | 0.697 | 1.817 | PASS |
|
||||
| 2026073002 | 0.676 | 1.783 | PASS |
|
||||
| 2026073003 | 0.658 | 1.501 | PASS |
|
||||
| **mean** | **0.677** | **1.700** | **6 / 6** |
|
||||
|
||||
raw metric 的三 seed mean 也从 reference 的 `2.8065 / 3.0651` 变为
|
||||
`1.1737 / 1.3450`。
|
||||
|
||||
`S_peak > 1` 不是 170% 因果贡献。它只表示在 log ratio 上,局部 uniform groups 6+7
|
||||
把 peak 推得比 all-65 uniform endpoint 还低。这是很直接的 non-additivity / interaction
|
||||
信号,也是协议坚持“不裁剪 score 到 [0,1]”的原因。
|
||||
|
||||
允许结论:
|
||||
|
||||
> groups 6+7 在本 diagnostic 中足以复现至少一半 global log-gap reduction。
|
||||
|
||||
禁止结论:
|
||||
|
||||
- “这 16 个 mixer 解释了 67.7% / 170.0% 的尖峰”;
|
||||
- “剩下 49 个 mixer 只贡献 32.3% / −70.0%”;
|
||||
- “group 6+7 是唯一原因”。
|
||||
|
||||
## 4. 反向 restoration 没有给出同样答案
|
||||
|
||||
从 all-uniform 背景出发,只把 groups 6+7 恢复为 detached-learned coefficients;
|
||||
其他 49 个 mixer 仍为 uniform:
|
||||
|
||||
| seed | `R_contrast` | `R_peak` | 50% × two metrics |
|
||||
|---:|---:|---:|---:|
|
||||
| 2026073001 | 0.649 PASS | 0.372 FAIL | FAIL |
|
||||
| 2026073002 | 0.621 PASS | 0.355 FAIL | FAIL |
|
||||
| 2026073003 | 0.680 PASS | 0.413 FAIL | FAIL |
|
||||
| **mean** | **0.650** | **0.380** | **3 / 6** |
|
||||
|
||||
raw metric 的三 seed mean 从 all-uniform 的 `0.7859 / 1.8964` 恢复为
|
||||
`1.7693 / 2.2658`。contrast 明显朝 reference 回升,但 peak 的 log-gap recovery
|
||||
没有一个 seed 达到 50%。
|
||||
|
||||
这说明同一个 scope 的作用强烈依赖其他 mixer 处于 learned 还是 uniform 背景:
|
||||
|
||||
- learned 背景中 uniformize groups 6+7,足以大幅压低 contrast 与 peak;
|
||||
- uniform 背景中 restore groups 6+7,足以恢复 contrast,却不足以恢复 peak;
|
||||
- 两个方向不对称,不能用单侧 sufficiency 替代双向 localization。
|
||||
|
||||
因此正式 verdict 是:
|
||||
|
||||
```text
|
||||
one_sided_evidence_localization_not_established
|
||||
```
|
||||
|
||||
不是 “almost passed”,也不因 `R_peak` mean 约 0.38 而软化 0.50 阈值。
|
||||
|
||||
## 5. 单 group 结果:同样显示交互
|
||||
|
||||
### 5.1 sufficiency
|
||||
|
||||
| scope | mean `S_contrast` | mean `S_peak` | 20% gate |
|
||||
|---|---:|---:|---:|
|
||||
| group 6 | 0.281 | 1.044 | PASS 6 / 6 |
|
||||
| group 7 | 0.438 | 0.843 | PASS 6 / 6 |
|
||||
|
||||
两个单 group 都在两个指标、三个 seed 通过 material local sufficiency。
|
||||
|
||||
但:
|
||||
|
||||
```text
|
||||
S(group6) + S(group7) ≠ S(groups6+7)
|
||||
```
|
||||
|
||||
尤其 peak 上,两个单 group 与联合 scope 都可能超过 global endpoint,不能按 mixer
|
||||
数量或 score 相加做贡献账。
|
||||
|
||||
### 5.2 restoration
|
||||
|
||||
| restored scope | mean `R_contrast` | mean `R_peak` | 20% gate |
|
||||
|---|---:|---:|---:|
|
||||
| group 6 | +0.248 | −0.146 | MIXED / FAIL 3 / 6 |
|
||||
| group 7 | +0.409 | −0.191 | MIXED / FAIL 3 / 6 |
|
||||
|
||||
恢复单个 group 时,contrast 在三 seed 都超过 20%,peak 却在三 seed 全为负:相对
|
||||
all-uniform,恢复一个 group 的 learned coefficients 反而让最高层 / 均值更低。
|
||||
|
||||
这不是 “group 没作用”,而是 effect direction 随 metric 与背景改变。它进一步反对
|
||||
简单、可加的局部归因故事。
|
||||
|
||||
## 6. attention vs MLP:只有 group 7 过闸
|
||||
|
||||
### group 6
|
||||
|
||||
MLP-only 在 6 个 branch cells 中赢 5 个;seed 2026073003 的 contrast
|
||||
`S=0.177 < 0.20`。因此:
|
||||
|
||||
```text
|
||||
group 6 branch dominance = NOT ESTABLISHED
|
||||
```
|
||||
|
||||
不能因为 margin 大、均值高,忽略 material threshold 的单格失败。
|
||||
|
||||
### group 7
|
||||
|
||||
| branch | mean `S_contrast` | mean `S_peak` |
|
||||
|---|---:|---:|
|
||||
| attention-only | 0.015 | 0.026 |
|
||||
| MLP-only | 0.426 | 0.823 |
|
||||
|
||||
MLP-only 自身 material,且在两个指标、三个 seed 都比 attention-only 高至少
|
||||
15 percentage points,因此:
|
||||
|
||||
```text
|
||||
group 7 MLP branch-dominant at the preregistered margin
|
||||
```
|
||||
|
||||
这是协议第 14 节的**次级、sufficiency-only、探索性**判定;没有 branch-level
|
||||
restoration,不得升级为第 13 节的双向 localization。
|
||||
|
||||
## 7. output / depth controls
|
||||
|
||||
| control | mean `S_contrast` | mean `S_peak` | verdict |
|
||||
|---|---:|---:|---|
|
||||
| output-only(1 mixer) | 0.131 | 0.187 | 0 / 6 at 50% |
|
||||
| all-depth(64 mixers) | 0.949 | 1.145 | near / beyond global endpoint |
|
||||
|
||||
output mixer 单独无法解释 global gap 的一半。all-depth 已复现绝大多数 contrast gap,
|
||||
peak 甚至超过 all-65 endpoint;把 output 与 depth scores 相加会产生负
|
||||
`interaction_residual`。该 residual 只是 bookkeeping,不预期为 0,不是统计交互检验。
|
||||
|
||||
## 8. 32-layer 谱的直观变化
|
||||
|
||||
三个 seed 的 reference peak 都在 layer 21;all-uniform peak 都迁到 layer 2。
|
||||
|
||||
groups 6+7 only:
|
||||
|
||||
- seed 1 peak → layer 5;
|
||||
- seed 2 peak → layer 25;
|
||||
- seed 3 peak → layer 6。
|
||||
|
||||
restore groups 6+7 on uniform background:
|
||||
|
||||
- 三 seed peak 都回到 layer 21;
|
||||
- 但 peak / mean 的恢复比例仍只有 0.355–0.413。
|
||||
|
||||
“peak layer 回来了”与“peak 强度恢复超过一半”不是同一判据。网站会同时展示谱与
|
||||
预注册 score,避免只凭最高点位置讲故事。
|
||||
|
||||
## 9. replay 与 artifact 链
|
||||
|
||||
seed 2026073001 从初始化完整重跑。排除 `run_kind`、timing 与 self canonical hash 后:
|
||||
|
||||
```text
|
||||
formal seed1 == replay
|
||||
compare SHA-256 = 7dbd15ad03fbd357c5d91e159706d63b24703722f76c492ed1dc733535d6b9cf
|
||||
```
|
||||
|
||||
它覆盖训练状态、六 checkpoints、14-mode 两端矩阵、六位置 gradient reductions、
|
||||
selector visits 与所有 gates,不只是 final BPC。
|
||||
|
||||
冻结物理 hashes:
|
||||
|
||||
| artifact | SHA-256 |
|
||||
|---|---|
|
||||
| manifest | `db01e92e…9139` |
|
||||
| runner | `b42879e2…b03d` |
|
||||
| analyzer | `e0921562…d02b` |
|
||||
| packager | `6a9ada0b…9196` |
|
||||
| aggregate | `bb0ec9fc…160e` |
|
||||
| compact | `3bb6c158…9c08` |
|
||||
| reproduction | `524a6883…5db` |
|
||||
|
||||
canonical hashes:
|
||||
|
||||
```text
|
||||
aggregate b86d119cd2f106e2cbee8a35760ed3244336a2fcfeb9178ea1e7dab13fc6f215
|
||||
compact 2aff9288f52d3d41bb1f59c64d9a07518ad3e2120c61615478087b24aaabd835
|
||||
reproduction 6f5d98fce6446fecc966dd2675f272f2c4f0c9a39a5741fabc4ffad6852ca7f4
|
||||
```
|
||||
|
||||
## 10. 两次 Grok Headless 审阅
|
||||
|
||||
### 结果前
|
||||
|
||||
session `019fb151-9627-76c1-b7d7-53012874f85c` 找出 selector API、output identity、
|
||||
“restore to learned”歧义、parent learned 调度、20% 分母和 negative control 六个硬问题。
|
||||
全部在预注册 freeze 前修正并留档。
|
||||
|
||||
### 结果后
|
||||
|
||||
session `019fb19d-94a3-7231-9a63-3a1ef33a9892` 只读对照 protocol、manifest、
|
||||
analyzer 与 aggregate,独立复算代表性 `G/S/R` cells:
|
||||
|
||||
- 阻断实现错误:0;
|
||||
- one-sided verdict:确认正确;
|
||||
- 不允许修改阈值;
|
||||
- 指出单 group restoration 的 `score_sign_split` reason 是跨指标汇总,因此 reason
|
||||
wording 略宽;`mixed / fail` 本身仍由 C/P pass split 独立成立,主 verdict 不受影响。
|
||||
|
||||
Grok 是方法学审稿人,不是论文或实验事实来源;正式证据仍是冻结代码与 raw outputs。
|
||||
|
||||
## 11. `A_log` 工件边界同步更新
|
||||
|
||||
本轮仍不是 K3 2.8T checkpoint forward。截止 2026-07-30 12:35 CST:
|
||||
|
||||
- official main 仍是 `9f62e4e9fffbd0a83ddd60e1c209d828994b3569`,96 vs 128
|
||||
mismatch 未修;
|
||||
- community PR #144 把 parameter 改成 128,但没有独立 forward 验证;
|
||||
- community PR #150 保留 96,并在加载时验证 / 裁掉 32 个全零尾项;提交者报告检查
|
||||
69 层并完成 disk-offloaded end-to-end generation;
|
||||
- 两个 PR 都未合并,Moonshot 尚未给出官方裁决。
|
||||
|
||||
因此“没有任何公开候选解释”已经过时;“官方 contract 已解决”同样不成立。
|
||||
|
||||
## 12. 最强允许结论
|
||||
|
||||
可以说:
|
||||
|
||||
> 在本缩小 Block AttnRes 模型的同前向 diagnostic backward 中,global learned-value
|
||||
> coefficient sensitivity 对 groups 6+7 的局部 uniformization 具有强 sufficiency;
|
||||
> restoration 只在 contrast 上超过一半,在 peak 上稳定不足一半,因此预注册的双向
|
||||
> localization 未建立。group 7 的 MLP mixer 在次级 sufficiency-only branch 判定中占优。
|
||||
|
||||
不能说:
|
||||
|
||||
- “证明 K3 的尖峰来自 group 6 / 7”;
|
||||
- “groups 6+7 解释了 67.7% / 170.0%”;
|
||||
- “MLP 是唯一原因”;
|
||||
- “只要把这些 mixer 训练成 uniform 就会更稳定”;
|
||||
- “复现了 K3 Figure 5(c)”;
|
||||
- “community PR #150 已经是官方 `A_log` 修复”。
|
||||
@@ -0,0 +1,64 @@
|
||||
# Round 07 Grok Headless 对抗审阅与处置
|
||||
|
||||
审阅日期:2026-07-30
|
||||
审阅会话:`019fb151-9627-76c1-b7d7-53012874f85c`
|
||||
身份:**外部模型的只读方法学审稿,不是论文证据源**
|
||||
|
||||
## 1. 调用边界
|
||||
|
||||
Grok CLI 使用 single/headless 方式读取:
|
||||
|
||||
- `research/K3_ATTNRES_LOCAL_PATH_SCOPING.md`
|
||||
- `research/K3_ATTNRES_LOCAL_PATH_PROTOCOL.md`
|
||||
- `experiments/k3/attnres_spike/train.py`
|
||||
|
||||
关闭 web search、禁止 subagents、使用 plan permission;它没有修改文件。审稿任务是找
|
||||
selector、公式、endpoint exactness、negative control、replay 与父 runner API 的冲突。
|
||||
|
||||
## 2. Blocking findings 与处置
|
||||
|
||||
| finding | 风险 | 处置 |
|
||||
|---|---|---|
|
||||
| 父 runner 只有 global mode,局部 selector 尚不可实现 | 不同实现可能选错 indices 或混入 full autograd | **采纳**:协议冻结 14 个 exact identity sets;所有矩阵 mode 只用同一个 `RoutedSourceBackward` |
|
||||
| 新 output `index=64` 与父 summary 的 `mixer_index=null` 冲突 | 改 schema 会破坏 Round 06 summary hash | **采纳**:父 `trace.mixers` 不动;新建平行 `selector_visits`,64 只作为新 schema alias |
|
||||
| “restore to learned” 会被误解成恢复 query/key/softmax autograd | 分母是 detached learned,实验会回答不同问题 | **采纳**:统一改为 “restore to detached-learned value coefficients” |
|
||||
| 14-mode 合同没明确保留六个 parent `learned` diagnostics | 可能削弱或破坏 Round 06 exactness | **采纳**:六 checkpoints 跑 parent learned;0/8000 追加 14-mode matrix |
|
||||
| 20% raw drop 没写明确分母 | gap gate 可能用不同公式 | **采纳**:冻结 `(X_ref-X_all)/X_ref`,并定义缺失格如何使 3×2 主 gate 失败 |
|
||||
| step-0 只做 custom-mode 互比不够 | 同一种 surrogate 错误可全体一致 | **采纳**:双 endpoint exact、learned-vs-detached 负控制、detached `×1/×2` loss-scale gate |
|
||||
|
||||
## 3. Non-blocking findings 与处置
|
||||
|
||||
全部采纳:
|
||||
|
||||
- 统一符号为 `X_ref = X_detached_learned`;
|
||||
- 在 protocol 正文列出 exact indices,而不只依赖 scoping 公式;
|
||||
- branch dominance 标为 sufficiency-only 次级探索;
|
||||
- interaction residual 明确“不预期为 0、不是检验”;
|
||||
- 明示 group 7 包含固定 `S` 外的 layers 26–28;
|
||||
- 精确定义 `mixed`;
|
||||
- replay 限定 same-host environment;
|
||||
- 所有 score / gate 只由单一 `analyze.py` 生成。
|
||||
|
||||
“单 group 20% 与 joint 50% 不按 mixer 数量成比例”保留为预注册决策;协议已禁止把它解释为
|
||||
per-mixer rate。
|
||||
|
||||
## 4. 复核结论
|
||||
|
||||
Grok 复核 14-mode arithmetic:
|
||||
|
||||
```text
|
||||
1 reference + 10 sufficiency + 3 restoration = 14
|
||||
```
|
||||
|
||||
uniform census `0 / 1 / 4 / 8 / 16 / 49 / 57 / 64 / 65` 与 64 depth + 1 output
|
||||
拓扑一致。真正风险是 identity/schema,不是计数;本轮修订已把两者写成独立 exact gate。
|
||||
|
||||
最终公式被审阅为内部自洽:
|
||||
|
||||
```text
|
||||
G_X = ln(X_ref / X_all)
|
||||
S_X(m) = ln(X_ref / X_m) / G_X
|
||||
R_X(r) = ln(X_r / X_all) / G_X
|
||||
```
|
||||
|
||||
审稿意见不会进入实验结果、论文事实或官网证据等级;它只用于在结果出现前强化协议。
|
||||
@@ -0,0 +1,539 @@
|
||||
# K3 Attention Residuals 局部 mixer 路径干预协议
|
||||
|
||||
协议 ID:`llm-atlas-k3-attnres-local-path-v1`
|
||||
冻结日期:2026-07-30
|
||||
协议状态:**结果前预注册 frozen;任何语义变更必须更换 protocol ID**
|
||||
父协议:`llm-atlas-k3-attnres-spike-path-v1`
|
||||
|
||||
## 0. 研究身份
|
||||
|
||||
本轮是 Round 06 结果后的**定向机制追踪**,不是盲发现。
|
||||
|
||||
已知:
|
||||
|
||||
- depth-32 / Block 的固定尖峰集合是 layers 21–25;
|
||||
- 全部 65 个 mixer 的 source value backward coefficients 从 learned 改成 uniform,
|
||||
在三 seed 上 material 地降低最终 `post_mlp_state` 的 spike contrast 和 peak;
|
||||
- 该全局干预保持 forward exact;
|
||||
- group 6 覆盖 layers 21–24,group 7 覆盖 layers 25–28。
|
||||
|
||||
未知:
|
||||
|
||||
- 全局下降是否主要集中在 group 6 / 7 的 16 个 depth mixers;
|
||||
- group 6 与 group 7 是否各自有稳定 effect;
|
||||
- attention 与 MLP mixer 是否能在预注册阈值下区分;
|
||||
- output mixer 是否解释了大量全局 effect;
|
||||
- 局部 sufficiency 与反向 restoration 是否给出一致证据。
|
||||
|
||||
前置证据、拓扑和同期 artifact audit 固定在
|
||||
`research/K3_ATTNRES_LOCAL_PATH_SCOPING.md`。任何结果不得倒写成事前未知。
|
||||
|
||||
## 1. 允许回答的问题
|
||||
|
||||
1. 在相同模型、batch、loss、activation 与 learned forward weights 下,只改变某个固定
|
||||
mixer scope 的 source-gradient coefficients,能复现多少全局 log gap?
|
||||
2. 从 all-uniform 背景只恢复 group 6 / 7 的 learned coefficients,能恢复多少
|
||||
global log gap?
|
||||
3. groups 6+7 是否在 sufficiency 与 restoration 两个方向、两个 spike 指标、三 seed
|
||||
同时通过 50% 阈值?
|
||||
4. 单独 group 6 或 group 7 是否在两个指标、三 seed 通过 20% 阈值?
|
||||
5. attention-only 与 MLP-only 是否达到预注册的 branch dominance 规则?
|
||||
6. output-only 与 all-depth 控制是否显示 effect 主要来自最终 output mixer?
|
||||
|
||||
## 2. 明确不回答的问题
|
||||
|
||||
- Kimi K3 2.8T checkpoint 的真实训练梯度;
|
||||
- 论文 Figure 5(c) 未公开 telemetry 的精确定义;
|
||||
- 哪个 layer、source 或 operator “产生”尖峰;
|
||||
- learned source weight 的语义归因;
|
||||
- 重新训练局部 uniform variant 的最终能力;
|
||||
- intervention effect 的可加性、Shapley value 或方差分解;
|
||||
- 三 seed 外的总体显著性、置信区间或 p-value;
|
||||
- K3 `A_log` 两个社区修复中哪个已经得到官方认可;
|
||||
- checkpoint conversion、推理正确性或部署可用性。
|
||||
|
||||
## 3. 冻结训练与数据合同
|
||||
|
||||
完整复用 Round 06:
|
||||
|
||||
| 字段 | 固定值 |
|
||||
|---|---|
|
||||
| architecture | Block AttnRes |
|
||||
| Transformer depth | 32 |
|
||||
| aggregation groups | 8 |
|
||||
| blocks / group | 4 |
|
||||
| depth mixers / output mixers | 64 / 1 |
|
||||
| width / heads / FFN | 192 / 6 / 768 |
|
||||
| context / vocabulary | 256 / byte-256 |
|
||||
| seeds | 2026073001 / 2026073002 / 2026073003 |
|
||||
| steps / batch | 8,000 / 32 |
|
||||
| target bytes / formal cell | 65,536,000 |
|
||||
| optimizer | AdamW |
|
||||
| peak / min LR | 3e-4 / 3e-5 |
|
||||
| warmup | 400 |
|
||||
| weight decay | 0.1 for ndim ≥ 2 |
|
||||
| betas / epsilon | 0.9, 0.95 / 1e-8 |
|
||||
| clip | global norm 1.0 |
|
||||
| forward | CUDA BF16 autocast |
|
||||
| residual accumulation | explicit FP32 |
|
||||
| diagnostic CE | fixed 16 × 256 token-mean FP32 CE |
|
||||
| parent diagnostic steps | 0 / 100 / 500 / 2,000 / 4,000 / 8,000 |
|
||||
| local matrix steps | 0 / 8,000 |
|
||||
|
||||
训练路径必须逐调用父 runner 原始 forward;局部 custom autograd 只能存在于 optimizer
|
||||
step 外的 diagnostic。数据 bytes、schedule、validation tensor、diagnostic tensor 与三
|
||||
seed 的每个 optimizer input 都必须与 Round 06 exact。
|
||||
|
||||
诊断调度明确分成两条:
|
||||
|
||||
- 六个 parent diagnostic steps 都运行原始 `learned` mode,用于 Round 06 等价;
|
||||
- step 0 / 8,000 另外运行下述 14-mode local matrix;
|
||||
- localization 公式只读取 14-mode matrix,绝不把 full-autograd `learned` 混入分母。
|
||||
|
||||
## 4. 正式矩阵与 replay
|
||||
|
||||
正式运行:
|
||||
|
||||
```text
|
||||
depth-32 / block / seed-2026073001
|
||||
depth-32 / block / seed-2026073002
|
||||
depth-32 / block / seed-2026073003
|
||||
```
|
||||
|
||||
另从初始化完整重跑:
|
||||
|
||||
```text
|
||||
replay / depth-32 / block / seed-2026073001
|
||||
```
|
||||
|
||||
正式三格处理 196,608,000 target bytes;含 replay 共 262,144,000 bytes。每格使用
|
||||
全新 Python 进程。最多并行两个进程;不能共享 model、optimizer、RNG 或 CUDA graph。
|
||||
wall-time 不进入数值复现合同。
|
||||
|
||||
## 5. 固定主对象与指标
|
||||
|
||||
主对象固定为最终 step 8,000:
|
||||
|
||||
```text
|
||||
position = post_mlp_state
|
||||
reduction = element_rms
|
||||
S = layers 21, 22, 23, 24, 25
|
||||
R = other 27 layers
|
||||
```
|
||||
|
||||
对 mode `m`:
|
||||
|
||||
```text
|
||||
C_m = mean(metric[S]) / mean(metric[R]) # spike contrast
|
||||
P_m = max(metric) / mean(metric) # peak normalized
|
||||
```
|
||||
|
||||
`C_m` 与 `P_m` 必须 finite 且严格大于 `1e-30`。后文统一用 `X_ref` 表示
|
||||
`X_detached_learned`;不用 `detached_reference` 等其他别名。不允许用其他位置、reduction、layer
|
||||
集合或 metric 替换主对象。32-layer raw spectrum、normalized spectrum、peak layer 和
|
||||
top-five layers 全量报告,但不参与主阈值。
|
||||
|
||||
## 6. 14 种冻结模式
|
||||
|
||||
所有模式调用同一个 parent learned forward。custom Function 的 forward 直接返回
|
||||
parent output,只有 backward 对 source tensors 使用选定 coefficients。
|
||||
|
||||
14-mode matrix **全部**使用 `RoutedSourceBackward`:每个 mixer 只在
|
||||
`stopgrad(w)` 与 `1/N` 两种 source value coefficients 中选择。完整
|
||||
query / key / softmax autograd 的 `learned` 不属于这 14 种模式,只用于训练和父诊断
|
||||
等价。实现不得把某个 restoration scope 切回 full-autograd `learned`。
|
||||
|
||||
### 6.1 reference
|
||||
|
||||
`detached_learned`
|
||||
|
||||
- 65 个 mixer 全部使用 learned `w` 作为 source value backward coefficients;
|
||||
- `w` detach,不走 query / key / softmax derivative path;
|
||||
- 必须 exact reproduce Round 06 的同名 mode。
|
||||
|
||||
### 6.2 learned 背景上的局部 uniform:sufficiency family
|
||||
|
||||
未选 mixer 使用 detached learned coefficients;选中 mixer 使用 `1/N`:
|
||||
|
||||
1. `uniform_group_6_only`
|
||||
2. `uniform_group_7_only`
|
||||
3. `uniform_groups_6_7_only`
|
||||
4. `uniform_group_6_attention_only`
|
||||
5. `uniform_group_6_mlp_only`
|
||||
6. `uniform_group_7_attention_only`
|
||||
7. `uniform_group_7_mlp_only`
|
||||
8. `uniform_output_only`
|
||||
9. `uniform_depth_all`
|
||||
10. `uniform_all`
|
||||
|
||||
`uniform_all` 必须 exact reproduce Round 06 的 `uniform_value_backward`。
|
||||
|
||||
### 6.3 all-uniform 背景上的 detached-learned restoration family
|
||||
|
||||
选中 scope 恢复 detached learned coefficients,其余保持 uniform:
|
||||
|
||||
1. `uniform_except_group_6`
|
||||
2. `uniform_except_group_7`
|
||||
3. `uniform_except_groups_6_7`
|
||||
|
||||
名字中的 `except` 表示该 scope **不是 uniform**。报告和网站必须同时展示人话标签
|
||||
“restore ... to detached-learned value coefficients”,避免误读。
|
||||
|
||||
## 7. selector 的唯一合同
|
||||
|
||||
为了同时满足新 selector audit 和 Round 06 endpoint exactness,保留两个互不混写的
|
||||
schema:
|
||||
|
||||
1. 父 `trace.mixers` summary **逐字段不变**;output 仍使用父 schema 的
|
||||
`mixer_index=null`,并继续参与父 `mixer_summary_sha256`;
|
||||
2. 新增平行 `selector_visits`,只用于 local mask audit,不写入父 summary。
|
||||
|
||||
`selector_visits` 的 depth mixer identity 用:
|
||||
|
||||
```text
|
||||
(kind="depth", index=0..63, layer=1..32,
|
||||
group=1..8, branch in {"attention","mlp"})
|
||||
```
|
||||
|
||||
`selector_visits` 的 output mixer identity 用:
|
||||
|
||||
```text
|
||||
(kind="output", index=64, layer=null, group=null, branch="output")
|
||||
```
|
||||
|
||||
这里 `index=64` 只是新 selector schema 的稳定别名,不得回写父 summary。
|
||||
|
||||
令 `D_i` 表示 `kind=depth,index=i`,`O` 表示 output。14 种 mode 的 uniform identity
|
||||
集合冻结如下:
|
||||
|
||||
| mode | exact uniform set |
|
||||
|---|---|
|
||||
| `detached_learned` | `∅` |
|
||||
| `uniform_group_6_only` | `{D40,…,D47}` |
|
||||
| `uniform_group_7_only` | `{D48,…,D55}` |
|
||||
| `uniform_groups_6_7_only` | `{D40,…,D55}` |
|
||||
| `uniform_group_6_attention_only` | `{D40,D42,D44,D46}` |
|
||||
| `uniform_group_6_mlp_only` | `{D41,D43,D45,D47}` |
|
||||
| `uniform_group_7_attention_only` | `{D48,D50,D52,D54}` |
|
||||
| `uniform_group_7_mlp_only` | `{D49,D51,D53,D55}` |
|
||||
| `uniform_output_only` | `{O}` |
|
||||
| `uniform_depth_all` | `{D0,…,D63}` |
|
||||
| `uniform_all` | `{D0,…,D63,O}` |
|
||||
| `uniform_except_group_6` | `{D0,…,D39,D48,…,D63,O}` |
|
||||
| `uniform_except_group_7` | `{D0,…,D47,D56,…,D63,O}` |
|
||||
| `uniform_except_groups_6_7` | `{D0,…,D39,D56,…,D63,O}` |
|
||||
|
||||
runner 必须把这些 set 编码为一个 frozen selector 函数;不能散落在 mode-specific
|
||||
if/else 中。manifest 同时保存 machine-readable exact index lists。runner 通过
|
||||
override `_mix` 或等价 hook 做 set lookup,并替换父 runner 中只接受三种 global mode
|
||||
的 mode validation、bundle loop 和相关 gate;训练 forward 继续直接调用父路径。
|
||||
|
||||
每次 forward 必须验证:
|
||||
|
||||
- exactly 65 个 mixer visits;
|
||||
- identity 不重复;
|
||||
- identity 顺序与 reference exact;
|
||||
- 父 `trace.mixers` schema 与 hash 路径没有新字段;
|
||||
- uniform census 与 manifest exact;
|
||||
- selected identity list 与 selector rule exact;
|
||||
- reference 的 uniform count 为 0;
|
||||
- branch-only 4,group-only 8,groups 6+7 为 16;
|
||||
- output-only 1,all-depth 64,all 65;
|
||||
- except-one-group 57,except-two-groups 49。
|
||||
|
||||
任一 gate 失败,cell invalid;不得只改结果 JSON。
|
||||
|
||||
## 8. forward identity 与 parent exactness
|
||||
|
||||
### 8.1 所有 14 模式的 forward identity
|
||||
|
||||
同 seed / step 相对 `detached_learned` 必须满足:
|
||||
|
||||
- logits tensor SHA-256 exact;
|
||||
- loss FP32 value exact;
|
||||
- 六位置 activation tensor hashes exact;
|
||||
- 65 个 mixer forward summaries exact。
|
||||
|
||||
任一 mode 失败,整格 invalid。
|
||||
|
||||
### 8.2 Round 06 endpoint exactness
|
||||
|
||||
对 step 0 / 8,000:
|
||||
|
||||
- `detached_learned` 的 logits、loss、六位置 activation、未改 schema 的 mixer
|
||||
summaries 和六位置 gradient reductions 必须与对应 Round 06
|
||||
raw output exact;
|
||||
- `uniform_all` 的同一组字段和六位置 gradient reductions必须与对应 Round 06
|
||||
`uniform_value_backward` exact;
|
||||
- 正式训练的 final model hash、optimizer hash、六个 validation BPC、training
|
||||
history 与六个 parent `learned` diagnostics 必须与 Round 06 exact;
|
||||
- 新 `selector_visits` 不参与旧 `mixer_summary_sha256`,而由独立 canonical hash
|
||||
和 exact-set gate 管理。
|
||||
|
||||
runner / protocol / scoping / manifest 物理 hash 在运行前冻结。父 raw 文件同时检查 physical
|
||||
SHA-256、canonical SHA-256、final model hash 和 final optimizer hash。
|
||||
|
||||
## 9. 初始化负控制
|
||||
|
||||
step 0 的 mixer query 为零,learned `w` 是 uniform。14 模式在六个位置的
|
||||
`element_rms` 必须:
|
||||
|
||||
- 32 个 raw values 全部 finite、strictly positive;
|
||||
- 相对 reference 的逐层 raw relative error `≤1e-6`;
|
||||
- normalized absolute error `≤1e-6`。
|
||||
|
||||
此外:
|
||||
|
||||
- `detached_learned` 与 `uniform_all` 必须分别与 Round 06 step-0 endpoint exact;
|
||||
- parent full-autograd `learned` 与 `detached_learned` 必须按 Round 06 负控制在
|
||||
`1e-6` tolerance 内一致;
|
||||
- `detached_learned` 另执行同一 loss 的 `×1 / ×2` backward,六位置、七 reductions
|
||||
都必须通过父协议相同的 scale 与 normalized-spectrum gate。
|
||||
|
||||
失败表示 selector 或 surrogate 没有隔离预期路径;正式结果无效。
|
||||
|
||||
## 10. global log gap
|
||||
|
||||
对每个 seed 和每个指标 `X ∈ {C,P}`:
|
||||
|
||||
```text
|
||||
G_X = ln(X_ref / X_uniform_all)
|
||||
relative_drop_X = (X_ref - X_uniform_all) / X_ref
|
||||
```
|
||||
|
||||
只有同时满足以下条件才允许解释局部比例:
|
||||
|
||||
1. `G_C > 0` 且 `G_P > 0`;
|
||||
2. 上式 `relative_drop_X ≥0.20`;
|
||||
3. Round 06 endpoint exactness 通过。
|
||||
|
||||
seed `s` 的 metric `X` 任一条件不满足,则该 `(s,X)` 称为
|
||||
`global gap not established`,不计算该格 `S_X / R_X`。groups 6+7 的主 gate 要求
|
||||
3 seed × 2 metrics 全部存在,因此任一 required cell 缺失都会使主 localization
|
||||
判定失败;仍公开 raw matrix,不使用事后替代分母。
|
||||
|
||||
log ratio 用于让相同的乘法变化在两个方向可比。所有归一化值按原值报告,**不裁剪到
|
||||
[0,1]**;负值表示反方向,超过 1 表示局部 intervention 超过 all-uniform endpoint。
|
||||
|
||||
## 11. sufficiency score
|
||||
|
||||
对 sufficiency mode `m`:
|
||||
|
||||
```text
|
||||
S_X(m) = ln(X_ref / X_m) / G_X
|
||||
```
|
||||
|
||||
### 11.1 groups 6+7 主判定
|
||||
|
||||
只有 `uniform_groups_6_7_only` 对 `C` 和 `P` 都满足:
|
||||
|
||||
```text
|
||||
S_X(m) ≥ 0.50
|
||||
```
|
||||
|
||||
且三个 formal seed 6 / 6 全部达标,才记为:
|
||||
|
||||
> groups 6+7 的 16 个 depth mixers 在本 diagnostic 中,足以复现至少一半
|
||||
> all-65 uniform intervention 的预注册 log-gap reduction。
|
||||
|
||||
任一失败记为 `not sufficient at the preregistered 50% threshold`。`mixed` 精确定义为:
|
||||
seed 通过/失败不一致、`C/P` 通过/失败不一致,或 score 的正负号跨 seed 不一致;可同时
|
||||
附加多个原因,不得降低阈值。
|
||||
|
||||
### 11.2 单 group
|
||||
|
||||
group 6 / group 7 分别对 `C` 和 `P`、三 seed 全部满足:
|
||||
|
||||
```text
|
||||
S_X(m) ≥ 0.20
|
||||
```
|
||||
|
||||
才称为 `material local sufficiency at the 20% threshold`。没过阈值不等于 effect 为零。
|
||||
|
||||
## 12. restoration score
|
||||
|
||||
对 restoration mode `r`:
|
||||
|
||||
```text
|
||||
R_X(r) = ln(X_r / X_uniform_all) / G_X
|
||||
```
|
||||
|
||||
### 12.1 groups 6+7 主判定
|
||||
|
||||
只有 `uniform_except_groups_6_7` 对 `C` 和 `P`、三 seed全部满足:
|
||||
|
||||
```text
|
||||
R_X(r) ≥ 0.50
|
||||
```
|
||||
|
||||
才称为:
|
||||
|
||||
> 从 all-uniform 背景只恢复 groups 6+7 的 learned coefficients,恢复了至少一半
|
||||
> 预注册 global log gap。
|
||||
|
||||
这仍是同前向 backward-rule restoration sensitivity,不是严格 causal necessity。
|
||||
|
||||
### 12.2 单 group
|
||||
|
||||
`uniform_except_group_6` / `uniform_except_group_7` 分别以 `≥0.20`、两个指标、三 seed
|
||||
作为 material restoration threshold。
|
||||
|
||||
## 13. localization 总闸门
|
||||
|
||||
只有以下两项同时通过:
|
||||
|
||||
1. groups 6+7 sufficiency:`S_C,S_P ≥0.50`,3 / 3 seeds;
|
||||
2. groups 6+7 restoration:`R_C,R_P ≥0.50`,3 / 3 seeds;
|
||||
|
||||
才允许写:
|
||||
|
||||
> 在本缩小模型、固定训练状态和 diagnostic backward 下,全局 value-coefficient
|
||||
> sensitivity 的主要部分 localization 到 group 6 / 7 mixer path。
|
||||
|
||||
即使通过,也必须紧邻注明:
|
||||
|
||||
- “主要部分”由 50% 双向阈值定义;
|
||||
- effect non-additive;
|
||||
- 不是唯一来源或 layer-origin;
|
||||
- 不是真实 K3 checkpoint 结论。
|
||||
|
||||
一侧通过一侧失败,统一写成 `one-sided evidence, localization not established`。
|
||||
|
||||
## 14. attention vs MLP branch 判定
|
||||
|
||||
每个 group 独立比较 attention-only 与 MLP-only sufficiency score。只有某 branch:
|
||||
|
||||
1. `S_C ≥0.20` 且 `S_P ≥0.20`;
|
||||
2. 在 `C` 与 `P` 上都比 sibling 高至少 `0.15`;
|
||||
3. 三 seed 全部满足前两项;
|
||||
|
||||
才称为 `branch-dominant at the preregistered margin`。
|
||||
|
||||
若 group-level sufficiency 未通过 20% 阈值,不允许宣称其内部 branch dominance。
|
||||
branch-only scores 可能交互、超加或相互抵消,不能相加成 group score。
|
||||
本节只有 sufficiency 方向,没有 branch-level restoration,属于预注册的次级探索性
|
||||
判定,证据层级低于 §13 双向 localization。
|
||||
|
||||
## 15. output 与 depth 控制
|
||||
|
||||
`uniform_output_only` 和 `uniform_depth_all` 不进入 group localization 主判定。
|
||||
|
||||
探索性报告:
|
||||
|
||||
```text
|
||||
S_X(output)
|
||||
S_X(depth_all)
|
||||
interaction_residual_X =
|
||||
1 - S_X(output) - S_X(depth_all)
|
||||
```
|
||||
|
||||
`interaction_residual` 只是 log-gap bookkeeping,不是统计交互估计或贡献分解。
|
||||
它不预期接近 0,也不是 hypothesis test。
|
||||
|
||||
只有 output-only 对两个指标、三 seed 都 `≥0.50`,才标记
|
||||
`output mixer alone captures at least half the global gap`。即使如此,也不否定
|
||||
groups 6+7;两者可能重叠、串联或超加。
|
||||
|
||||
## 16. 报告顺序与反 cherry-picking
|
||||
|
||||
固定报告顺序:
|
||||
|
||||
1. input / parent / endpoint exactness;
|
||||
2. step-0 negative control;
|
||||
3. 每 seed 的 raw `C` / `P` 矩阵;
|
||||
4. global gaps;
|
||||
5. groups 6+7 sufficiency;
|
||||
6. groups 6+7 restoration;
|
||||
7. localization gate;
|
||||
8. single-group scores;
|
||||
9. branch scores;
|
||||
10. output / depth controls;
|
||||
11. full 32-layer spectra;
|
||||
12. replay;
|
||||
13. limitations。
|
||||
|
||||
所有 14 modes、两个指标、三个 seed 都公开。不得只展示通过阈值的 scope。不得用跨 seed
|
||||
均值替代 3 / 3 gate;均值只用于视觉摘要。
|
||||
|
||||
## 17. replay 与复现闸门
|
||||
|
||||
seed 2026073001 从初始化独立 replay,比较去除以下字段后的 canonical content:
|
||||
|
||||
- `run_kind`;
|
||||
- wall-clock timing;
|
||||
- output path;
|
||||
- self canonical hash。
|
||||
|
||||
至少以下字段必须 exact:
|
||||
|
||||
- input tensor hashes;
|
||||
- initial/final model 与 optimizer hashes;
|
||||
- evaluations / training history;
|
||||
- parent diagnostics;
|
||||
- 14-mode step-0 / step-8,000 forward hashes;
|
||||
- selector census / identities;
|
||||
- 六位置 raw gradient reductions;
|
||||
- global gaps / local scores / gates。
|
||||
|
||||
若正式 seed1 与 replay 不 exact,Round 07 数值结论无效。
|
||||
replay 固定在与 formal 相同 host、GPU、Python、PyTorch、CUDA 和
|
||||
`CUBLAS_WORKSPACE_CONFIG` 环境;本协议不声称跨硬件 bit exact。
|
||||
|
||||
## 18. 预期失败与停止规则
|
||||
|
||||
以下任一项使 cell invalid:
|
||||
|
||||
- CUDA deterministic contract 未开启;
|
||||
- parent manifest / runner / protocol / scoping / raw hash 不匹配;
|
||||
- 训练等价失败;
|
||||
- diagnostic 改变 optimizer state;
|
||||
- forward identity 失败;
|
||||
- selector identity / census 失败;
|
||||
- Round 06 endpoint exactness 失败;
|
||||
- step-0 negative control 失败;
|
||||
- raw gradient missing、non-finite 或 non-positive;
|
||||
- global gap denominator 不成立。
|
||||
|
||||
程序错误修复必须:
|
||||
|
||||
1. 保存失败日志;
|
||||
2. 修改 runner;
|
||||
3. 更新 runner hash;
|
||||
4. 明确判断协议语义是否改变;
|
||||
5. 若改变 selector、mode、metric、threshold 或 aggregation,创建新 protocol ID;
|
||||
6. 全部受影响 cell 从初始化重跑。
|
||||
|
||||
## 19. 结果语言边界
|
||||
|
||||
允许:
|
||||
|
||||
- “在同前向 diagnostic backward 下,uniformizing scope X 改变了固定尖峰指标”;
|
||||
- “groups 6+7 在预注册 50% 双向阈值下建立 / 未建立 localization”;
|
||||
- “branch effect mixed / below threshold”;
|
||||
- “这是一项 reduced-model mechanism probe”。
|
||||
|
||||
禁止:
|
||||
|
||||
- “证明 K3 的尖峰来自第 6 组”;
|
||||
- “这些 mixer 贡献了 X% 梯度”;
|
||||
- “group effect 加总为 100%”;
|
||||
- “uniform mixer 更适合训练”;
|
||||
- “复现了 Figure 5(c)”;
|
||||
- “验证了 K3 2.8T checkpoint”;
|
||||
- “社区 PR #144 或 #150 已成为官方修复”。
|
||||
|
||||
## 20. 冻结清单
|
||||
|
||||
在任何 formal 结果产生前必须完成:
|
||||
|
||||
- [x] scoping 文件完成;
|
||||
- [x] protocol 状态改为 frozen;
|
||||
- [ ] 14 modes 与 selector census 写入 manifest;
|
||||
- [ ] thresholds / formulas 写入 manifest;
|
||||
- [ ] Round 06 父 artifact physical / canonical hashes 写入 manifest;
|
||||
- [ ] runner、protocol、scoping、manifest hashes 固定;
|
||||
- [x] Grok Headless 对抗审阅完成,采纳/拒绝理由留档;
|
||||
- [ ] step-0 smoke 全门通过;
|
||||
- [ ] formal 命令与环境写入 README;
|
||||
- [ ] 单一 `analyze.py` 实现所有 score / gate,网站只消费其冻结输出;
|
||||
- [ ] protocol commit 早于 formal result commit。
|
||||
@@ -0,0 +1,197 @@
|
||||
# K3 Attention Residuals 局部 mixer 路径:Round 07 前置定位
|
||||
|
||||
研究日期:2026-07-30
|
||||
阶段身份:**定向 scoping,不是 Round 07 预注册结果**
|
||||
上游协议:`llm-atlas-k3-attnres-spike-path-v1`
|
||||
|
||||
## 1. 已知到什么程度
|
||||
|
||||
Round 06 在同一个 depth-32 / Block AttnRes 缩小模型上,把前向保持为 learned
|
||||
weights,只改写 mixer 的反向规则。三 seed 的最终 `post_mlp_state / element_rms`
|
||||
结果为:
|
||||
|
||||
| backward rule | spike contrast(3-seed mean) | 相对 detached learned | peak normalized(3-seed mean) | 相对 detached learned |
|
||||
|---|---:|---:|---:|---:|
|
||||
| learned | 2.754 | — | 4.803 | — |
|
||||
| detached learned | 2.812 | reference | 4.956 | reference |
|
||||
| uniform value / all 65 mixers | 0.837 | **−70.2%** | 3.122 | **−37.0%** |
|
||||
|
||||
其中:
|
||||
|
||||
- `learned → detached learned` 没有降低尖峰,contrast 反而平均增加约 2.0%;
|
||||
- `detached learned → uniform value backward` 在 contrast 和 peak 上都 3 / 3 seed
|
||||
超过预注册的 20% material threshold;
|
||||
- 三种模式的 logits、loss、六位置 activation 和 mixer forward 摘要全部 exact;
|
||||
- 这证明的是**全局 backward-rule sensitivity**,不是训练变体,也不是局部归因。
|
||||
|
||||
因此 Round 07 不再重复问“learned value coefficients 是否重要”,而是问:
|
||||
|
||||
> 65 个 mixer 全局改写带来的下降,主要能否由尖峰邻近的 group 6 / 7
|
||||
> depth mixers 复现,并能否从反方向恢复?
|
||||
|
||||
## 2. 固定拓扑,而不是结果后挑层
|
||||
|
||||
depth 32 的 Block AttnRes 有 8 个 aggregation groups,每组 4 个 Transformer
|
||||
blocks。每层有 attention 和 MLP 两个 depth mixers,合计 64 个;模型末尾还有一个
|
||||
独立 output mixer,合计 65 个 intervention nodes。
|
||||
|
||||
对 1-based layer `l`:
|
||||
|
||||
```text
|
||||
group = floor((l - 1) / 4) + 1
|
||||
attention mixer index = 2 × (l - 1) # 0-based
|
||||
MLP mixer index = 2 × (l - 1) + 1 # 0-based
|
||||
```
|
||||
|
||||
所以:
|
||||
|
||||
| scope | layers | 0-based depth mixer indices | mixer count |
|
||||
|---|---:|---:|---:|
|
||||
| group 6 | 21–24 | 40–47 | 8 |
|
||||
| group 7 | 25–28 | 48–55 | 8 |
|
||||
| groups 6+7 | 21–28 | 40–55 | 16 |
|
||||
| output | — | separate node | 1 |
|
||||
|
||||
Round 05 已在看过数据后冻结尖峰集合 `S = layers 21–25`。它覆盖完整 group 6 和
|
||||
group 7 的首层。因此 Round 07 明确是**定向邻域追踪**,不能称为盲发现;group 6 / 7
|
||||
也不能结果后替换成更好看的范围。
|
||||
|
||||
## 3. 为什么需要两个方向
|
||||
|
||||
只在 learned 背景把 group 6 / 7 改成 uniform,回答的是:
|
||||
|
||||
> 只改这段是否足以复现全局干预的一大部分下降?
|
||||
|
||||
但 mixer 路径有串联、分流和 nonlinear interaction,单侧结果可能被其他 learned
|
||||
路径补偿。反过来,在 all-uniform 背景只把 group 6 / 7 恢复为 learned,回答的是:
|
||||
|
||||
> 只恢复这段是否足以让尖峰朝 reference 回升?
|
||||
|
||||
两种值都不是“贡献百分比”,也不要求相加为 100%。Round 07 用相同的 global log gap
|
||||
归一化两种方向,只把双向、跨 seed 稳定的结果称为 localization evidence。
|
||||
|
||||
## 4. 冻结候选范围
|
||||
|
||||
从 `detached_learned` 背景出发的 sufficiency scopes:
|
||||
|
||||
1. group 6;
|
||||
2. group 7;
|
||||
3. groups 6+7;
|
||||
4. group 6 attention-only;
|
||||
5. group 6 MLP-only;
|
||||
6. group 7 attention-only;
|
||||
7. group 7 MLP-only;
|
||||
8. output-only;
|
||||
9. all 64 depth mixers;
|
||||
10. all 65 mixers。
|
||||
|
||||
从 all-uniform 背景出发的 restoration scopes:
|
||||
|
||||
1. restore group 6 to detached-learned value coefficients;
|
||||
2. restore group 7 to detached-learned value coefficients;
|
||||
3. restore groups 6+7 to detached-learned value coefficients。
|
||||
|
||||
加上 `detached_learned` reference,共 14 种模式。预期 uniform selector census 为:
|
||||
|
||||
| mode | uniform mixers |
|
||||
|---|---:|
|
||||
| detached reference | 0 |
|
||||
| group branch only | 4 |
|
||||
| group only | 8 |
|
||||
| groups 6+7 | 16 |
|
||||
| output only | 1 |
|
||||
| all depth | 64 |
|
||||
| all | 65 |
|
||||
| uniform except group 6 / 7 | 57 |
|
||||
| uniform except groups 6+7 | 49 |
|
||||
|
||||
每次 diagnostic 都必须保存实际选中的 mixer identity;不能只信 mode 名称。
|
||||
group 7 的 scope 包含 layers 26–28,它们不在固定尖峰集合 `S=21–25` 中,所以
|
||||
group 7 是预先定义的**完整邻接 group intervention**,不是 spike-layer-only
|
||||
intervention。
|
||||
|
||||
## 5. 不把局部 intervention 误译成什么
|
||||
|
||||
即使 groups 6+7 双向通过,结论也只限于:
|
||||
|
||||
- 固定训练状态;
|
||||
- 固定 diagnostic batch 与 loss;
|
||||
- 固定 `S = 21–25` 指标;
|
||||
- 同前向、替代 source-gradient coefficient 的 diagnostic backward。
|
||||
|
||||
它不等于:
|
||||
|
||||
- 这些层“产生”了尖峰;
|
||||
- group 6 / 7 是唯一原因;
|
||||
- 真实 K3 checkpoint 有相同梯度路径;
|
||||
- 把 mixer 训练成 uniform 会有同样结果;
|
||||
- 局部 effect 可加,或可解释成方差分解;
|
||||
- 论文 Figure 5(c) 的未公开 telemetry 已被复现。
|
||||
|
||||
保留 output-only 和 all-depth 两个控制,是为了看清最终 readout 与 depth path 的关系;
|
||||
它们不进入 group 6 / 7 localization 的主判定。
|
||||
|
||||
## 6. 同期 artifact 状态审计:`A_log`
|
||||
|
||||
这一问题与缩小实验的局部梯度机制**相互独立**,但会限制任何真实 K3 checkpoint
|
||||
验证,因此在冻结 Round 07 前重新检查官方模型仓库。
|
||||
|
||||
截至 **2026-07-30 12:35 CST**:
|
||||
|
||||
- 官方 Hugging Face main commit 仍为
|
||||
`9f62e4e9fffbd0a83ddd60e1c209d828994b3569`;
|
||||
- main 的 `modeling_kimi_linear.py` 仍以 `num_heads=96` 初始化 `A_log`;
|
||||
- 已发布 checkpoint 中该张量的公开 shape 是 `[128]`,与 main 存在加载不匹配;
|
||||
- 官方 main 尚未合并修复或给出 conversion contract。
|
||||
|
||||
同时出现了两个**未合并、互相竞争的社区 PR**:
|
||||
|
||||
### PR #144:把参数改成 128
|
||||
|
||||
- 一行把初始化从 `self.num_heads` 改为 `self.head_dim`;
|
||||
- 提交者报告所有 shards 能加载;
|
||||
- 提交者明确说没有独立验证 forward;
|
||||
- 它把 checkpoint shape 当作权威语义。
|
||||
|
||||
### PR #150:保留 96,加载时验证并裁零尾
|
||||
|
||||
- 保持模型参数为 `[num_heads]=[96]`;
|
||||
- `_load_from_state_dict` 检查 `[96:128]` 全为零后再裁掉;
|
||||
- 提交者报告检查了 69 个 KDA 层,所有 32 项尾部都 exact zero;
|
||||
- 提交者还报告经过 disk-offloaded MoE 的完整生成;
|
||||
- 这些 checkpoint 全量扫描和生成是**提交者报告**,本项目没有下载约 1.56 TB
|
||||
权重独立复核;本项目只核对了 PR diff、main 代码路径和 PR 状态。
|
||||
|
||||
PR #150 进一步指出,forward 中 `v` 被 reshape 为 96 heads,kernel 随后接收
|
||||
`A_log`;若直接采用 #144 的 128 元素参数,现有 `view(H, 1)` 路径会在 96 heads 下
|
||||
失败。这个论证比单看 checkpoint shape 更完整,但在官方合并或独立复核前,仍必须标成
|
||||
高可信社区解释,而不是 Kimi 官方结论。
|
||||
|
||||
当前准确状态应写成:
|
||||
|
||||
> official main 仍然不匹配;社区已有两个竞争性候选修复,其中 #150 提供了更完整的
|
||||
> checkpoint-tail 与 end-to-end 证据,但尚无官方裁决。
|
||||
|
||||
来源:
|
||||
|
||||
- [Kimi-K3 official main](https://huggingface.co/moonshotai/Kimi-K3/tree/main)
|
||||
- [main `modeling_kimi_linear.py`](https://huggingface.co/moonshotai/Kimi-K3/blob/main/modeling_kimi_linear.py)
|
||||
- [community PR #144](https://huggingface.co/moonshotai/Kimi-K3/discussions/144)
|
||||
- [community PR #150](https://huggingface.co/moonshotai/Kimi-K3/discussions/150)
|
||||
|
||||
## 7. Round 07 的可证伪问题
|
||||
|
||||
Round 07 将:
|
||||
|
||||
1. exact replay Round 06 的三 seed 训练;
|
||||
2. exact reproduce Round 06 的 `detached_learned` 与 `uniform_all` 两个端点;
|
||||
3. 在 step 0 对 14 种模式做负控制,在 step 8,000 做正式矩阵;
|
||||
4. 同时测 `spike_contrast` 与 `peak_normalized`;
|
||||
5. 用 groups 6+7 的 sufficiency 与 restoration 两个方向预注册 50% log-gap
|
||||
localization threshold;
|
||||
6. 用单 group 的 20% 阈值和 attention-vs-MLP 的 15 percentage-point margin
|
||||
作更细分的层级判定;
|
||||
7. 从初始化完整 replay seed 2026073001。
|
||||
|
||||
完整模式、公式、失败规则和复现合同见
|
||||
`research/K3_ATTNRES_LOCAL_PATH_PROTOCOL.md`。
|
||||
@@ -228,7 +228,7 @@ if (numeric(reliability.initial.passAt) <= numeric(reliability.k2.passAt) || num
|
||||
if (reliability.nonIdempotent.sideRisk === "LOW") failures.push("非幂等写操作风险没有提升");
|
||||
if (numeric(rl.wait.utilization) >= numeric(rl.full.utilization) || numeric(rl.wait.lostWork) <= numeric(rl.full.lostWork)) failures.push("wait-all 长尾/重算方向异常");
|
||||
if (!rl.wait.takeaway.includes("wait-all") || rl.keyboardSelected !== "rl" || rl.keyboardVisible !== "rl") failures.push("长程 RL 解释或键盘导航异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("尖峰不是出生时就有") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
|
||||
if (home.paperCount !== "486" || papers.total !== 486 || !papers.hasAgentFilter || papers.agentVisible < 52) failures.push("论文库 Agent 标签或论文总数异常");
|
||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
||||
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||
|
||||
@@ -226,7 +226,7 @@ if (!update.steps[0].includes("Fixed preference")) failures.push("DPO 更新流
|
||||
if (!recipe.family.includes("Multi-effort") || !recipe.regime.includes("9 RL experts") || !recipe.constraints.includes("verbosity")) failures.push("K3 配方合同异常");
|
||||
if (!recipe.path.some((step) => step.includes("3 domains × 3 efforts")) || !recipe.path.some((step) => step.includes("MOPD"))) failures.push("K3 配方路径异常");
|
||||
if (recipe.keyboardSelected !== "recipe" || recipe.keyboardVisible !== "recipe") failures.push("实验 tab 键盘导航异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("尖峰不是出生时就有") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
|
||||
if (home.paperCount !== "486" || papers.total !== 486 || !papers.hasAlignmentFilter || papers.alignmentVisible < 35) failures.push("论文库后训练标签或论文总数异常");
|
||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
||||
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||
|
||||
@@ -234,7 +234,7 @@ if (layout.navLinks !== 20 || mobile.mobileLinks !== 20 || home.navLinks !== 20)
|
||||
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
|
||||
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
||||
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("尖峰不是出生时就有") || home.firstHref !== "/k3/") {
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") {
|
||||
failures.push("首页 Transformer 新章入口异常");
|
||||
}
|
||||
if (home.paperCount !== "486") failures.push(`首页论文总数异常:${home.paperCount}`);
|
||||
|
||||
@@ -1318,7 +1318,7 @@ if (completionDepth.tasks.panel !== "tasks" || completionDepth.tasks.mathCards !
|
||||
if (completionDepth.hidden.panel !== "hidden" || completionDepth.hidden.stages !== 29 || completionDepth.hidden.selected !== "layer_07" || completionDepth.hidden.points !== 29 || completionDepth.hidden.exact === "1,537 / 1,537" || numeric(completionDepth.hidden.relative) <= 0) failures.push("29 阶段隐藏状态曲线或交互异常");
|
||||
if (completionDepth.router.panel !== "router" || completionDepth.router.layers !== 26 || completionDepth.router.selected !== "layer 24" || completionDepth.router.points !== 26 || !completionDepth.router.ordered.includes("%") || !completionDepth.router.setExact.includes("%") || numeric(completionDepth.router.tv) <= 0 || completionDepth.router.reproCards !== 4 || completionDepth.router.links !== 4) failures.push("26 层 MoE 路由曲线或复跑证据异常");
|
||||
if (completionDepth.keyboardSelected !== "tasks" || completionDepth.keyboardVisible !== "tasks") failures.push("完成度与全深度实验键盘 tab 导航异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("尖峰不是出生时就有") || home.firstHref !== "/k3/" || home.paperCount !== "486") failures.push("首页 K3 首发入口或论文数异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/" || home.paperCount !== "486") failures.push("首页 K3 首发入口或论文数异常");
|
||||
if (papers.total !== 486 || !papers.hasFilter || papers.visible < 20 || !papers.hasCoder || !papers.hasEngram) failures.push("论文库 DeepSeek 聚光异常");
|
||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4 || mobile.artifactTabs !== 13 || mobile.behaviorTabs !== 4 || mobile.behaviorSources !== 16 || mobile.behaviorEdges !== 10 || mobile.behaviorDeviceCells !== 29 || mobile.completionDepthTabs !== 4 || mobile.completionDepthHiddenStages !== 29 || mobile.completionDepthRouterLayers !== 26 || mobile.artifactHeatCells !== 64 || mobile.corpusCohorts !== 3 || mobile.lengthDeltaCards !== 4 || mobile.templateLayers !== 6 || mobile.templateScopes !== 2 || mobile.templateModes !== 2 || mobile.templateDomainCards !== 4 || mobile.templateDepthCells !== 24 || mobile.historyLayers !== 6 || mobile.historyScopes !== 2 || mobile.historyModes !== 2 || mobile.historyEffects !== 3 || mobile.historyDomainCards !== 4 || mobile.historyDepthCells !== 24 || mobile.distanceLayers !== 6 || mobile.distanceScopes !== 2 || mobile.distanceModes !== 2 || mobile.distanceContrasts !== 2 || mobile.distanceDomainCards !== 4 || mobile.distanceDepthCells !== 24 || mobile.boundaryLayers !== 6 || mobile.boundaryScopes !== 2 || mobile.boundaryModes !== 2 || mobile.boundaryContrasts !== 3 || mobile.boundaryTokenCards !== 4 || mobile.boundaryDomainCards !== 4 || mobile.boundaryDepthCells !== 24 || mobile.roleLayers !== 6 || mobile.roleScopes !== 2 || mobile.roleModes !== 2 || mobile.roleContrasts !== 3 || mobile.roleLevelCards !== 4 || mobile.roleDomainCards !== 4 || mobile.roleDepthCells !== 24 || mobile.specialLayers !== 6 || mobile.specialScopes !== 2 || mobile.specialModes !== 2 || mobile.specialContrasts !== 4 || mobile.specialTokenCards !== 4 || mobile.specialDomainCards !== 4 || mobile.specialDepthCells !== 24 || mobile.roleBlockLayers !== 6 || mobile.roleBlockScopes !== 2 || mobile.roleBlockModes !== 2 || mobile.roleBlockEffects !== 3 || mobile.roleBlockMatrixCards !== 4 || mobile.roleBlockDomainCards !== 4 || mobile.roleBlockDepthCells !== 24) failures.push("移动端导航或实验异常");
|
||||
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
||||
|
||||
@@ -277,7 +277,7 @@ if (numeric(system.initial.success) <= numeric(system.initial.model) || numeric(
|
||||
if (numeric(system.cheap.success) >= numeric(system.initial.success) || numeric(system.cheap.cost) !== 4) failures.push("低预算没有降低成功率 / 成本");
|
||||
if (numeric(system.locked.unsafe) !== 0 || numeric(system.locked.overrefusal) <= numeric(system.initial.overrefusal)) failures.push("安全壳没有展现危险服从 / 过拒权衡");
|
||||
if (system.keyboardSelected !== "judge" || system.keyboardVisible !== "judge") failures.push("实验键盘 tab 导航异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("尖峰不是出生时就有") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
|
||||
if (home.paperCount !== "486" || home.topicCount !== "17" || papers.total !== 486 || !papers.hasFilter || papers.visible < 80) failures.push("首页 / 论文库评测索引异常");
|
||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
||||
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
||||
|
||||
@@ -264,7 +264,7 @@ if (!fleet.k3.avoided.includes("320K") || fleet.k3.shortSlo !== "PROTECTED") fai
|
||||
if (!fleet.failed.state.includes("SECONDARY RE-PREFILL") || !fleet.failed.recompute.includes("FAILED PRIMARY")) failures.push("缓存故障没有触发原子失效后的重算");
|
||||
if (fleet.bursty.shortSlo !== "VIOLATED") failures.push("平均并发阈值没有暴露长请求突发");
|
||||
if (fleet.keyboardSelected !== "phase" || fleet.keyboardVisible !== "phase") failures.push("实验键盘 tab 导航异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("尖峰不是出生时就有") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
|
||||
if (home.paperCount !== "486" || papers.total !== 486 || !papers.hasFilter || papers.visible !== 46) failures.push("论文库推理服务标签或总数异常");
|
||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
||||
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
||||
|
||||
@@ -0,0 +1,292 @@
|
||||
import { writeFileSync } from "node:fs";
|
||||
|
||||
const cdpPort = process.env.CDP_PORT ?? "9230";
|
||||
const baseUrl = process.env.SITE_URL ?? "http://127.0.0.1:4329";
|
||||
const pages = await fetch(`http://127.0.0.1:${cdpPort}/json/list`).then((response) => response.json());
|
||||
const page = pages.find((entry) => entry.type === "page");
|
||||
if (!page) throw new Error(`CDP ${cdpPort} 没有可用页面`);
|
||||
|
||||
const socket = new WebSocket(page.webSocketDebuggerUrl);
|
||||
await new Promise((resolve, reject) => {
|
||||
socket.addEventListener("open", resolve, { once: true });
|
||||
socket.addEventListener("error", reject, { once: true });
|
||||
});
|
||||
|
||||
let nextId = 0;
|
||||
const pending = new Map();
|
||||
const exceptions = [];
|
||||
socket.addEventListener("message", (event) => {
|
||||
const message = JSON.parse(event.data);
|
||||
if (message.id && pending.has(message.id)) {
|
||||
const { resolve, reject } = pending.get(message.id);
|
||||
pending.delete(message.id);
|
||||
if (message.error) reject(new Error(message.error.message));
|
||||
else resolve(message.result);
|
||||
}
|
||||
if (message.method === "Runtime.exceptionThrown") {
|
||||
exceptions.push(message.params.exceptionDetails.exception?.description ?? message.params.exceptionDetails.text);
|
||||
}
|
||||
});
|
||||
|
||||
const command = (method, params = {}) => new Promise((resolve, reject) => {
|
||||
const id = ++nextId;
|
||||
pending.set(id, { resolve, reject });
|
||||
socket.send(JSON.stringify({ id, method, params }));
|
||||
});
|
||||
const pause = (milliseconds) => new Promise((resolve) => setTimeout(resolve, milliseconds));
|
||||
const evaluate = async (expression) => {
|
||||
const result = await command("Runtime.evaluate", { expression, returnByValue: true, awaitPromise: true });
|
||||
if (result.exceptionDetails) throw new Error(result.exceptionDetails.exception?.description ?? result.exceptionDetails.text);
|
||||
return result.result.value;
|
||||
};
|
||||
const navigate = async (path) => {
|
||||
await command("Page.navigate", { url: `${baseUrl}${path}` });
|
||||
for (let attempt = 0; attempt < 100; attempt += 1) {
|
||||
await pause(100);
|
||||
if (await evaluate("document.readyState === 'complete'")) return;
|
||||
}
|
||||
throw new Error(`${path} 加载超时`);
|
||||
};
|
||||
const screenshot = async (path) => {
|
||||
const result = await command("Page.captureScreenshot", { format: "png", captureBeyondViewport: false });
|
||||
writeFileSync(path, Buffer.from(result.data, "base64"));
|
||||
};
|
||||
|
||||
await command("Page.enable");
|
||||
await command("Runtime.enable");
|
||||
await command("Emulation.setDeviceMetricsOverride", {
|
||||
width: 1440,
|
||||
height: 1100,
|
||||
deviceScaleFactor: 1,
|
||||
mobile: false,
|
||||
});
|
||||
await navigate("/k3/");
|
||||
|
||||
const desktop = await evaluate(`(() => {
|
||||
const root = document.querySelector("[data-local-path-lab]");
|
||||
root.scrollIntoView({ block: "start", behavior: "instant" });
|
||||
window.scrollBy(0, -78);
|
||||
const text = (selector) => root.querySelector(selector)?.textContent.trim();
|
||||
const panel = () => root.querySelector("[data-local-panel]:not([hidden])")?.dataset.localPanel;
|
||||
const setSelect = (selector, value) => {
|
||||
const node = root.querySelector(selector);
|
||||
node.value = value;
|
||||
node.dispatchEvent(new Event("change", { bubbles: true }));
|
||||
};
|
||||
const findMatrix = (needle) => [...root.querySelectorAll(".matrix-bars article")]
|
||||
.find((node) => node.textContent.includes(needle))?.textContent.replace(/\\s+/g, " ").trim();
|
||||
|
||||
const initial = {
|
||||
panel: panel(),
|
||||
tabs: root.querySelectorAll("[data-local-tab]").length,
|
||||
panels: root.querySelectorAll("[data-local-panel]").length,
|
||||
ledger: root.querySelectorAll(".local-ledger article").length,
|
||||
groups: root.querySelectorAll(".group-map article.target").length,
|
||||
spikeLayers: root.querySelectorAll(".group-map i.spike").length,
|
||||
boundary: root.textContent.includes("LOCALIZATION NOT ESTABLISHED") &&
|
||||
root.textContent.includes("不是贡献率") &&
|
||||
root.textContent.includes("不是 K3 checkpoint"),
|
||||
};
|
||||
|
||||
root.querySelector('[data-local-tab="matrix"]').click();
|
||||
const matrixInitial = {
|
||||
panel: panel(),
|
||||
rows: root.querySelectorAll(".matrix-bars article").length,
|
||||
state: text("[data-local-matrix-state]"),
|
||||
group67: findMatrix("仅 Groups 6+7 uniform"),
|
||||
restoration67: findMatrix("恢复 Groups 6+7"),
|
||||
};
|
||||
setSelect("[data-local-matrix-seed]", "2026073002");
|
||||
root.querySelector('[data-local-matrix-metric="peak_normalized"]').click();
|
||||
const matrixChanged = {
|
||||
rows: root.querySelectorAll(".matrix-bars article").length,
|
||||
state: text("[data-local-matrix-state]"),
|
||||
group67: findMatrix("仅 Groups 6+7 uniform"),
|
||||
restoration67: findMatrix("恢复 Groups 6+7"),
|
||||
};
|
||||
|
||||
root.querySelector('[data-local-tab="dual"]').click();
|
||||
const dual = {
|
||||
panel: panel(),
|
||||
rows: root.querySelectorAll(".dual-table tbody tr").length,
|
||||
good: root.querySelectorAll(".dual-table td.good").length,
|
||||
bad: root.querySelectorAll(".dual-table td.bad").length,
|
||||
verdict: text(".verdict-banner"),
|
||||
mean: root.querySelector(".dual-table tr.mean")?.textContent.replace(/\\s+/g, " ").trim(),
|
||||
};
|
||||
|
||||
root.querySelector('[data-local-tab="branch"]').click();
|
||||
const branchInitial = {
|
||||
panel: panel(),
|
||||
state: text("[data-local-branch-state]"),
|
||||
values: [...root.querySelectorAll("[data-local-branch-value]")].map((node) => node.textContent.trim()),
|
||||
passed: root.dataset.branchPassed,
|
||||
};
|
||||
root.querySelector('[data-local-branch-group="7"]').click();
|
||||
const branchChanged = {
|
||||
state: text("[data-local-branch-state]"),
|
||||
values: [...root.querySelectorAll("[data-local-branch-value]")].map((node) => node.textContent.trim()),
|
||||
labels: [...root.querySelectorAll("[data-local-branch-label]")].map((node) => node.textContent.trim()),
|
||||
passed: root.dataset.branchPassed,
|
||||
};
|
||||
|
||||
root.querySelector('[data-local-tab="spectrum"]').click();
|
||||
const spectrumInitial = {
|
||||
panel: panel(),
|
||||
points: root.querySelectorAll("[data-local-spectrum-points] circle").length,
|
||||
line: root.querySelector("[data-local-spectrum-line]").getAttribute("points"),
|
||||
state: text("[data-local-spectrum-state]"),
|
||||
count: text("[data-local-spectrum-count]"),
|
||||
contrast: text("[data-local-spectrum-contrast]"),
|
||||
peak: text("[data-local-spectrum-peak]"),
|
||||
layer: text("[data-local-spectrum-layer]"),
|
||||
};
|
||||
setSelect("[data-local-spectrum-seed]", "2026073002");
|
||||
setSelect("[data-local-spectrum-mode]", "uniform_groups_6_7_only");
|
||||
const spectrumGroup67 = {
|
||||
points: root.querySelectorAll("[data-local-spectrum-points] circle").length,
|
||||
line: root.querySelector("[data-local-spectrum-line]").getAttribute("points"),
|
||||
state: text("[data-local-spectrum-state]"),
|
||||
count: text("[data-local-spectrum-count]"),
|
||||
contrast: text("[data-local-spectrum-contrast]"),
|
||||
peak: text("[data-local-spectrum-peak]"),
|
||||
layer: text("[data-local-spectrum-layer]"),
|
||||
selectedPeakRadius: root.querySelector('[data-local-spectrum-points] circle[data-layer="25"]')?.getAttribute("r"),
|
||||
};
|
||||
setSelect("[data-local-spectrum-seed]", "2026073003");
|
||||
setSelect("[data-local-spectrum-mode]", "uniform_all");
|
||||
const spectrumAll = {
|
||||
points: root.querySelectorAll("[data-local-spectrum-points] circle").length,
|
||||
state: text("[data-local-spectrum-state]"),
|
||||
count: text("[data-local-spectrum-count]"),
|
||||
layer: text("[data-local-spectrum-layer]"),
|
||||
selectedPeakRadius: root.querySelector('[data-local-spectrum-points] circle[data-layer="2"]')?.getAttribute("r"),
|
||||
};
|
||||
|
||||
const first = root.querySelector('[data-local-tab="scope"]');
|
||||
first.focus();
|
||||
first.dispatchEvent(new KeyboardEvent("keydown", { key: "ArrowRight", bubbles: true }));
|
||||
const keyboard = {
|
||||
selected: root.querySelector('[data-local-tab][aria-selected="true"]').dataset.localTab,
|
||||
panel: panel(),
|
||||
};
|
||||
|
||||
return {
|
||||
initial, matrixInitial, matrixChanged, dual, branchInitial, branchChanged,
|
||||
spectrumInitial, spectrumGroup67, spectrumAll, keyboard,
|
||||
documentOverflow: document.documentElement.scrollWidth - document.documentElement.clientWidth,
|
||||
rootOverflow: root.scrollWidth - root.clientWidth,
|
||||
};
|
||||
})()`);
|
||||
await pause(180);
|
||||
await screenshot("/tmp/llm-atlas-k3-attnres-local-path-desktop.png");
|
||||
|
||||
await command("Emulation.setDeviceMetricsOverride", {
|
||||
width: 390,
|
||||
height: 844,
|
||||
deviceScaleFactor: 1,
|
||||
mobile: true,
|
||||
});
|
||||
await navigate("/k3/");
|
||||
const mobile = await evaluate(`(() => {
|
||||
const root = document.querySelector("[data-local-path-lab]");
|
||||
root.scrollIntoView({ block: "start", behavior: "instant" });
|
||||
window.scrollBy(0, -64);
|
||||
root.querySelector('[data-local-tab="dual"]').click();
|
||||
const dual = {
|
||||
panel: root.querySelector("[data-local-panel]:not([hidden])")?.dataset.localPanel,
|
||||
rows: root.querySelectorAll(".dual-table tbody tr").length,
|
||||
wrapperOverflow: root.querySelector(".dual-table-wrap").scrollWidth -
|
||||
root.querySelector(".dual-table-wrap").clientWidth,
|
||||
};
|
||||
root.querySelector('[data-local-tab="spectrum"]').click();
|
||||
return {
|
||||
tabs: root.querySelectorAll("[data-local-tab]").length,
|
||||
ledger: root.querySelectorAll(".local-ledger article").length,
|
||||
visiblePanel: root.querySelector("[data-local-panel]:not([hidden])")?.dataset.localPanel,
|
||||
points: root.querySelectorAll("[data-local-spectrum-points] circle").length,
|
||||
dual,
|
||||
documentOverflow: document.documentElement.scrollWidth - document.documentElement.clientWidth,
|
||||
rootOverflow: root.scrollWidth - root.clientWidth,
|
||||
};
|
||||
})()`);
|
||||
await pause(180);
|
||||
await screenshot("/tmp/llm-atlas-k3-attnres-local-path-mobile.png");
|
||||
|
||||
const report = { desktop, mobile, exceptions };
|
||||
console.log(JSON.stringify(report, null, 2));
|
||||
|
||||
const numeric = (value) => Number.parseFloat(value.replace("−", "-").replace("×", ""));
|
||||
const failures = [];
|
||||
if (desktop.initial.panel !== "scope" || desktop.initial.tabs !== 5 || desktop.initial.panels !== 5 ||
|
||||
desktop.initial.ledger !== 6 || desktop.initial.groups !== 2 || desktop.initial.spikeLayers !== 5) {
|
||||
failures.push("五视图、六项账本或预注册路径地图结构异常");
|
||||
}
|
||||
if (!desktop.initial.boundary) failures.push("localization / contribution / reduced-model claim boundary 缺失");
|
||||
if (desktop.matrixInitial.panel !== "matrix" || desktop.matrixInitial.rows !== 14 ||
|
||||
!desktop.matrixInitial.state.includes("3-SEED MEAN") ||
|
||||
!desktop.matrixInitial.group67.includes("+0.677") ||
|
||||
!desktop.matrixInitial.restoration67.includes("+0.650")) {
|
||||
failures.push("14-mask mean contrast 矩阵异常");
|
||||
}
|
||||
if (desktop.matrixChanged.rows !== 14 || !desktop.matrixChanged.state.includes("2026073002") ||
|
||||
!desktop.matrixChanged.state.includes("PEAK / MEAN") ||
|
||||
!desktop.matrixChanged.group67.includes("+1.783") ||
|
||||
!desktop.matrixChanged.restoration67.includes("+0.355")) {
|
||||
failures.push("matrix seed / metric 切换异常");
|
||||
}
|
||||
if (desktop.dual.panel !== "dual" || desktop.dual.rows !== 4 || desktop.dual.good !== 9 ||
|
||||
desktop.dual.bad !== 6 || !desktop.dual.verdict.includes("LOCALIZATION NOT ESTABLISHED") ||
|
||||
!desktop.dual.mean.includes("0.677") || !desktop.dual.mean.includes("0.380") ||
|
||||
!desktop.dual.mean.includes("3 / 6")) {
|
||||
failures.push("双向主门表格或冻结判定异常");
|
||||
}
|
||||
if (desktop.branchInitial.panel !== "branch" || desktop.branchInitial.passed !== "false" ||
|
||||
!desktop.branchInitial.state.includes("5 / 6") ||
|
||||
desktop.branchChanged.passed !== "true" || !desktop.branchChanged.state.includes("6 / 6 PASS") ||
|
||||
desktop.branchChanged.labels.some((value) => value !== "GROUP 7") ||
|
||||
Math.abs(numeric(desktop.branchChanged.values[0]) - .015) > .001 ||
|
||||
Math.abs(numeric(desktop.branchChanged.values[1]) - .026) > .001 ||
|
||||
Math.abs(numeric(desktop.branchChanged.values[2]) - .426) > .001 ||
|
||||
Math.abs(numeric(desktop.branchChanged.values[3]) - .823) > .001) {
|
||||
failures.push("group 6 / 7 branch gate 切换异常");
|
||||
}
|
||||
if (desktop.spectrumInitial.panel !== "spectrum" || desktop.spectrumInitial.points !== 32 ||
|
||||
desktop.spectrumInitial.count !== "0 / 65" || numeric(desktop.spectrumInitial.contrast) !== 3.093 ||
|
||||
numeric(desktop.spectrumInitial.peak) !== 3.241 || desktop.spectrumInitial.layer !== "21") {
|
||||
failures.push("reference 32 层谱异常");
|
||||
}
|
||||
if (desktop.spectrumGroup67.points !== 32 || desktop.spectrumGroup67.line === desktop.spectrumInitial.line ||
|
||||
!desktop.spectrumGroup67.state.includes("2026073002") ||
|
||||
desktop.spectrumGroup67.count !== "16 / 65" ||
|
||||
numeric(desktop.spectrumGroup67.contrast) !== 1.263 ||
|
||||
numeric(desktop.spectrumGroup67.peak) !== 1.326 ||
|
||||
desktop.spectrumGroup67.layer !== "25" || desktop.spectrumGroup67.selectedPeakRadius !== "4.5") {
|
||||
failures.push("groups 6+7 spectrum seed / mode / peak 切换异常");
|
||||
}
|
||||
if (desktop.spectrumAll.points !== 32 || desktop.spectrumAll.count !== "65 / 65" ||
|
||||
desktop.spectrumAll.layer !== "2" || desktop.spectrumAll.selectedPeakRadius !== "4.5") {
|
||||
failures.push("all-uniform spectrum selector census 或 peak 异常");
|
||||
}
|
||||
if (desktop.keyboard.selected !== "matrix" || desktop.keyboard.panel !== "matrix") {
|
||||
failures.push("键盘 tab 导航异常");
|
||||
}
|
||||
if (desktop.documentOverflow > 1 || desktop.rootOverflow > 1 ||
|
||||
mobile.documentOverflow > 1 || mobile.rootOverflow > 1) {
|
||||
failures.push("桌面或移动端出现文档级横向溢出");
|
||||
}
|
||||
if (mobile.tabs !== 5 || mobile.ledger !== 6 || mobile.visiblePanel !== "spectrum" ||
|
||||
mobile.points !== 32 || mobile.dual.panel !== "dual" || mobile.dual.rows !== 4 ||
|
||||
mobile.dual.wrapperOverflow <= 0) {
|
||||
failures.push("移动端交互结构或局部可滚动表格异常");
|
||||
}
|
||||
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||
|
||||
if (failures.length) {
|
||||
console.error(`\nFAIL\n- ${failures.join("\n- ")}`);
|
||||
process.exitCode = 1;
|
||||
} else {
|
||||
console.log("\nPASS K3 AttnRes local-path browser regression");
|
||||
}
|
||||
|
||||
socket.close();
|
||||
@@ -0,0 +1,106 @@
|
||||
import { createHash } from "node:crypto";
|
||||
import { readdirSync, readFileSync } from "node:fs";
|
||||
|
||||
const hash = (bytes) => createHash("sha256").update(bytes).digest("hex");
|
||||
const read = (path) => {
|
||||
const bytes = readFileSync(new URL(path, import.meta.url));
|
||||
return { bytes, json: JSON.parse(bytes), sha256: hash(bytes) };
|
||||
};
|
||||
const aggregate = read("../src/data/k3-attnres-local-path.json");
|
||||
const compact = read("../src/data/k3-attnres-local-path-compact.json");
|
||||
const reproduction = read("../experiments/k3/attnres_local_path/reproduction.json");
|
||||
const manifest = read("../experiments/k3/attnres_local_path/manifest.json");
|
||||
const rawDirectory = new URL("../experiments/k3/attnres_local_path/results/raw/", import.meta.url);
|
||||
const failures = [];
|
||||
const expect = (condition, message) => {
|
||||
if (!condition) failures.push(message);
|
||||
};
|
||||
const close = (actual, expected, tolerance = 1e-15) =>
|
||||
Math.abs(actual - expected) <= tolerance;
|
||||
|
||||
expect(aggregate.sha256 === "bb0ec9fce5b30d50ad5c50c4b95af7a892d614f205a2c20cfc2662125e10160e", "aggregate physical SHA-256 changed");
|
||||
expect(compact.sha256 === "3bb6c15815f37d2109386fb87f9642246d41e636242ca3a9ba3c17ae8f419c08", "compact physical SHA-256 changed");
|
||||
expect(reproduction.sha256 === "524a6883c08941c1be908834d7bcf915eb32f0bdb501aae09f3cec761fdd65db", "reproduction physical SHA-256 changed");
|
||||
expect(manifest.sha256 === "db01e92ef2cf0896212fcd529429bd94a344de0e1195db7f87b9a56dc3449139", "manifest physical SHA-256 changed");
|
||||
|
||||
expect(aggregate.json.canonical_sha256_without_self === "b86d119cd2f106e2cbee8a35760ed3244336a2fcfeb9178ea1e7dab13fc6f215", "aggregate canonical SHA-256 changed");
|
||||
expect(compact.json.canonical_sha256_without_self === "2aff9288f52d3d41bb1f59c64d9a07518ad3e2120c61615478087b24aaabd835", "compact canonical SHA-256 changed");
|
||||
expect(reproduction.json.canonical_sha256_without_self === "6f5d98fce6446fecc966dd2675f272f2c4f0c9a39a5741fabc4ffad6852ca7f4", "reproduction canonical SHA-256 changed");
|
||||
|
||||
expect(compact.json.protocol_id === "llm-atlas-k3-attnres-local-path-v1", "protocol identity mismatch");
|
||||
expect(compact.json.study.seeds.length === 3, "formal seed count changed");
|
||||
expect(compact.json.study.steps === 8000, "formal step budget changed");
|
||||
expect(compact.json.study.modes.length === 14, "matrix mode count changed");
|
||||
expect(compact.json.study.spike_layers.join(",") === "21,22,23,24,25", "fixed spike set changed");
|
||||
expect(compact.json.study.formal_target_bytes === 196608000, "formal target-byte count changed");
|
||||
expect(compact.json.study.total_target_bytes_with_replay === 262144000, "total target-byte count changed");
|
||||
expect(readdirSync(rawDirectory).filter((name) => name.endsWith(".json")).length === 4, "raw run count is not four");
|
||||
|
||||
for (const [name, expected] of Object.entries(reproduction.json.raw_files)) {
|
||||
const raw = read(`../experiments/k3/attnres_local_path/results/raw/${name}`);
|
||||
expect(raw.sha256 === expected.file_sha256, `${name} physical hash mismatch`);
|
||||
expect(raw.json.canonical_sha256_without_self === expected.canonical_sha256, `${name} canonical hash mismatch`);
|
||||
expect(raw.json.round06_equivalence.passed, `${name} Round 06 equivalence failed`);
|
||||
}
|
||||
|
||||
expect(compact.json.hashes.aggregate_canonical_sha256 === aggregate.json.canonical_sha256_without_self, "compact→aggregate canonical link mismatch");
|
||||
expect(compact.json.hashes.reproduction_canonical_sha256 === reproduction.json.canonical_sha256_without_self, "compact→reproduction canonical link mismatch");
|
||||
expect(reproduction.json.aggregate.canonical_sha256 === aggregate.json.canonical_sha256_without_self, "reproduction→aggregate canonical link mismatch");
|
||||
expect(reproduction.json.replay_gate.passed, "full replay is not exact");
|
||||
expect(reproduction.json.replay_gate.frozen_compare_sha256 === "7dbd15ad03fbd357c5d91e159706d63b24703722f76c492ed1dc733535d6b9cf", "replay compare hash changed");
|
||||
expect(reproduction.json.post_result_grok_review.blocking_errors === 0, "post-result audit reports a blocking error");
|
||||
expect(reproduction.json.post_result_grok_review.localization_status_confirmed, "post-result audit did not confirm the status");
|
||||
|
||||
const gates = compact.json.gates;
|
||||
expect(gates.all_input_and_parent_gates_passed, "input or parent gate failed");
|
||||
expect(gates.global_gap.passed && gates.global_gap.passed_cells === 6, "global gap gate changed");
|
||||
expect(gates.sufficiency.groups_6_7.passed && gates.sufficiency.groups_6_7.passed_cells === 6, "groups 6+7 sufficiency gate failed");
|
||||
expect(!gates.restoration.groups_6_7.passed && gates.restoration.groups_6_7.passed_cells === 3, "groups 6+7 restoration verdict changed");
|
||||
expect(!gates.localization.passed, "localization unexpectedly passed");
|
||||
expect(gates.localization.status === "one_sided_evidence_localization_not_established", "localization status changed");
|
||||
expect(gates.sufficiency.group_6.passed && gates.sufficiency.group_7.passed, "single-group sufficiency gate changed");
|
||||
expect(!gates.restoration.group_6.passed && !gates.restoration.group_7.passed, "single-group restoration unexpectedly passed");
|
||||
expect(!gates.sufficiency.output_half_gap.passed && gates.sufficiency.output_half_gap.passed_cells === 0, "output half-gap control changed");
|
||||
expect(!gates.branch_dominance.group_6.passed, "group 6 branch dominance unexpectedly passed");
|
||||
expect(gates.branch_dominance.group_7.passed && gates.branch_dominance.group_7.dominant_branch === "mlp", "group 7 MLP dominance changed");
|
||||
|
||||
expect(close(compact.json.means.sufficiency.uniform_groups_6_7_only.spike_contrast, 0.6769080107044138), "groups 6+7 mean contrast sufficiency changed");
|
||||
expect(close(compact.json.means.sufficiency.uniform_groups_6_7_only.peak_normalized, 1.7003398291502192), "groups 6+7 mean peak sufficiency changed");
|
||||
expect(close(compact.json.means.restoration.uniform_except_groups_6_7.spike_contrast, 0.6499896714884171), "groups 6+7 mean contrast restoration changed");
|
||||
expect(close(compact.json.means.restoration.uniform_except_groups_6_7.peak_normalized, 0.3801045652660336), "groups 6+7 mean peak restoration changed");
|
||||
expect(close(compact.json.means.sufficiency.uniform_output_only.spike_contrast, 0.13066777032743535), "output-only mean contrast score changed");
|
||||
expect(close(compact.json.means.sufficiency.uniform_output_only.peak_normalized, 0.18724852586461135), "output-only mean peak score changed");
|
||||
|
||||
for (const spectrum of compact.json.final_spectra) {
|
||||
expect(Object.keys(spectrum.modes).length === 14, `seed ${spectrum.seed} spectrum mode count changed`);
|
||||
expect(spectrum.modes.detached_learned.uniform_count === 0, `seed ${spectrum.seed} reference census changed`);
|
||||
expect(spectrum.modes.uniform_groups_6_7_only.uniform_count === 16, `seed ${spectrum.seed} groups 6+7 census changed`);
|
||||
expect(spectrum.modes.uniform_except_groups_6_7.uniform_count === 49, `seed ${spectrum.seed} restoration census changed`);
|
||||
expect(spectrum.modes.uniform_all.uniform_count === 65, `seed ${spectrum.seed} all-uniform census changed`);
|
||||
for (const mode of Object.values(spectrum.modes)) {
|
||||
expect(mode.normalized.length === 32, `seed ${spectrum.seed} normalized spectrum length changed`);
|
||||
}
|
||||
}
|
||||
|
||||
if (failures.length) {
|
||||
console.error(`FAIL K3 AttnRes local-path data\n- ${failures.join("\n- ")}`);
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
console.log(JSON.stringify({
|
||||
protocol: compact.json.protocol_id,
|
||||
modes: compact.json.study.modes.length,
|
||||
formalRuns: compact.json.study.seeds.length,
|
||||
replayExact: reproduction.json.replay_gate.passed,
|
||||
globalGap: gates.global_gap,
|
||||
sufficiency: gates.sufficiency.groups_6_7,
|
||||
restoration: gates.restoration.groups_6_7,
|
||||
localization: gates.localization,
|
||||
branch: gates.branch_dominance,
|
||||
hashes: {
|
||||
aggregate: aggregate.sha256,
|
||||
compact: compact.sha256,
|
||||
reproduction: reproduction.sha256,
|
||||
},
|
||||
}, null, 2));
|
||||
console.log("PASS K3 AttnRes local-path frozen data");
|
||||
@@ -85,6 +85,9 @@ const overview = await evaluate(`(() => ({
|
||||
gradientPanels: document.querySelectorAll("[data-gradient-panel]").length,
|
||||
spikeTabs: document.querySelectorAll("[data-spike-tab]").length,
|
||||
spikePanels: document.querySelectorAll("[data-spike-panel]").length,
|
||||
localPathTabs: document.querySelectorAll("[data-local-tab]").length,
|
||||
localPathPanels: document.querySelectorAll("[data-local-panel]").length,
|
||||
localPathVerdict: document.querySelector("#attnres-local-path")?.textContent.includes("localization 未建立"),
|
||||
nativeVisionCorrected: document.body.textContent.includes("MoonViT‑V2 从头训练") &&
|
||||
document.body.textContent.includes("同一个 next-token prediction objective"),
|
||||
staleVisionClaim: document.body.textContent.includes("先固定语言模型训练视觉组件"),
|
||||
@@ -296,8 +299,9 @@ const mobile = await evaluate(`(() => {
|
||||
attnresTabs: document.querySelectorAll("[data-attnres-tab]").length,
|
||||
gradientTabs: document.querySelectorAll("[data-gradient-tab]").length,
|
||||
spikeTabs: document.querySelectorAll("[data-spike-tab]").length,
|
||||
localPathTabs: document.querySelectorAll("[data-local-tab]").length,
|
||||
offenders: [...document.querySelectorAll("body *")]
|
||||
.filter((node) => !node.closest(".paper-chain, .spec-table-wrap, .cache-strip, .architecture-explorer, [data-k3-lab], [data-k3-artifact-lab], [data-attnres-lab], [data-gradient-lab], [data-spike-lab]"))
|
||||
.filter((node) => !node.closest(".paper-chain, .spec-table-wrap, .cache-strip, .architecture-explorer, [data-k3-lab], [data-k3-artifact-lab], [data-attnres-lab], [data-gradient-lab], [data-spike-lab], [data-local-path-lab]"))
|
||||
.filter((node) => node.getBoundingClientRect().right > document.documentElement.clientWidth + 1)
|
||||
.slice(0, 15)
|
||||
.map((node) => ({
|
||||
@@ -328,7 +332,7 @@ console.log(JSON.stringify(report, null, 2));
|
||||
const numeric = (text) => Number.parseFloat(text.replaceAll(",", "").replace("−", "-"));
|
||||
const failures = [];
|
||||
if (!overview.title.includes("因果环节")) failures.push("K3 二轮标题异常");
|
||||
if (overview.sections !== 35 || overview.tocLinks !== 35) failures.push("34 个编号专题加阅读链的目录结构异常");
|
||||
if (overview.sections !== 36 || overview.tocLinks !== 36) failures.push("35 个编号专题加阅读链的目录结构异常");
|
||||
if (overview.ledgers !== 32 || overview.reportMap !== 9) failures.push("32 张问题账或报告地图异常");
|
||||
if (overview.figureAtlas !== 21 || overview.paperLinks !== 100 || overview.paperGroups < 12) failures.push("图表审计或 100 节点阅读链异常");
|
||||
if (overview.labTabs !== 8 || overview.labPanels !== 8) failures.push("八联实验结构异常");
|
||||
@@ -336,6 +340,7 @@ if (overview.artifactTabs !== 4 || overview.artifactPanels !== 4 || overview.art
|
||||
if (overview.attnresTabs !== 5 || overview.attnresPanels !== 5) failures.push("AttnRes 独立实验五视图异常");
|
||||
if (overview.gradientTabs !== 5 || overview.gradientPanels !== 5) failures.push("AttnRes 梯度定义扩展五视图异常");
|
||||
if (overview.spikeTabs !== 5 || overview.spikePanels !== 5) failures.push("AttnRes 尖峰路径五视图异常");
|
||||
if (overview.localPathTabs !== 5 || overview.localPathPanels !== 5 || !overview.localPathVerdict) failures.push("AttnRes 局部路径五视图或冻结判定异常");
|
||||
if (!overview.nativeVisionCorrected || overview.staleVisionClaim) failures.push("原生多模态纠错未生效或旧错误残留");
|
||||
if (overview.documentOverflow > 1 || mobile.documentOverflow > 1) failures.push("桌面或移动端存在文档级横向溢出");
|
||||
if (labs.memoryInitial.panel !== "memory" || numeric(labs.memoryInitial.additiveError) <= numeric(labs.memoryInitial.deltaError)) failures.push("Delta memory 初始递推异常");
|
||||
@@ -360,7 +365,7 @@ if (artifacts.parameterChanged.shape !== "[96,128] F32" || !artifacts.parameterC
|
||||
if (artifacts.reproductionInitial.panel !== "reproduction" || numeric(artifacts.reproductionInitial.speedup) !== 1.85 || numeric(artifacts.reproductionInitial.localMean) < 2.6 || !artifacts.reproductionInitial.exactSuite || numeric(artifacts.reproductionInitial.cv) < 2) failures.push("FlashKDA H20、本机 exact suite 或 router 初始探针异常");
|
||||
if (numeric(artifacts.reproductionChanged.speedup) !== 3.27 || numeric(artifacts.reproductionChanged.flash) !== 0.7064 || numeric(artifacts.reproductionChanged.localMean) >= numeric(artifacts.reproductionInitial.localMean) || !artifacts.reproductionChanged.localMode.includes("FP32 state") || numeric(artifacts.reproductionChanged.cv) <= numeric(artifacts.reproductionInitial.cv) || numeric(artifacts.reproductionChanged.zero) <= numeric(artifacts.reproductionInitial.zero)) failures.push("GB200 benchmark、本机 varlen/state 或 synthetic router counterexample 未更新");
|
||||
if (artifacts.keyboardSelected !== "tensors" || artifacts.keyboardVisible !== "tensors") failures.push("开放工件键盘 tab 导航异常");
|
||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 8 || mobile.artifactTabs !== 4 || mobile.artifactLayers !== 93 || mobile.attnresTabs !== 5 || mobile.gradientTabs !== 5 || mobile.spikeTabs !== 5) failures.push("移动端导航或实验异常");
|
||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 8 || mobile.artifactTabs !== 4 || mobile.artifactLayers !== 93 || mobile.attnresTabs !== 5 || mobile.gradientTabs !== 5 || mobile.spikeTabs !== 5 || mobile.localPathTabs !== 5) failures.push("移动端导航或实验异常");
|
||||
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
||||
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||
|
||||
|
||||
@@ -246,7 +246,7 @@ if (ocr.unreported.status !== "OUT OF EVIDENCE" || ocr.unreported.accuracy !== "
|
||||
if (loop.toolsStart.state !== "OPEN" || loop.toolsEnd.state !== "VERIFIED" || loop.toolsEnd.evidence !== "97%" || loop.toolsEnd.tools !== "3") failures.push("vision-in-the-loop 终局异常");
|
||||
if (loop.cotEnd.state !== "FAILED" || !loop.cotEnd.takeaway.includes("不能凭空增加")) failures.push("文字 CoT 与新观察没有分开");
|
||||
if (loop.keyboardSelected !== "connector" || loop.keyboardVisible !== "connector") failures.push("实验键盘 tab 导航异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("尖峰不是出生时就有") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
|
||||
if (home.paperCount !== "486" || papers.total !== 486 || !papers.hasFilter || papers.multimodalVisible < 59) failures.push("论文库多模态标签或总数异常");
|
||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
||||
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
||||
|
||||
@@ -277,7 +277,7 @@ if (layout.navLinks !== 20 || mobile.mobileLinks !== 20 || home.navLinks !== 20)
|
||||
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
|
||||
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
||||
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("尖峰不是出生时就有") || home.firstHref !== "/k3/") {
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") {
|
||||
failures.push("首页 Transformer 新章入口异常");
|
||||
}
|
||||
if (home.paperCount !== "486") failures.push(`首页论文总数异常:${home.paperCount}`);
|
||||
|
||||
@@ -289,7 +289,7 @@ if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentO
|
||||
}
|
||||
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
||||
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("尖峰不是出生时就有")) failures.push("首页 K3 首发入口异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强")) failures.push("首页 K3 首发入口异常");
|
||||
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
|
||||
|
||||
socket.close();
|
||||
|
||||
@@ -288,7 +288,7 @@ if (numeric(residual.attnres.states) !== 9 || !residual.attnres.routeExplain.inc
|
||||
if (!residual.clamp.activation.includes("V4") || !residual.clamp.bound.includes("100")) failures.push("DeepSeek-V4 clamp 展示异常");
|
||||
if (!residual.situ.activation.includes("KIMI") || !residual.situ.bound.includes("100")) failures.push("K3 SiTU 上界展示异常");
|
||||
if (residual.keyboardSelected !== "position" || residual.keyboardVisible !== "position") failures.push("实验键盘 tab 导航异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("尖峰不是出生时就有") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
|
||||
if (home.paperCount !== "486" || home.topicCount !== "17" || papers.total !== 486 || !papers.hasFilter || papers.visible < 30) failures.push("首页 / 论文库表示索引异常");
|
||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
||||
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
||||
|
||||
@@ -273,7 +273,7 @@ if (layout.navLinks !== 20 || mobile.mobileLinks !== 20 || home.navLinks !== 20)
|
||||
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
|
||||
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
||||
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("尖峰不是出生时就有") || home.firstHref !== "/k3/") {
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") {
|
||||
failures.push("首页 Transformer 新章入口异常");
|
||||
}
|
||||
if (home.paperCount !== "486") failures.push(`首页论文总数异常:${home.paperCount}`);
|
||||
|
||||
@@ -233,7 +233,7 @@ if (layout.articleSections !== 16 || layout.paperLinks !== 37 || layout.labTabs
|
||||
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
|
||||
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
|
||||
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("尖峰不是出生时就有")) failures.push("首页 K3 首发入口异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强")) failures.push("首页 K3 首发入口异常");
|
||||
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
|
||||
|
||||
socket.close();
|
||||
|
||||
@@ -236,7 +236,7 @@ if (block.family.trim() !== "Hybrid MoE" || !block.kv.includes("3 KDA : 1 Gated
|
||||
if (!block.path.some((step) => step.includes("KDA × 3")) || !block.note.includes("AttnRes")) failures.push("K3 Block 路径异常");
|
||||
if (block.context.trim() !== "128K" || numeric(block.mha) !== 400 || numeric(block.kda) !== 1) failures.push("KV 成本缩放异常");
|
||||
if (block.keyboardSelected !== "block" || block.keyboardVisible !== "block") failures.push("实验 tab 键盘导航异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("尖峰不是出生时就有") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
|
||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
|
||||
if (home.paperCount !== "486" || papers.total !== 486 || papers.transformerVisible < 30) failures.push("论文库或首页论文数量异常");
|
||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
|
||||
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||
|
||||
@@ -334,7 +334,7 @@ const benchmarkDevices = [
|
||||
<p><span>TOP-16 OVERLAP</span><b>{probe.membership_overlap_mean.toFixed(2)} / 16</b></p>
|
||||
</div>
|
||||
</div>
|
||||
<div class="boundary"><b>X/S/U boundary</b><p>本机 kernel 实测只验证公开 FlashKDA API 与合成合法 shape;没有加载 K3 checkpoint,也不解决 checkpoint `A_log [128]` 与 API `[96]` 的冲突。Router counterexample 仍只证明 hidden distribution 不可省略。</p></div>
|
||||
<div class="boundary"><b>X/S/U boundary</b><p>本机 kernel 实测只验证公开 FlashKDA API 与合成合法 shape;没有加载 K3 checkpoint。官方 main 仍存在 `A_log [128]` 与 API `[96]` 的冲突;社区 #144 与 #150 给出两种未合并候选,本站不替官方裁决。Router counterexample 仍只证明 hidden distribution 不可省略。</p></div>
|
||||
</section>
|
||||
|
||||
<footer class="evidence-strip">
|
||||
|
||||
@@ -0,0 +1,707 @@
|
||||
---
|
||||
import rawLab from "@/data/k3-attnres-local-path-compact.json";
|
||||
|
||||
const lab = rawLab as any;
|
||||
const json = JSON.stringify(lab).replaceAll("<", "\\u003c");
|
||||
const seeds = lab.study.seeds as number[];
|
||||
const shortHash = (value: string) => `${value.slice(0, 10)}…${value.slice(-8)}`;
|
||||
const modeLabels: Record<string, string> = {
|
||||
detached_learned: "REFERENCE · 全部 detached learned",
|
||||
uniform_group_6_only: "仅 Group 6 uniform",
|
||||
uniform_group_7_only: "仅 Group 7 uniform",
|
||||
uniform_groups_6_7_only: "仅 Groups 6+7 uniform",
|
||||
uniform_group_6_attention_only: "Group 6 · Attention only",
|
||||
uniform_group_6_mlp_only: "Group 6 · MLP only",
|
||||
uniform_group_7_attention_only: "Group 7 · Attention only",
|
||||
uniform_group_7_mlp_only: "Group 7 · MLP only",
|
||||
uniform_output_only: "仅 Output mixer uniform",
|
||||
uniform_depth_all: "64 个 depth mixers uniform",
|
||||
uniform_all: "全部 65 个 mixers uniform",
|
||||
uniform_except_group_6: "恢复 Group 6 → detached learned",
|
||||
uniform_except_group_7: "恢复 Group 7 → detached learned",
|
||||
uniform_except_groups_6_7: "恢复 Groups 6+7 → detached learned",
|
||||
};
|
||||
const score = (family: "sufficiency" | "restoration", mode: string, seed: number, metric: string) =>
|
||||
lab.scores[family][mode][String(seed)][metric].score;
|
||||
---
|
||||
|
||||
<figure class="local-lab" data-local-path-lab>
|
||||
<figcaption>
|
||||
<span>ROUND 07 / LOCAL MIXER PATHS</span>
|
||||
<div>
|
||||
<h3>把“全局敏感”缩到 16 个 mixer:为什么单侧很强,双向门仍然不让过?</h3>
|
||||
<p>同一 forward · 14 个冻结 mask · 3 seeds · sufficiency × restoration · 完整 replay</p>
|
||||
</div>
|
||||
<em>REDUCED-MODEL DIAGNOSTIC</em>
|
||||
</figcaption>
|
||||
|
||||
<div class="local-ledger">
|
||||
<article><span>TOPOLOGY</span><b>64 + 1</b><p>depth mixers + output</p></article>
|
||||
<article><span>MATRIX</span><b>14 modes</b><p>exact selector sets</p></article>
|
||||
<article><span>GLOBAL GAP</span><b>6 / 6</b><p>two metrics × three seeds</p></article>
|
||||
<article class="pass"><span>SUFFICIENCY</span><b>6 / 6</b><p>Groups 6+7 ≥ 50%</p></article>
|
||||
<article class="warn"><span>RESTORATION</span><b>3 / 6</b><p>contrast pass · peak fail</p></article>
|
||||
<article class="boundary"><span>LOCALIZATION</span><b>NOT ESTABLISHED</b><p>one-sided evidence</p></article>
|
||||
</div>
|
||||
|
||||
<div class="local-tabs" role="tablist" aria-label="Round 07 局部 mixer 路径实验视图">
|
||||
<button type="button" role="tab" data-local-tab="scope" aria-selected="true">01 / PATH MAP</button>
|
||||
<button type="button" role="tab" data-local-tab="matrix" aria-selected="false">02 / 14 MASKS</button>
|
||||
<button type="button" role="tab" data-local-tab="dual" aria-selected="false">03 / TWO-WAY GATE</button>
|
||||
<button type="button" role="tab" data-local-tab="branch" aria-selected="false">04 / BRANCH × OUTPUT</button>
|
||||
<button type="button" role="tab" data-local-tab="spectrum" aria-selected="false">05 / SPECTRUM × AUDIT</button>
|
||||
</div>
|
||||
|
||||
<section class="local-panel" data-local-panel="scope">
|
||||
<div class="panel-lead">
|
||||
<div><span>I / FIXED TOPOLOGY</span><h4>先画清 65 个 intervention nodes,再谈“局部”</h4></div>
|
||||
<p>
|
||||
layer 21–25 是 Round 05 看过数据后冻结的 spike set;Round 07 预先选择完整
|
||||
group 6 / 7。group 7 还包含 S 外的 layers 26–28,所以不是事后只挑尖峰层。
|
||||
</p>
|
||||
</div>
|
||||
<div class="group-map">
|
||||
{[1,2,3,4,5,6,7,8].map((group) => (
|
||||
<article class:list={{ target: group === 6 || group === 7 }}>
|
||||
<span>GROUP {group}</span>
|
||||
<b>L{(group - 1) * 4 + 1}–{group * 4}</b>
|
||||
<div>
|
||||
{[0,1,2,3].map((offset) => {
|
||||
const layer = (group - 1) * 4 + offset + 1;
|
||||
return <i class:list={{ spike: layer >= 21 && layer <= 25 }}>{layer}</i>;
|
||||
})}
|
||||
</div>
|
||||
<small>8 MIXERS</small>
|
||||
</article>
|
||||
))}
|
||||
<i class="map-arrow">→</i>
|
||||
<article class="output-node"><span>OUTPUT</span><b>#65</b><p>9 sources</p></article>
|
||||
</div>
|
||||
<div class="scope-key">
|
||||
<span><i class="target"></i>预注册局部 scope:groups 6+7 / 16 mixers</span>
|
||||
<span><i class="spike"></i>固定 spike set:layers 21–25</span>
|
||||
<span><i class="plain"></i>其他 49 mixers</span>
|
||||
</div>
|
||||
<div class="direction-pair">
|
||||
<article>
|
||||
<span>A / SUFFICIENCY</span>
|
||||
<div><b>其他 49:learned</b><i>+</i><b class="accent">G6+7:uniform</b></div>
|
||||
<code>Sₓ = ln(Xref / Xm) ÷ Gₓ</code>
|
||||
<p>只改这 16 个,能否复现全局下降的一半?</p>
|
||||
</article>
|
||||
<i>⇄</i>
|
||||
<article>
|
||||
<span>B / RESTORATION</span>
|
||||
<div><b>其他 49:uniform</b><i>+</i><b class="accent">G6+7:learned</b></div>
|
||||
<code>Rₓ = ln(Xr / Xall) ÷ Gₓ</code>
|
||||
<p>只恢复这 16 个,能否把全局 gap 恢复一半?</p>
|
||||
</article>
|
||||
</div>
|
||||
<div class="plain-rule">
|
||||
<b>为什么要两个方向?</b>
|
||||
<p>mixer 路径非线性交互。一个 scope 在 learned 背景“足够强”,不代表它在 uniform 背景“恢复得回来”。</p>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section class="local-panel" data-local-panel="matrix" hidden>
|
||||
<div class="panel-lead">
|
||||
<div><span>II / FROZEN MATRIX</span><h4>把 14 个 mask 全部摆出来,不只展示通过的 scope</h4></div>
|
||||
<p>条形长度是 global log gap 的归一化 score,不是贡献率;负值和大于 1 都保留。</p>
|
||||
</div>
|
||||
<div class="local-controls">
|
||||
<label>SEED
|
||||
<select data-local-matrix-seed>
|
||||
<option value="mean">3-SEED MEAN</option>
|
||||
{seeds.map((seed) => <option value={String(seed)}>{seed}</option>)}
|
||||
</select>
|
||||
</label>
|
||||
<div>
|
||||
<button type="button" data-local-matrix-metric="spike_contrast" aria-pressed="true">SPIKE CONTRAST</button>
|
||||
<button type="button" data-local-matrix-metric="peak_normalized" aria-pressed="false">PEAK / MEAN</button>
|
||||
</div>
|
||||
<span data-local-matrix-state>3-SEED MEAN · SPIKE CONTRAST</span>
|
||||
</div>
|
||||
<div class="matrix-axis"><span>−</span><i></i><b>0</b><i></i><b>.5</b><i></i><b>1.0</b><i></i><span>2.0+</span></div>
|
||||
<div class="matrix-bars" data-local-matrix-bars></div>
|
||||
<div class="matrix-note">
|
||||
<article><b>S</b><p>learned 背景 → scope uniform</p></article>
|
||||
<article><b>R</b><p>uniform 背景 → scope learned</p></article>
|
||||
<article><b>>1</b><p>超过 all-uniform endpoint;不是 >100% 贡献</p></article>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section class="local-panel" data-local-panel="dual" hidden>
|
||||
<div class="panel-lead">
|
||||
<div><span>III / PREREGISTERED VERDICT</span><h4>contrast 双向过线;peak 只在 sufficiency 方向过线</h4></div>
|
||||
<p>主门要求两个方向、两个指标、三个 seed 全部 ≥ .50;均值只用于视觉摘要。</p>
|
||||
</div>
|
||||
<div class="verdict-banner">
|
||||
<span>ONE-SIDED EVIDENCE</span>
|
||||
<b>LOCALIZATION NOT ESTABLISHED</b>
|
||||
<p>SUFFICIENCY 6 / 6 PASS · RESTORATION 3 / 6 FAIL</p>
|
||||
</div>
|
||||
<div class="dual-table-wrap">
|
||||
<table class="dual-table">
|
||||
<thead><tr><th>SEED</th><th>S / CONTRAST</th><th>S / PEAK</th><th>R / CONTRAST</th><th>R / PEAK</th><th>DUAL</th></tr></thead>
|
||||
<tbody>
|
||||
{seeds.map((seed) => {
|
||||
const sc = score("sufficiency", "uniform_groups_6_7_only", seed, "spike_contrast");
|
||||
const sp = score("sufficiency", "uniform_groups_6_7_only", seed, "peak_normalized");
|
||||
const rc = score("restoration", "uniform_except_groups_6_7", seed, "spike_contrast");
|
||||
const rp = score("restoration", "uniform_except_groups_6_7", seed, "peak_normalized");
|
||||
return (
|
||||
<tr>
|
||||
<th>{seed}</th>
|
||||
<td class="good">{sc.toFixed(3)} ✓</td>
|
||||
<td class="good">{sp.toFixed(3)} ✓</td>
|
||||
<td class="good">{rc.toFixed(3)} ✓</td>
|
||||
<td class="bad">{rp.toFixed(3)} ×</td>
|
||||
<td class="bad">FAIL</td>
|
||||
</tr>
|
||||
);
|
||||
})}
|
||||
<tr class="mean">
|
||||
<th>MEAN</th>
|
||||
<td>{lab.means.sufficiency.uniform_groups_6_7_only.spike_contrast.toFixed(3)}</td>
|
||||
<td>{lab.means.sufficiency.uniform_groups_6_7_only.peak_normalized.toFixed(3)}</td>
|
||||
<td>{lab.means.restoration.uniform_except_groups_6_7.spike_contrast.toFixed(3)}</td>
|
||||
<td>{lab.means.restoration.uniform_except_groups_6_7.peak_normalized.toFixed(3)}</td>
|
||||
<td>3 / 6</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
</div>
|
||||
<div class="dual-visual">
|
||||
<article class="pass">
|
||||
<span>LEARNED BACKGROUND</span><b>只 uniform G6+7</b>
|
||||
<div><i style="--score:.677"></i><em>CONTRAST .677</em></div>
|
||||
<div><i style="--score:1"></i><em>PEAK 1.700</em></div>
|
||||
<p>16 个 mixer 足以把两项都推过 50% global log gap。</p>
|
||||
</article>
|
||||
<article class="fail">
|
||||
<span>UNIFORM BACKGROUND</span><b>只 restore G6+7</b>
|
||||
<div><i style="--score:.650"></i><em>CONTRAST .650</em></div>
|
||||
<div><i style="--score:.380"></i><em>PEAK .380</em></div>
|
||||
<p>contrast 回升;peak 三 seed 都稳定停在 50% 以下。</p>
|
||||
</article>
|
||||
</div>
|
||||
<div class="boundary-pair">
|
||||
<article class="yes"><span>可以说</span><b>强局部 sufficiency</b><p>Groups 6+7 在 learned 背景足以复现全局下降的大部分。</p></article>
|
||||
<article class="no"><span>不能说</span><b>尖峰 localization 到 G6+7</b><p>预注册 restoration peak 门没有通过;双向证据不闭合。</p></article>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section class="local-panel" data-local-panel="branch" hidden>
|
||||
<div class="panel-lead">
|
||||
<div><span>IV / SECONDARY CONTROLS</span><h4>group 7 的 MLP 分支过闸;group 6 差一格,output 远不到一半</h4></div>
|
||||
<p>branch 判定只有 sufficiency 方向,证据等级低于双向 localization。</p>
|
||||
</div>
|
||||
<div class="local-controls">
|
||||
<div>
|
||||
<button type="button" data-local-branch-group="6" aria-pressed="true">GROUP 6</button>
|
||||
<button type="button" data-local-branch-group="7" aria-pressed="false">GROUP 7</button>
|
||||
</div>
|
||||
<span data-local-branch-state>GROUP 6 · NO DOMINANT BRANCH · MLP 5 / 6</span>
|
||||
</div>
|
||||
<div class="branch-chart">
|
||||
<article>
|
||||
<span>ATTENTION-ONLY</span><b data-local-branch-label="attention">GROUP 6</b>
|
||||
<div><em>CONTRAST</em><i data-local-branch-bar="attention:spike_contrast"></i><strong data-local-branch-value="attention:spike_contrast">.082</strong></div>
|
||||
<div><em>PEAK</em><i data-local-branch-bar="attention:peak_normalized"></i><strong data-local-branch-value="attention:peak_normalized">.246</strong></div>
|
||||
</article>
|
||||
<article class="mlp">
|
||||
<span>MLP-ONLY</span><b data-local-branch-label="mlp">GROUP 6</b>
|
||||
<div><em>CONTRAST</em><i data-local-branch-bar="mlp:spike_contrast"></i><strong data-local-branch-value="mlp:spike_contrast">.250</strong></div>
|
||||
<div><em>PEAK</em><i data-local-branch-bar="mlp:peak_normalized"></i><strong data-local-branch-value="mlp:peak_normalized">.950</strong></div>
|
||||
</article>
|
||||
</div>
|
||||
<div class="branch-verdicts">
|
||||
<article><span>GROUP 6</span><b>MLP 5 / 6</b><p>seed 3 contrast S=.177;低于 material .20,不能宣布 dominance。</p></article>
|
||||
<article class="pass"><span>GROUP 7</span><b>MLP 6 / 6 PASS</b><p>两个指标、三个 seed 自身 material,且领先 attention 至少 15pp。</p></article>
|
||||
</div>
|
||||
<div class="control-grid">
|
||||
<article><span>OUTPUT ONLY · 1 MIXER</span><b>.131 / .187</b><p>contrast / peak mean sufficiency;0 / 6 达到 50%。</p></article>
|
||||
<article><span>ALL DEPTH · 64 MIXERS</span><b>.949 / 1.145</b><p>depth path 已复现绝大多数 gap;peak 超过 global endpoint。</p></article>
|
||||
<article class="warning"><span>INTERACTION RESIDUAL</span><b>NEGATIVE</b><p>只作 log-gap bookkeeping;不是加和分解或 hypothesis test。</p></article>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section class="local-panel" data-local-panel="spectrum" hidden>
|
||||
<div class="panel-lead">
|
||||
<div><span>V / RAW SHAPE × REPLAY</span><h4>同一个 score 背后,32 层峰值究竟移到哪里?</h4></div>
|
||||
<p>谱按每个 mode 自己的 layer mean 归一化。位置变化帮助理解,但不替代双向 score gate。</p>
|
||||
</div>
|
||||
<div class="local-controls spectrum-controls">
|
||||
<label>SEED
|
||||
<select data-local-spectrum-seed>
|
||||
{seeds.map((seed) => <option value={String(seed)}>{seed}</option>)}
|
||||
</select>
|
||||
</label>
|
||||
<label>MODE
|
||||
<select data-local-spectrum-mode>
|
||||
{lab.study.modes.map((mode: string) => <option value={mode}>{modeLabels[mode]}</option>)}
|
||||
</select>
|
||||
</label>
|
||||
<span data-local-spectrum-state>2026073001 · REFERENCE</span>
|
||||
</div>
|
||||
<div class="spectrum-layout">
|
||||
<div class="chart-shell">
|
||||
<header><b>POST-MLP · NORMALIZED ELEMENT RMS</b><span>layer mean = 1</span></header>
|
||||
<svg viewBox="0 0 940 350" role="img" aria-label="局部 mixer mask 下的 32 层梯度谱" data-local-spectrum-chart>
|
||||
<rect class="spike-zone" x="0" y="26" width="0" height="280" data-local-spectrum-zone></rect>
|
||||
<g data-local-spectrum-grid></g>
|
||||
<polyline class="series" points="" data-local-spectrum-line></polyline>
|
||||
<g data-local-spectrum-points></g>
|
||||
</svg>
|
||||
<div class="spectrum-key"><span><i></i>normalized layer gradient</span><span><i class="zone"></i>S = layers 21–25</span></div>
|
||||
</div>
|
||||
<aside>
|
||||
<span>SELECTED READOUT</span>
|
||||
<b data-local-spectrum-label>REFERENCE</b>
|
||||
<dl>
|
||||
<div><dt>UNIFORM MIXERS</dt><dd data-local-spectrum-count>0 / 65</dd></div>
|
||||
<div><dt>SPIKE CONTRAST</dt><dd data-local-spectrum-contrast>3.093×</dd></div>
|
||||
<div><dt>PEAK / MEAN</dt><dd data-local-spectrum-peak>3.241×</dd></div>
|
||||
<div><dt>PEAK LAYER</dt><dd data-local-spectrum-layer>21</dd></div>
|
||||
</dl>
|
||||
</aside>
|
||||
</div>
|
||||
<div class="spectrum-story">
|
||||
<article><span>REFERENCE</span><b>L21 · 3 / 3</b><p>三个 seed 的最高层都是 layer 21。</p></article>
|
||||
<i>→</i>
|
||||
<article><span>G6+7 UNIFORM</span><b>L5 / L25 / L6</b><p>peak 强度压到 global endpoint 以下,但位置不统一。</p></article>
|
||||
<i>→</i>
|
||||
<article><span>RESTORE G6+7</span><b>L21 · 3 / 3</b><p>位置回归不等于强度恢复超过 50%。</p></article>
|
||||
</div>
|
||||
<div class="audit-grid">
|
||||
<article><span>ROUND 06 EQUIVALENCE</span><b>3 / 3 exact</b><p>model、optimizer、history、BPC 与 parent diagnostics。</p></article>
|
||||
<article><span>FULL REPLAY</span><b>canonical exact</b><p><code>{shortHash(lab.replay.frozen_compare_sha256)}</code></p></article>
|
||||
<article><span>FORWARD IDENTITY</span><b>14 / 14</b><p>logits、loss、activations 与 parent summaries exact。</p></article>
|
||||
<article class="boundary"><span>CLAIM BOUNDARY</span><b>diagnostic only</b><p>不是 K3 checkpoint、训练变体或 additive attribution。</p></article>
|
||||
</div>
|
||||
<div class="hash-strip">
|
||||
<span>PROTOCOL <code>{lab.protocol_id}</code></span>
|
||||
<span>AGGREGATE <code>{shortHash(lab.hashes.aggregate_canonical_sha256)}</code></span>
|
||||
<span>COMPACT <code>{shortHash(lab.canonical_sha256_without_self)}</code></span>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<script is:inline type="application/json" data-local-payload set:html={json}></script>
|
||||
</figure>
|
||||
|
||||
<script>
|
||||
const initializeLocalPathLab = (root: HTMLElement) => {
|
||||
if (root.dataset.ready === "true") return;
|
||||
root.dataset.ready = "true";
|
||||
const payload = root.querySelector<HTMLScriptElement>("[data-local-payload]");
|
||||
if (!payload) return;
|
||||
const data = JSON.parse(payload.textContent || "{}");
|
||||
const modeLabels: Record<string, string> = {
|
||||
detached_learned: "REFERENCE · 全部 detached learned",
|
||||
uniform_group_6_only: "仅 Group 6 uniform",
|
||||
uniform_group_7_only: "仅 Group 7 uniform",
|
||||
uniform_groups_6_7_only: "仅 Groups 6+7 uniform",
|
||||
uniform_group_6_attention_only: "Group 6 · Attention only",
|
||||
uniform_group_6_mlp_only: "Group 6 · MLP only",
|
||||
uniform_group_7_attention_only: "Group 7 · Attention only",
|
||||
uniform_group_7_mlp_only: "Group 7 · MLP only",
|
||||
uniform_output_only: "仅 Output mixer uniform",
|
||||
uniform_depth_all: "64 个 depth mixers uniform",
|
||||
uniform_all: "全部 65 个 mixers uniform",
|
||||
uniform_except_group_6: "恢复 Group 6 → detached learned",
|
||||
uniform_except_group_7: "恢复 Group 7 → detached learned",
|
||||
uniform_except_groups_6_7: "恢复 Groups 6+7 → detached learned",
|
||||
};
|
||||
const metricLabels: Record<string, string> = {
|
||||
spike_contrast: "SPIKE CONTRAST",
|
||||
peak_normalized: "PEAK / MEAN",
|
||||
};
|
||||
const tabs = [...root.querySelectorAll<HTMLButtonElement>("[data-local-tab]")];
|
||||
const panels = [...root.querySelectorAll<HTMLElement>("[data-local-panel]")];
|
||||
const activate = (name: string, focus = false) => {
|
||||
tabs.forEach((tab) => {
|
||||
const selected = tab.dataset.localTab === name;
|
||||
tab.setAttribute("aria-selected", String(selected));
|
||||
if (selected && focus) tab.focus();
|
||||
});
|
||||
panels.forEach((panel) => {
|
||||
panel.hidden = panel.dataset.localPanel !== name;
|
||||
});
|
||||
};
|
||||
tabs.forEach((tab, index) => {
|
||||
tab.addEventListener("click", () => activate(tab.dataset.localTab || "scope"));
|
||||
tab.addEventListener("keydown", (event) => {
|
||||
if (!["ArrowLeft", "ArrowRight", "Home", "End"].includes(event.key)) return;
|
||||
event.preventDefault();
|
||||
let next = index;
|
||||
if (event.key === "ArrowRight") next = (index + 1) % tabs.length;
|
||||
if (event.key === "ArrowLeft") next = (index - 1 + tabs.length) % tabs.length;
|
||||
if (event.key === "Home") next = 0;
|
||||
if (event.key === "End") next = tabs.length - 1;
|
||||
activate(tabs[next].dataset.localTab || "scope", true);
|
||||
});
|
||||
});
|
||||
|
||||
let matrixMetric = "spike_contrast";
|
||||
const matrixSeed = root.querySelector<HTMLSelectElement>("[data-local-matrix-seed]");
|
||||
const matrixBars = root.querySelector<HTMLElement>("[data-local-matrix-bars]");
|
||||
const matrixState = root.querySelector<HTMLElement>("[data-local-matrix-state]");
|
||||
const family = (mode: string) => mode.startsWith("uniform_except_") ? "restoration" : "sufficiency";
|
||||
const matrixValue = (mode: string, seed: string, metric: string) => {
|
||||
if (mode === "detached_learned") return 0;
|
||||
const kind = family(mode);
|
||||
if (seed === "mean") return data.means[kind][mode][metric];
|
||||
return data.scores[kind][mode][seed][metric].score;
|
||||
};
|
||||
const rawValue = (mode: string, seed: string, metric: string) => {
|
||||
if (seed === "mean") return data.means.mode_metrics[mode][metric];
|
||||
return data.formal_cells.find((cell: any) => String(cell.seed) === seed).mode_metrics[mode][metric];
|
||||
};
|
||||
const uniformCount = (mode: string) => data.final_spectra[0].modes[mode].uniform_count;
|
||||
const renderMatrix = () => {
|
||||
if (!matrixBars || !matrixSeed || !matrixState) return;
|
||||
const selectedSeed = matrixSeed.value;
|
||||
matrixState.textContent = `${selectedSeed === "mean" ? "3-SEED MEAN" : selectedSeed} · ${metricLabels[matrixMetric]}`;
|
||||
matrixBars.innerHTML = "";
|
||||
data.study.modes.forEach((mode: string) => {
|
||||
const value = matrixValue(mode, selectedSeed, matrixMetric);
|
||||
const row = document.createElement("article");
|
||||
const kind = family(mode);
|
||||
row.dataset.family = mode === "detached_learned" ? "reference" : kind;
|
||||
const magnitude = Math.min(Math.abs(value) / 2, 1) * 100;
|
||||
row.innerHTML = `
|
||||
<div><span>${mode === "detached_learned" ? "REF" : kind === "sufficiency" ? "S" : "R"}</span><b>${modeLabels[mode]}</b><small>${uniformCount(mode)} uniform</small></div>
|
||||
<div class="matrix-track"><i class="${value < 0 ? "negative" : ""}" style="--magnitude:${magnitude}%"></i></div>
|
||||
<strong>${value >= 0 ? "+" : "−"}${Math.abs(value).toFixed(3)}</strong>
|
||||
<em>RAW ${rawValue(mode, selectedSeed, matrixMetric).toFixed(3)}×</em>`;
|
||||
matrixBars.appendChild(row);
|
||||
});
|
||||
};
|
||||
matrixSeed?.addEventListener("change", renderMatrix);
|
||||
root.querySelectorAll<HTMLButtonElement>("[data-local-matrix-metric]").forEach((button) => {
|
||||
button.addEventListener("click", () => {
|
||||
matrixMetric = button.dataset.localMatrixMetric || "spike_contrast";
|
||||
root.querySelectorAll<HTMLButtonElement>("[data-local-matrix-metric]").forEach((item) =>
|
||||
item.setAttribute("aria-pressed", String(item === button)));
|
||||
renderMatrix();
|
||||
});
|
||||
});
|
||||
renderMatrix();
|
||||
|
||||
const renderBranch = (group: string) => {
|
||||
root.querySelectorAll<HTMLButtonElement>("[data-local-branch-group]").forEach((button) =>
|
||||
button.setAttribute("aria-pressed", String(button.dataset.localBranchGroup === group)));
|
||||
const gate = data.gates.branch_dominance[`group_${group}`];
|
||||
const state = root.querySelector<HTMLElement>("[data-local-branch-state]");
|
||||
if (state) state.textContent = group === "7"
|
||||
? "GROUP 7 · MLP DOMINANT · 6 / 6 PASS"
|
||||
: "GROUP 6 · NO DOMINANT BRANCH · MLP 5 / 6";
|
||||
(["attention", "mlp"] as const).forEach((branch) => {
|
||||
root.querySelectorAll<HTMLElement>(`[data-local-branch-label="${branch}"]`).forEach((node) => node.textContent = `GROUP ${group}`);
|
||||
(["spike_contrast", "peak_normalized"] as const).forEach((metric) => {
|
||||
const mode = `uniform_group_${group}_${branch}_only`;
|
||||
const value = data.means.sufficiency[mode][metric];
|
||||
const bar = root.querySelector<HTMLElement>(`[data-local-branch-bar="${branch}:${metric}"]`);
|
||||
const label = root.querySelector<HTMLElement>(`[data-local-branch-value="${branch}:${metric}"]`);
|
||||
if (bar) bar.style.setProperty("--branch-score", String(Math.min(value, 1.2) / 1.2));
|
||||
if (label) label.textContent = value.toFixed(3);
|
||||
});
|
||||
});
|
||||
root.dataset.branchPassed = String(gate.passed);
|
||||
};
|
||||
root.querySelectorAll<HTMLButtonElement>("[data-local-branch-group]").forEach((button) =>
|
||||
button.addEventListener("click", () => renderBranch(button.dataset.localBranchGroup || "6")));
|
||||
renderBranch("6");
|
||||
|
||||
const spectrumSeed = root.querySelector<HTMLSelectElement>("[data-local-spectrum-seed]");
|
||||
const spectrumMode = root.querySelector<HTMLSelectElement>("[data-local-spectrum-mode]");
|
||||
const spectrumState = root.querySelector<HTMLElement>("[data-local-spectrum-state]");
|
||||
const chart = root.querySelector<SVGElement>("[data-local-spectrum-chart]");
|
||||
const grid = root.querySelector<SVGGElement>("[data-local-spectrum-grid]");
|
||||
const line = root.querySelector<SVGPolylineElement>("[data-local-spectrum-line]");
|
||||
const points = root.querySelector<SVGGElement>("[data-local-spectrum-points]");
|
||||
const zone = root.querySelector<SVGRectElement>("[data-local-spectrum-zone]");
|
||||
const svg = (name: string) => document.createElementNS("http://www.w3.org/2000/svg", name);
|
||||
const renderSpectrum = () => {
|
||||
if (!spectrumSeed || !spectrumMode || !chart || !grid || !line || !points || !zone) return;
|
||||
const seed = spectrumSeed.value;
|
||||
const mode = spectrumMode.value;
|
||||
const run = data.final_spectra.find((item: any) => String(item.seed) === seed);
|
||||
const selected = run.modes[mode];
|
||||
const values = selected.normalized;
|
||||
const left = 56, right = 908, top = 28, bottom = 306;
|
||||
const maximum = Math.max(2, ...values) * 1.08;
|
||||
const x = (index: number) => left + index * (right - left) / 31;
|
||||
const y = (value: number) => bottom - value / maximum * (bottom - top);
|
||||
grid.innerHTML = "";
|
||||
[0, .5, 1, 1.5, 2].filter((value) => value <= maximum).forEach((value) => {
|
||||
const guide = svg("line");
|
||||
guide.setAttribute("x1", String(left));
|
||||
guide.setAttribute("x2", String(right));
|
||||
guide.setAttribute("y1", String(y(value)));
|
||||
guide.setAttribute("y2", String(y(value)));
|
||||
grid.appendChild(guide);
|
||||
const label = svg("text");
|
||||
label.setAttribute("x", "48");
|
||||
label.setAttribute("y", String(y(value) + 4));
|
||||
label.setAttribute("text-anchor", "end");
|
||||
label.textContent = value.toFixed(1);
|
||||
grid.appendChild(label);
|
||||
});
|
||||
zone.setAttribute("x", String(x(20) - 8));
|
||||
zone.setAttribute("width", String(x(24) - x(20) + 16));
|
||||
line.setAttribute("points", values.map((value: number, index: number) => `${x(index)},${y(value)}`).join(" "));
|
||||
points.innerHTML = "";
|
||||
values.forEach((value: number, index: number) => {
|
||||
const point = svg("circle");
|
||||
point.setAttribute("cx", String(x(index)));
|
||||
point.setAttribute("cy", String(y(value)));
|
||||
point.setAttribute("r", index + 1 === selected.peak_layer ? "4.5" : "2.3");
|
||||
point.dataset.layer = String(index + 1);
|
||||
points.appendChild(point);
|
||||
});
|
||||
if (spectrumState) spectrumState.textContent = `${seed} · ${modeLabels[mode]}`;
|
||||
const bind = (selector: string, value: string) => {
|
||||
const node = root.querySelector<HTMLElement>(selector);
|
||||
if (node) node.textContent = value;
|
||||
};
|
||||
bind("[data-local-spectrum-label]", modeLabels[mode]);
|
||||
bind("[data-local-spectrum-count]", `${selected.uniform_count} / 65`);
|
||||
bind("[data-local-spectrum-contrast]", `${selected.spike_contrast.toFixed(3)}×`);
|
||||
bind("[data-local-spectrum-peak]", `${selected.peak_normalized.toFixed(3)}×`);
|
||||
bind("[data-local-spectrum-layer]", String(selected.peak_layer));
|
||||
};
|
||||
spectrumSeed?.addEventListener("change", renderSpectrum);
|
||||
spectrumMode?.addEventListener("change", renderSpectrum);
|
||||
renderSpectrum();
|
||||
};
|
||||
|
||||
document.querySelectorAll<HTMLElement>("[data-local-path-lab]").forEach(initializeLocalPathLab);
|
||||
</script>
|
||||
|
||||
<style>
|
||||
.local-lab {
|
||||
--ink: #e8e5dd;
|
||||
--muted: #9a9890;
|
||||
--line: rgba(232,229,221,.14);
|
||||
--panel: #121413;
|
||||
--panel-2: #171a18;
|
||||
--green: #8aae8d;
|
||||
--copper: #c98a58;
|
||||
--red: #c77768;
|
||||
margin: 42px 0 0;
|
||||
border: 1px solid var(--line);
|
||||
background: #0d0f0e;
|
||||
color: var(--ink);
|
||||
overflow: hidden;
|
||||
}
|
||||
.local-lab figcaption { display: grid; grid-template-columns: auto 1fr auto; gap: 24px; align-items: start; padding: 26px 28px; border-bottom: 1px solid var(--line); }
|
||||
.local-lab figcaption > span, .panel-lead span { color: var(--copper); font: .58rem/1.3 var(--mono); letter-spacing: .1em; }
|
||||
.local-lab figcaption h3 { margin: 0; color: var(--ink); font-size: 1.08rem; line-height: 1.35; }
|
||||
.local-lab figcaption p { margin: 8px 0 0; color: var(--muted); font: .62rem/1.5 var(--mono); }
|
||||
.local-lab figcaption em { color: var(--muted); font: normal .55rem var(--mono); letter-spacing: .08em; }
|
||||
.local-ledger { display: grid; grid-template-columns: repeat(6, 1fr); border-bottom: 1px solid var(--line); }
|
||||
.local-ledger article { min-width: 0; padding: 16px 14px; border-right: 1px solid var(--line); }
|
||||
.local-ledger article:last-child { border-right: 0; }
|
||||
.local-ledger span { color: var(--muted); font: .52rem var(--mono); letter-spacing: .08em; }
|
||||
.local-ledger b { display: block; margin-top: 10px; font: 800 .78rem var(--mono); }
|
||||
.local-ledger p { margin: 7px 0 0; color: var(--muted); font: .55rem/1.4 var(--mono); }
|
||||
.local-ledger .pass b { color: var(--green); }
|
||||
.local-ledger .warn b { color: var(--copper); }
|
||||
.local-ledger .boundary { background: rgba(199,119,104,.07); }
|
||||
.local-ledger .boundary b { color: var(--red); font-size: .68rem; }
|
||||
.local-tabs { display: grid; grid-template-columns: repeat(5, 1fr); border-bottom: 1px solid var(--line); }
|
||||
.local-tabs button, .local-controls button { appearance: none; border: 0; border-right: 1px solid var(--line); background: transparent; color: var(--muted); cursor: pointer; font: .57rem var(--mono); letter-spacing: .06em; }
|
||||
.local-tabs button { padding: 16px 10px; }
|
||||
.local-tabs button[aria-selected="true"], .local-controls button[aria-pressed="true"] { background: rgba(201,138,88,.11); color: var(--copper); box-shadow: inset 0 -2px var(--copper); }
|
||||
.local-panel { padding: 28px; }
|
||||
.panel-lead { display: grid; grid-template-columns: 1fr minmax(260px, .7fr); gap: 34px; align-items: end; margin-bottom: 26px; }
|
||||
.panel-lead h4 { margin: 9px 0 0; color: var(--ink); font-size: 1rem; line-height: 1.4; }
|
||||
.panel-lead > p { margin: 0; color: var(--muted); font-size: .68rem; line-height: 1.7; }
|
||||
.group-map { display: grid; grid-template-columns: repeat(8, minmax(72px,1fr)) 22px minmax(82px,.8fr); align-items: stretch; border: 1px solid var(--line); }
|
||||
.group-map article { position: relative; min-width: 0; padding: 14px 10px; border-right: 1px solid var(--line); background: var(--panel); }
|
||||
.group-map article.target { background: rgba(201,138,88,.11); box-shadow: inset 0 3px var(--copper); }
|
||||
.group-map span { color: var(--muted); font: .5rem var(--mono); }
|
||||
.group-map b { display: block; margin: 12px 0; font: .7rem var(--mono); }
|
||||
.group-map article > div { display: grid; grid-template-columns: repeat(2, 1fr); gap: 4px; }
|
||||
.group-map article i { display: grid; place-items: center; aspect-ratio: 1; border: 1px solid var(--line); color: var(--muted); font: normal .55rem var(--mono); }
|
||||
.group-map article i.spike { border-color: var(--copper); background: var(--copper); color: #16110d; }
|
||||
.group-map small { display: block; margin-top: 11px; color: var(--muted); font: .45rem var(--mono); }
|
||||
.map-arrow { display: grid; place-items: center; color: var(--muted); font-style: normal; }
|
||||
.group-map .output-node { border-left: 1px solid var(--line); border-right: 0; background: rgba(138,174,141,.08); }
|
||||
.group-map .output-node b { color: var(--green); font-size: 1rem; }
|
||||
.group-map .output-node p { color: var(--muted); font: .5rem var(--mono); }
|
||||
.scope-key { display: flex; flex-wrap: wrap; gap: 18px; margin: 14px 0 24px; color: var(--muted); font: .54rem var(--mono); }
|
||||
.scope-key span { display: inline-flex; align-items: center; gap: 7px; }
|
||||
.scope-key i { width: 12px; height: 12px; border: 1px solid var(--line); }
|
||||
.scope-key i.target { background: rgba(201,138,88,.2); border-color: var(--copper); }
|
||||
.scope-key i.spike { background: var(--copper); }
|
||||
.scope-key i.plain { background: var(--panel); }
|
||||
.direction-pair { display: grid; grid-template-columns: 1fr auto 1fr; gap: 18px; align-items: center; }
|
||||
.direction-pair > article { padding: 20px; border: 1px solid var(--line); background: var(--panel); }
|
||||
.direction-pair > article > span { color: var(--copper); font: .55rem var(--mono); }
|
||||
.direction-pair article div { display: flex; gap: 10px; align-items: center; margin: 18px 0; }
|
||||
.direction-pair article div b { padding: 9px 10px; border: 1px solid var(--line); font: .6rem var(--mono); }
|
||||
.direction-pair article div b.accent { border-color: var(--copper); color: var(--copper); }
|
||||
.direction-pair article div i, .direction-pair > i { color: var(--muted); font-style: normal; }
|
||||
.direction-pair code { color: var(--green); font-size: .65rem; }
|
||||
.direction-pair p { margin: 12px 0 0; color: var(--muted); font-size: .63rem; }
|
||||
.plain-rule { margin-top: 20px; padding: 17px 20px; border-left: 3px solid var(--copper); background: rgba(201,138,88,.07); }
|
||||
.plain-rule b { font-size: .68rem; }
|
||||
.plain-rule p { margin: 6px 0 0; color: var(--muted); font-size: .65rem; line-height: 1.6; }
|
||||
.local-controls { display: flex; gap: 12px; align-items: stretch; min-height: 40px; margin-bottom: 18px; border: 1px solid var(--line); }
|
||||
.local-controls label { display: flex; align-items: center; gap: 10px; padding: 0 12px; border-right: 1px solid var(--line); color: var(--muted); font: .54rem var(--mono); }
|
||||
.local-controls select { max-width: 310px; border: 0; background: transparent; color: var(--ink); font: .58rem var(--mono); }
|
||||
.local-controls select option { background: #151715; }
|
||||
.local-controls > div { display: flex; }
|
||||
.local-controls button { padding: 0 16px; border-left: 1px solid var(--line); }
|
||||
.local-controls > span { margin-left: auto; display: flex; align-items: center; padding: 0 14px; color: var(--muted); font: .55rem var(--mono); }
|
||||
.matrix-axis { display: grid; grid-template-columns: auto 1fr auto 1fr auto 1fr auto 1fr auto; gap: 8px; align-items: center; padding: 0 170px 8px 245px; color: var(--muted); font: .47rem var(--mono); }
|
||||
.matrix-axis i { height: 1px; background: var(--line); }
|
||||
.matrix-bars { border: 1px solid var(--line); }
|
||||
.matrix-bars :global(article) { display: grid; grid-template-columns: 235px 1fr 66px 86px; gap: 12px; align-items: center; min-height: 44px; padding: 5px 12px; border-bottom: 1px solid var(--line); }
|
||||
.matrix-bars :global(article:last-child) { border-bottom: 0; }
|
||||
.matrix-bars :global(article[data-family="restoration"]) { background: rgba(138,174,141,.035); }
|
||||
.matrix-bars :global(article > div:first-child) { display: grid; grid-template-columns: 24px 1fr auto; gap: 8px; align-items: center; min-width: 0; }
|
||||
.matrix-bars :global(article > div:first-child span) { display: grid; place-items: center; width: 22px; height: 22px; border: 1px solid var(--line); color: var(--copper); font: .52rem var(--mono); }
|
||||
.matrix-bars :global(article[data-family="restoration"] > div:first-child span) { color: var(--green); }
|
||||
.matrix-bars :global(article > div:first-child b) { overflow: hidden; color: var(--ink); text-overflow: ellipsis; white-space: nowrap; font: .58rem var(--mono); }
|
||||
.matrix-bars :global(article > div:first-child small) { color: var(--muted); font: .45rem var(--mono); }
|
||||
.matrix-bars :global(.matrix-track) { position: relative; height: 12px; background: rgba(232,229,221,.045); }
|
||||
.matrix-bars :global(.matrix-track::after) { content: ""; position: absolute; left: 13%; top: -4px; bottom: -4px; width: 1px; background: rgba(232,229,221,.4); }
|
||||
.matrix-bars :global(.matrix-track i) { position: absolute; left: 13%; top: 2px; height: 8px; width: calc(var(--magnitude) * .87); background: var(--copper); }
|
||||
.matrix-bars :global(article[data-family="restoration"] .matrix-track i) { background: var(--green); }
|
||||
.matrix-bars :global(.matrix-track i.negative) { right: 87%; left: auto; width: calc(var(--magnitude) * .13); background: var(--red) !important; }
|
||||
.matrix-bars :global(strong) { text-align: right; color: var(--ink); font: .64rem var(--mono); }
|
||||
.matrix-bars :global(em) { color: var(--muted); font: normal .48rem var(--mono); }
|
||||
.matrix-note { display: grid; grid-template-columns: repeat(3, 1fr); margin-top: 14px; border: 1px solid var(--line); }
|
||||
.matrix-note article { display: grid; grid-template-columns: 30px 1fr; gap: 10px; align-items: center; padding: 12px; border-right: 1px solid var(--line); }
|
||||
.matrix-note article:last-child { border-right: 0; }
|
||||
.matrix-note b { color: var(--copper); font: .75rem var(--mono); }
|
||||
.matrix-note p { margin: 0; color: var(--muted); font-size: .56rem; }
|
||||
.verdict-banner { display: grid; grid-template-columns: auto 1fr auto; align-items: center; gap: 18px; padding: 18px 20px; border: 1px solid rgba(199,119,104,.5); background: rgba(199,119,104,.08); }
|
||||
.verdict-banner span, .verdict-banner p { color: var(--muted); font: .54rem var(--mono); }
|
||||
.verdict-banner b { color: var(--red); font: 800 .9rem var(--mono); }
|
||||
.verdict-banner p { margin: 0; text-align: right; }
|
||||
.dual-table-wrap { margin-top: 18px; overflow-x: auto; border: 1px solid var(--line); }
|
||||
.dual-table { width: 100%; border-collapse: collapse; font: .59rem var(--mono); }
|
||||
.dual-table th, .dual-table td { padding: 13px 11px; border-right: 1px solid var(--line); border-bottom: 1px solid var(--line); text-align: right; }
|
||||
.dual-table th:first-child { text-align: left; }
|
||||
.dual-table thead { color: var(--muted); }
|
||||
.dual-table .good { color: var(--green); }
|
||||
.dual-table .bad { color: var(--red); }
|
||||
.dual-table .mean { background: rgba(232,229,221,.04); }
|
||||
.dual-visual { display: grid; grid-template-columns: 1fr 1fr; gap: 16px; margin-top: 18px; }
|
||||
.dual-visual article { padding: 18px; border: 1px solid var(--line); background: var(--panel); }
|
||||
.dual-visual span { color: var(--muted); font: .52rem var(--mono); }
|
||||
.dual-visual b { display: block; margin: 9px 0 16px; font: .75rem var(--mono); }
|
||||
.dual-visual article > div { position: relative; height: 22px; margin: 8px 0; background: rgba(232,229,221,.05); }
|
||||
.dual-visual article > div::after { content: "50%"; position: absolute; left: 50%; top: -14px; bottom: -2px; border-left: 1px dashed var(--muted); color: var(--muted); font: .42rem var(--mono); }
|
||||
.dual-visual article i { display: block; width: calc(min(var(--score), 1) * 100%); height: 100%; background: var(--green); }
|
||||
.dual-visual article.fail i { background: var(--copper); }
|
||||
.dual-visual em { position: absolute; left: 9px; top: 5px; color: #0d0f0e; font: normal 800 .52rem var(--mono); }
|
||||
.dual-visual p { margin: 13px 0 0; color: var(--muted); font-size: .62rem; line-height: 1.55; }
|
||||
.boundary-pair { display: grid; grid-template-columns: 1fr 1fr; gap: 16px; margin-top: 18px; }
|
||||
.boundary-pair article { padding: 18px; border: 1px solid var(--line); }
|
||||
.boundary-pair span { color: var(--muted); font: .52rem var(--mono); }
|
||||
.boundary-pair b { display: block; margin: 9px 0; font-size: .72rem; }
|
||||
.boundary-pair p { margin: 0; color: var(--muted); font-size: .61rem; line-height: 1.5; }
|
||||
.boundary-pair .yes { border-color: rgba(138,174,141,.45); }
|
||||
.boundary-pair .yes b { color: var(--green); }
|
||||
.boundary-pair .no { border-color: rgba(199,119,104,.45); }
|
||||
.boundary-pair .no b { color: var(--red); }
|
||||
.branch-chart { display: grid; grid-template-columns: 1fr 1fr; gap: 16px; }
|
||||
.branch-chart article { padding: 18px; border: 1px solid var(--line); background: var(--panel); }
|
||||
.branch-chart article > span { color: var(--muted); font: .52rem var(--mono); }
|
||||
.branch-chart article > b { display: block; margin: 8px 0 18px; font: .75rem var(--mono); }
|
||||
.branch-chart article > div { display: grid; grid-template-columns: 74px 1fr 46px; gap: 9px; align-items: center; margin: 10px 0; }
|
||||
.branch-chart em { color: var(--muted); font: normal .5rem var(--mono); }
|
||||
.branch-chart i { height: 14px; width: calc(var(--branch-score, .1) * 100%); background: var(--copper); }
|
||||
.branch-chart .mlp i { background: var(--green); }
|
||||
.branch-chart strong { text-align: right; font: .58rem var(--mono); }
|
||||
.branch-verdicts, .control-grid, .audit-grid { display: grid; gap: 14px; margin-top: 16px; }
|
||||
.branch-verdicts { grid-template-columns: 1fr 1fr; }
|
||||
.control-grid { grid-template-columns: repeat(3, 1fr); }
|
||||
.branch-verdicts article, .control-grid article, .audit-grid article { padding: 16px; border: 1px solid var(--line); }
|
||||
.branch-verdicts span, .control-grid span, .audit-grid span { color: var(--muted); font: .5rem var(--mono); }
|
||||
.branch-verdicts b, .control-grid b, .audit-grid b { display: block; margin: 9px 0; font: .7rem var(--mono); }
|
||||
.branch-verdicts p, .control-grid p, .audit-grid p { margin: 0; color: var(--muted); font-size: .59rem; line-height: 1.5; }
|
||||
.branch-verdicts .pass { border-color: rgba(138,174,141,.45); }
|
||||
.branch-verdicts .pass b { color: var(--green); }
|
||||
.control-grid .warning { border-color: rgba(201,138,88,.4); }
|
||||
.spectrum-controls label:nth-child(2) { flex: 1; }
|
||||
.spectrum-controls label:nth-child(2) select { width: 100%; max-width: none; }
|
||||
.spectrum-layout { display: grid; grid-template-columns: 1fr 220px; gap: 16px; }
|
||||
.chart-shell { border: 1px solid var(--line); background: var(--panel); }
|
||||
.chart-shell header { display: flex; justify-content: space-between; padding: 12px 15px; border-bottom: 1px solid var(--line); }
|
||||
.chart-shell header b, .chart-shell header span { font: .53rem var(--mono); }
|
||||
.chart-shell header span { color: var(--muted); }
|
||||
.chart-shell svg { display: block; width: 100%; height: auto; }
|
||||
[data-local-spectrum-grid] :global(line) { stroke: var(--line); stroke-width: 1; }
|
||||
[data-local-spectrum-grid] :global(text) { fill: var(--muted); font: 10px var(--mono); }
|
||||
.spike-zone { fill: rgba(201,138,88,.08); }
|
||||
.series { fill: none; stroke: var(--green); stroke-width: 2; }
|
||||
[data-local-spectrum-points] :global(circle) { fill: var(--green); stroke: #0d0f0e; stroke-width: 1; }
|
||||
.spectrum-key { display: flex; gap: 18px; padding: 10px 14px; border-top: 1px solid var(--line); color: var(--muted); font: .5rem var(--mono); }
|
||||
.spectrum-key span { display: flex; gap: 6px; align-items: center; }
|
||||
.spectrum-key i { width: 14px; height: 2px; background: var(--green); }
|
||||
.spectrum-key i.zone { height: 10px; background: rgba(201,138,88,.25); }
|
||||
.spectrum-layout aside { padding: 18px; border: 1px solid var(--line); background: var(--panel); }
|
||||
.spectrum-layout aside > span { color: var(--muted); font: .5rem var(--mono); }
|
||||
.spectrum-layout aside > b { display: block; margin: 10px 0 18px; font: .67rem/1.4 var(--mono); }
|
||||
.spectrum-layout dl { margin: 0; }
|
||||
.spectrum-layout dl div { display: flex; justify-content: space-between; gap: 8px; padding: 10px 0; border-top: 1px solid var(--line); }
|
||||
.spectrum-layout dt, .spectrum-layout dd { margin: 0; font: .53rem var(--mono); }
|
||||
.spectrum-layout dt { color: var(--muted); }
|
||||
.spectrum-layout dd { color: var(--green); }
|
||||
.spectrum-story { display: grid; grid-template-columns: 1fr auto 1fr auto 1fr; gap: 12px; align-items: center; margin-top: 16px; }
|
||||
.spectrum-story article { min-height: 105px; padding: 15px; border: 1px solid var(--line); }
|
||||
.spectrum-story span { color: var(--muted); font: .5rem var(--mono); }
|
||||
.spectrum-story b { display: block; margin: 9px 0; font: .7rem var(--mono); }
|
||||
.spectrum-story p { margin: 0; color: var(--muted); font-size: .58rem; line-height: 1.5; }
|
||||
.spectrum-story > i { color: var(--muted); font-style: normal; }
|
||||
.audit-grid { grid-template-columns: repeat(4, 1fr); }
|
||||
.audit-grid .boundary { border-color: rgba(199,119,104,.45); }
|
||||
.hash-strip { display: flex; flex-wrap: wrap; gap: 18px; margin-top: 14px; padding: 12px 14px; border: 1px solid var(--line); color: var(--muted); font: .48rem var(--mono); }
|
||||
.hash-strip code { color: var(--green); }
|
||||
|
||||
@media (max-width: 980px) {
|
||||
.local-ledger { grid-template-columns: repeat(3, 1fr); }
|
||||
.local-ledger article:nth-child(3) { border-right: 0; }
|
||||
.local-tabs { grid-template-columns: repeat(3, 1fr); }
|
||||
.group-map { grid-template-columns: repeat(4, 1fr); }
|
||||
.group-map .map-arrow { display: none; }
|
||||
.group-map .output-node { grid-column: span 4; border-left: 0; border-top: 1px solid var(--line); }
|
||||
.matrix-bars :global(article) { grid-template-columns: 210px 1fr 58px; }
|
||||
.matrix-bars :global(em) { display: none; }
|
||||
.matrix-axis { display: none; }
|
||||
.spectrum-layout { grid-template-columns: 1fr; }
|
||||
.audit-grid { grid-template-columns: repeat(2, 1fr); }
|
||||
}
|
||||
@media (max-width: 680px) {
|
||||
.local-lab figcaption { grid-template-columns: 1fr; padding: 20px; }
|
||||
.local-lab figcaption > span { order: -1; }
|
||||
.local-ledger { grid-template-columns: repeat(2, 1fr); }
|
||||
.local-ledger article:nth-child(3) { border-right: 1px solid var(--line); }
|
||||
.local-ledger article:nth-child(even) { border-right: 0; }
|
||||
.local-tabs { display: flex; overflow-x: auto; }
|
||||
.local-tabs button { min-width: 145px; }
|
||||
.local-panel { padding: 20px 14px; }
|
||||
.panel-lead { grid-template-columns: 1fr; gap: 14px; }
|
||||
.group-map { grid-template-columns: repeat(2, 1fr); }
|
||||
.group-map .output-node { grid-column: span 2; }
|
||||
.direction-pair { grid-template-columns: 1fr; }
|
||||
.direction-pair > i { transform: rotate(90deg); text-align: center; }
|
||||
.local-controls { flex-wrap: wrap; }
|
||||
.local-controls > span { width: 100%; min-height: 34px; border-top: 1px solid var(--line); }
|
||||
.matrix-bars :global(article) { grid-template-columns: 1fr 58px; padding: 9px; }
|
||||
.matrix-bars :global(article > div:first-child) { grid-column: 1 / -1; }
|
||||
.matrix-bars :global(.matrix-track) { min-width: 0; }
|
||||
.matrix-note, .dual-visual, .boundary-pair, .branch-chart, .branch-verdicts, .control-grid, .audit-grid { grid-template-columns: 1fr; }
|
||||
.verdict-banner { grid-template-columns: 1fr; }
|
||||
.verdict-banner p { text-align: left; }
|
||||
.dual-table { min-width: 670px; }
|
||||
.spectrum-story { grid-template-columns: 1fr; }
|
||||
.spectrum-story > i { transform: rotate(90deg); text-align: center; }
|
||||
}
|
||||
</style>
|
||||
@@ -334,7 +334,7 @@ const reductionLabels: Record<string, string> = {
|
||||
<i>→</i>
|
||||
<article class="accent"><span>3 / INTERVENE</span><b>value route −70.2%</b><p>全局反向规则干预支持路径敏感性。</p></article>
|
||||
<i>→</i>
|
||||
<article><span>4 / NEXT GATE</span><b>局部 group 6 / 7</b><p>下一轮冻结局部 mixer 的干预矩阵。</p></article>
|
||||
<article><span>4 / LOCAL FOLLOW-UP</span><b>单侧 6 / 6 · 双向 3 / 6</b><p>groups 6+7 的 sufficiency 通过;restoration peak 未过,localization 未建立。</p></article>
|
||||
</div>
|
||||
<div class="audit-grid">
|
||||
<article><span>ROUND 05 EQUIVALENCE</span><b>3 / 3 exact</b><p>model、optimizer、history、六个 BPC 与全部 post-MLP 数组一致。</p></article>
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -128,20 +128,20 @@ const paths = [
|
||||
<div class="release-grid">
|
||||
<a class="release-card k3-release" href="/k3/">
|
||||
<div>
|
||||
<p class="eyebrow"><span>NEW / K3 ROUND 06</span> SPIKE TRAJECTORY · SAME-FORWARD BACKWARD PATH</p>
|
||||
<h2>尖峰不是出生时就有:它在训练中形成,并对 value 路径敏感</h2>
|
||||
<p class="eyebrow"><span>NEW / K3 ROUND 07</span> LOCAL MIXER PATH · BIDIRECTIONAL GATE</p>
|
||||
<h2>16 个局部 mixer 单侧证据很强,但双向定位仍然没有闭合</h2>
|
||||
<p>
|
||||
严格复用 Round 05 的 depth-32 Block 正式格,追踪六个训练时点、六个张量位置与四种
|
||||
reduction。layer 21–25 的尖峰在 step 500 后形成;全局 uniform value-backward
|
||||
让 contrast 平均下降 70.2%,但只判为反向路径敏感性,不冒充因果贡献或训练变体。
|
||||
以 14 个冻结 mask 同时检查 sufficiency 与 restoration:groups 6+7 在 learned
|
||||
背景的 6 / 6 格全部超过 50% global log gap;但从 uniform 背景恢复时,peak 三个
|
||||
seed 全部未过线。因此只能报告 one-sided evidence,不能宣布尖峰已定位到这 16 个 mixer。
|
||||
</p>
|
||||
</div>
|
||||
<dl>
|
||||
<div><dt>FORMAL</dt><dd>3 seeds × 8,000 steps</dd></div>
|
||||
<div><dt>ROBUST</dt><dd>4 reductions · 12/12</dd></div>
|
||||
<div><dt>REPLAY</dt><dd>16 frozen groups exact</dd></div>
|
||||
<div><dt>MATRIX</dt><dd>14 masks × 3 seeds</dd></div>
|
||||
<div><dt>SUFFICIENCY</dt><dd>6 / 6 pass</dd></div>
|
||||
<div><dt>RESTORATION</dt><dd>3 / 6 fail</dd></div>
|
||||
</dl>
|
||||
<span class="release-arrow" aria-hidden="true">进入训练轨迹、六位置谱与反向路径干预 →</span>
|
||||
<span class="release-arrow" aria-hidden="true">进入局部路径图、双向门与 32 层原始谱 →</span>
|
||||
</a>
|
||||
<a class="release-card deepseek-release" href="/deepseek/">
|
||||
<div>
|
||||
|
||||
@@ -3,6 +3,7 @@ import BaseLayout from "@/layouts/BaseLayout.astro";
|
||||
import ArchitectureExplorer from "@/components/ArchitectureExplorer.astro";
|
||||
import K3ArtifactLab from "@/components/K3ArtifactLab.astro";
|
||||
import K3AttnResGradientLab from "@/components/K3AttnResGradientLab.astro";
|
||||
import K3AttnResLocalPathLab from "@/components/K3AttnResLocalPathLab.astro";
|
||||
import K3AttnResSpikeLab from "@/components/K3AttnResSpikeLab.astro";
|
||||
import K3AttnResTraceLab from "@/components/K3AttnResTraceLab.astro";
|
||||
import K3ReportLab from "@/components/K3ReportLab.astro";
|
||||
@@ -42,7 +43,8 @@ const toc = [
|
||||
["30", "attnres-reduced", "AttnRes 缩小机制实验"],
|
||||
["31", "attnres-gradient", "梯度定义与深度扩展"],
|
||||
["32", "attnres-spike", "尖峰轨迹与反向路径"],
|
||||
["33", "audit", "21 张图表审计"],
|
||||
["33", "attnres-local-path", "局部 mixer 双向干预"],
|
||||
["34", "audit", "21 张图表审计"],
|
||||
["↳", "papers", "100 节点阅读链"],
|
||||
];
|
||||
|
||||
@@ -111,13 +113,13 @@ const paperGroups = [
|
||||
|
||||
<BaseLayout
|
||||
title="Kimi K3 技术报告完整深读:架构、训练、RL、系统与评测"
|
||||
description="用三十二张问题账、二十一张图表审计、八个机制实验、四个开放工件视图、三轮十五个 AttnRes 独立实验视图与一百个一手阅读节点,逐节读懂 Kimi K3。"
|
||||
description="用三十二张问题账、二十一张图表审计、八个机制实验、四个开放工件视图、四轮二十个 AttnRes 独立实验视图与一百个一手阅读节点,逐节读懂 Kimi K3。"
|
||||
section="k3"
|
||||
>
|
||||
<header class="page-hero k3-hero">
|
||||
<div class="page-hero-inner">
|
||||
<div>
|
||||
<p class="eyebrow"><span>ANCHOR REPORT / ROUND 06</span> KIMI K3 · REPORT → ARTIFACTS → INDEPENDENT PROBE</p>
|
||||
<p class="eyebrow"><span>ANCHOR REPORT / ROUND 07</span> KIMI K3 · REPORT → ARTIFACTS → INDEPENDENT PROBE</p>
|
||||
<h1>不把报告压成摘要<br />把每个因果环节<br />重新展开</h1>
|
||||
<p class="lead">
|
||||
K3 同时扩展序列、深度、宽度、视觉与 Agent 轨迹。真正值得读的不是 2.8T 这个最大数字,
|
||||
@@ -131,7 +133,7 @@ const paperGroups = [
|
||||
<div><dt>READING</dt><dd>100 个一手 / 官方节点</dd></div>
|
||||
<div><dt>MODEL</dt><dd>2.78T total / 104.2B active</dd></div>
|
||||
<div><dt>ARTIFACTS</dt><dd>96 shards · 497,220 tensors</dd></div>
|
||||
<div><dt>STATUS</dt><dd>K3 六轮 · 尖峰路径审计</dd></div>
|
||||
<div><dt>STATUS</dt><dd>K3 七轮 · 局部路径审计</dd></div>
|
||||
</dl>
|
||||
</div>
|
||||
</header>
|
||||
@@ -866,13 +868,14 @@ const paperGroups = [
|
||||
<article><span>O / OBSERVED</span><b>1.4196 TiB tensor data</b><p>96 shards、497,220 entries;不是运行显存,也不是参数量口径。</p></article>
|
||||
<article><span>D / CLOSED LOOP</span><b>69 KDA · 24 MLA · 92 MoE</b><p>配置、tensor names 与 header shape 三方闭合。</p></article>
|
||||
<article><span>X / RTX 5090</span><b>exact 6/6 · max error 0</b><p>官方 torch reference;fixed BF16 mean 2.6210 ms。</p></article>
|
||||
<article class="warning"><span>U / UNRESOLVED</span><b>A_log [128] ≠ expected [96]</b><p>checkpoint 与公开代码 / kernel API 的形状冲突保留在主视区,不擅自解释。</p></article>
|
||||
<article class="warning"><span>U / MAIN UNRESOLVED</span><b>A_log [128] ≠ expected [96]</b><p>main 未修;#144 改成 128,#150 验零后裁成 96,两个社区 PR 都未合并。</p></article>
|
||||
</div>
|
||||
<K3ArtifactLab />
|
||||
<div class="hero-actions">
|
||||
<a class="button primary" href="https://huggingface.co/moonshotai/Kimi-K3">打开官方开放权重</a>
|
||||
<a class="button" href="https://github.com/MoonshotAI/FlashKDA">打开 FlashKDA 官方实现</a>
|
||||
<a class="button" href="https://github.com/MoonshotAI/FlashKDA/blob/master/BENCHMARK_GB200.md">核对作者 GB200 benchmark</a>
|
||||
<a class="button" href="https://huggingface.co/moonshotai/Kimi-K3/discussions/150">审阅社区 PR #150</a>
|
||||
<a class="button" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/experiments/k3/flashkda">复跑本站 RTX 5090 探针</a>
|
||||
</div>
|
||||
</section>
|
||||
@@ -881,8 +884,8 @@ const paperGroups = [
|
||||
<p class="eyebrow"><span>30</span> REDUCED ATTENTION RESIDUALS STUDY</p>
|
||||
<h2>真实 K3 权重还不能诚实 forward;先把一个可证伪的 AttnRes 问题完整做完</h2>
|
||||
<p class="lede">
|
||||
checkpoint 的 <code>A_log [128]</code> 与 config、remote code、FlashKDA、vLLM 和 SGLang
|
||||
期望的 96 heads 仍没有公开转换合同。本轮不裁剪权重冒充 K3,而是预注册一个从零训练的缩小实验:
|
||||
checkpoint 的 <code>A_log [128]</code> 与 main 代码期望的 96 heads 仍没有官方转换合同;
|
||||
社区 #144 / #150 提出相反修复,均未合并。本轮不把候选 patch 冒充 K3 官方 forward,而是预注册一个从零训练的缩小实验:
|
||||
相同 16-block Transformer、相同数据窗口与相同初始化,只改变 residual source 的读取拓扑。
|
||||
</p>
|
||||
<div class="artifact-callout">
|
||||
@@ -949,8 +952,33 @@ const paperGroups = [
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section class="article-section" id="attnres-local-path">
|
||||
<p class="eyebrow"><span>33</span> LOCAL MIXER PATH × BIDIRECTIONAL GATE</p>
|
||||
<h2>全局 value-route 很敏感;能不能把它诚实地缩到 group 6 / 7?</h2>
|
||||
<p class="lede">
|
||||
第七轮先冻结 14 种 same-forward mask:在 learned 背景只把固定 scope 改成 uniform,
|
||||
测 sufficiency;再从 all-uniform 背景只恢复同一 scope 的 detached-learned coefficients,
|
||||
测 restoration。groups 6+7 的 16 个 depth mixers 在 sufficiency 的两个指标、三个 seed
|
||||
全部复现至少一半 global log gap;但 restoration 只在 contrast 通过,peak 三 seed 均低于 50%。
|
||||
因而主结论不是“定位成功”,而是强单侧证据与未闭合的双向 localization。
|
||||
</p>
|
||||
<div class="artifact-callout">
|
||||
<article><span>F / FROZEN</span><b>14 masks · 65 visits</b><p>每次 forward 审计 exact identity set、顺序、唯一性与 census。</p></article>
|
||||
<article><span>X / SUFFICIENCY</span><b>6 / 6 PASS</b><p>groups 6+7 mean S:contrast .677;peak 1.700。</p></article>
|
||||
<article class="warning"><span>X / RESTORATION</span><b>3 / 6 FAIL</b><p>contrast mean .650;peak 仅 .380,三 seed 均未过 .50。</p></article>
|
||||
<article class="warning"><span>B / VERDICT</span><b>one-sided evidence</b><p>双向 localization 未建立;group 7 MLP 只作次级 sufficiency 发现。</p></article>
|
||||
</div>
|
||||
<K3AttnResLocalPathLab />
|
||||
<div class="hero-actions">
|
||||
<a class="button primary" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/research/K3_ATTNRES_LOCAL_PATH_AUDIT.md">阅读完整结果审计</a>
|
||||
<a class="button" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/research/K3_ATTNRES_LOCAL_PATH_PROTOCOL.md">核对预注册协议</a>
|
||||
<a class="button" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/research/K3_ATTNRES_LOCAL_PATH_GROK_REVIEW.md">查看结果前对抗审阅</a>
|
||||
<a class="button" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/experiments/k3/attnres_local_path">复跑矩阵与 replay</a>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section class="article-section" id="audit">
|
||||
<p class="eyebrow"><span>33</span> FIGURE & TABLE AUDIT</p>
|
||||
<p class="eyebrow"><span>34</span> FIGURE & TABLE AUDIT</p>
|
||||
<h2>Figure 1–16、Table 1–5:每张图究竟支持什么,不能支持什么</h2>
|
||||
<div class="figure-atlas">
|
||||
{k3FigureAtlas.map(([id, report, title, contract]) => (
|
||||
|
||||
@@ -9,7 +9,7 @@ const researching = chapters.filter((chapter) => ["researching", "drafting"].inc
|
||||
const workstreams = [
|
||||
{ label: "研究框架与规范", value: 83, next: "给 Scaling 与推理专题补逐篇图表/实验精读层级" },
|
||||
{ label: "网站设计系统", value: 89, next: "打印样式与更多通用可视化组件" },
|
||||
{ label: "Kimi K3 深读", value: 99, next: "对 group 6 / 7 做局部 mixer backward 干预;等待 A_log 官方转换合同" },
|
||||
{ label: "Kimi K3 深读", value: 99, next: "设计前向训练变体,并等待 A_log 社区方案的官方裁决" },
|
||||
{ label: "语言模型前史", value: 78, next: "逐图精读 Kneser–Ney、LSTM 与 Bahdanau,并加入真实小语料复现" },
|
||||
{ label: "Transformer 基础", value: 79, next: "逐图精读多头电路、Pre/Post-LN 与真实 kernel / KV 配置" },
|
||||
{ label: "表示、位置与残差高速公路", value: 81, next: "加入真实 hidden-state / norm traces、长上下文位置外推复现与更多深层稳定性消融" },
|
||||
@@ -50,7 +50,7 @@ const workstreams = [
|
||||
<div><dt>OVERALL</dt><dd>专题平均 {average}%</dd></div>
|
||||
<div><dt>READABLE</dt><dd>{published} 个首版可读专题</dd></div>
|
||||
<div><dt>ACTIVE</dt><dd>{researching} 个研究/写作中</dd></div>
|
||||
<div><dt>UPDATED</dt><dd>2026-07-30 12:20 CST</dd></div>
|
||||
<div><dt>UPDATED</dt><dd>2026-07-30 15:20 CST</dd></div>
|
||||
<div><dt>MODE</dt><dd>持续迭代,不锁死版本</dd></div>
|
||||
</dl>
|
||||
</div>
|
||||
@@ -97,7 +97,7 @@ const workstreams = [
|
||||
<article><span>✓</span><h3>K3 报告已结构化拆解</h3><p>47 页报告目录、151 条参考来源和架构/后训练/系统主线已经提取。</p></article>
|
||||
<article><span>✓</span><h3>17 专题知识图</h3><p>从语言模型基础到评测安全,包含先修依赖和三条贯穿案例。</p></article>
|
||||
<article><span>✓</span><h3>编辑式网站系统</h3><p>响应式导航、章节模板、侧栏、进度、论文链和证据提示组件。</p></article>
|
||||
<article><span>✓</span><h3>一百零四个原创交互视图</h3><p>K3 三轴图、八联报告实验、四联开放工件实验与三轮十五联 AttnRes 独立实验,DeepSeek 四联公式实验、十三联 Base 工件实验、Chat 行为、completion/full-depth、multi-seed、cross-source 与 task-bootstrap CRN 五轮实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
|
||||
<article><span>✓</span><h3>一百零九个原创交互视图</h3><p>K3 三轴图、八联报告实验、四联开放工件实验与四轮二十联 AttnRes 独立实验,DeepSeek 四联公式实验、十三联 Base 工件实验、Chat 行为、completion/full-depth、multi-seed、cross-source 与 task-bootstrap CRN 五轮实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
|
||||
<article><span>✓</span><h3>十七篇首版长文</h3><p>K3、语言模型前史、Transformer、表示/位置/残差、DeepSeek、Scaling、数据工程、长上下文、MoE、后训练、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全专题。</p></article>
|
||||
<article><span>✓</span><h3>语言模型前史深度专题</h3><p>八张独立问题账、33 个正式节点、20 段长文与概率—向量—记忆—对齐四联实验。</p></article>
|
||||
<article><span>✓</span><h3>Transformer 深度专题</h3><p>十张独立问题账、40 个正式节点、21 段正文与 QKV—Mask—多头位置—Block 成本四联实验。</p></article>
|
||||
@@ -108,6 +108,7 @@ const workstreams = [
|
||||
<article><span>✓</span><h3>Kimi K3 四轮 AttnRes 独立实验</h3><p>冻结三结构 × 三 seed 的 9 个 2,000-step 格;Full / Block 相对 Baseline 的平均 paired delta 为 −0.01457 / −0.04247 BPC,但核心参数梯度 CV 没有复现论文叙述。指定正式格全新进程八字段 exact,五视图同时展示结果、反证、成本与 claim boundary。</p></article>
|
||||
<article><span>✓</span><h3>Kimi K3 五轮梯度定义与深度扩展</h3><p>先确认 Figure 5 没有公开唯一 gradient telemetry 合同,再冻结 16/32 blocks × Baseline/Block × 3 seeds 的 12 个 8,000-step 格。Block 的首尾失衡 6/6 改善但全层 CV 6/6 恶化,两个深度都判为 mixed;指定 32 层格完整重训的模型、优化器与全部冻结字段 exact。</p></article>
|
||||
<article><span>✓</span><h3>Kimi K3 六轮尖峰轨迹与反向路径</h3><p>严格复用 Round 05 depth-32 Block 的三个正式格:尖峰在 step 500 后形成,六个位置 3/3 seed 可见,四种 reduction 12/12 格稳健。切断 key/softmax 源梯度没有降低尖峰;uniform value-backward 让 contrast 平均下降 70.2%,只判为全局 backward-rule sensitivity。完整 replay 的 16 组冻结字段 exact。</p></article>
|
||||
<article><span>✓</span><h3>Kimi K3 七轮局部路径双向审计</h3><p>冻结 14 个 same-forward mask,把 groups 6+7 的 16 个 depth mixers 同时放进 sufficiency 与 restoration 两个方向。充分性 6/6 过 50%,恢复性却只有 contrast 3/3 通过、peak 0/3 通过,因此正式状态为 one-sided evidence / localization not established。三个正式格、完整 replay、selector 与 forward identity 全部 exact。</p></article>
|
||||
<article><span>✓</span><h3>FlashKDA RTX 5090 执行闸门</h3><p>隔离 CUDA 13.0 / glibc 2.39 编译 sm_120a wheel;6/6 官方参考逐元素相等,并完成 fixed / varlen、三种 state mode 的 1,800 个 CUDA Event samples。</p></article>
|
||||
<article><span>✓</span><h3>Scaling Laws 深度专题</h3><p>九张账、29 个一手节点、DeepSeek/Kimi 双谱系与曲面—部署—复用—涌现四联实验。</p></article>
|
||||
<article><span>✓</span><h3>数据工程深度专题</h3><p>十二张账、31 个一手节点、DeepSeek/Kimi 双谱系与流水线—去重—混合—改写四联实验。</p></article>
|
||||
@@ -136,7 +137,7 @@ const workstreams = [
|
||||
</div>
|
||||
<div class="queue-table">
|
||||
<div class="head"><b>优先级</b><b>专题</b><b>本轮交付</b><b>完成闸门</b></div>
|
||||
<div><span>P0</span><strong>K3 六轮后续</strong><p>group 6 / 7 局部 mixer intervention matrix → 前向训练变体 → 等待 A_log 官方合同后进入真实 checkpoint forward</p><em>局部机制 + 工件边界</em></div>
|
||||
<div><span>P0</span><strong>K3 七轮后续</strong><p>前向训练变体 → 非加性局部交互地图 → 等待 A_log 社区候选的官方裁决后进入真实 checkpoint forward</p><em>局部机制 + 工件边界</em></div>
|
||||
<div><span>P0</span><strong>DeepSeek 八轮后续</strong><p>干预式 mediation → SM90 FlashMLA / FP8 / pipeline traces → R1-like RL 小模型复现</p><em>运行证据 + 独立复现</em></div>
|
||||
<div><span>P0</span><strong>Transformer 二轮</strong><p>多头电路逐图 → Pre/Post-LN 真实 traces → Flash/KV 配置与 kernel 对照</p><em>逐图笔记 + 实测边界</em></div>
|
||||
<div><span>P0</span><strong>表示、位置与残差二轮</strong><p>真实 hidden-state / norm traces → 长上下文位置外推 → mHC / AttnRes 深层稳定性消融</p><em>可复现实验 + 逐图笔记</em></div>
|
||||
@@ -218,7 +219,7 @@ const workstreams = [
|
||||
<div><time>2026-07-29</time><b>A_log 形状冲突保持未决</b><p>checkpoint 的 [128] 与 config / remote code / FlashKDA API 期待的 [96] 并列展示;不宣布权重损坏,也不把 channel-wise 假设写成真实 forward。</p></div>
|
||||
<div><time>2026-07-29</time><b>FlashKDA 编译与执行永久分两道闸门</b><p>容器产出 sm_120a wheel 只证明可编译;RTX 5090 的 6/6 official-reference exact suite 通过后,才把证据升级为本机执行 X。</p></div>
|
||||
<div><time>2026-07-29</time><b>作者表与 RTX 5090 表永久分账</b><p>H20 / GB200 保持 O;本站只报告独立环境、协议、300 samples/mode 和延迟分布,未跑本机 FLA 就不写本机 speedup。</p></div>
|
||||
<div><time>2026-07-30</time><b>K3 权重冲突不靠裁剪“解决”</b><p>HF / FlashKDA / vLLM / SGLang 仍没有公开 A_log 128→96 转换;真实 K3 forward 继续标为未决。</p></div>
|
||||
<div><time>2026-07-30</time><b>K3 权重冲突仍没有官方裁决</b><p>官方 main 仍保留 128↔96 不匹配;社区 #144 把 head 数改成 128,#150 验证尾部全零后按 96 裁入,两案都未合并。真实 K3 forward 继续标为未决。</p></div>
|
||||
<div><time>2026-07-30</time><b>AttnRes 缩小实验先冻结、后运行</b><p>三结构共享公共主干、初始化、窗口与优化器;只按三个 paired seed 和预注册 −0.010 BPC 阈值给出本协议内方向判断。</p></div>
|
||||
<div><time>2026-07-30</time><b>支持结果与梯度反结果同时进入主视区</b><p>Full / Block 的最终 BPC 同向改善;核心参数 gradient RMS CV 却高于 Baseline,不换指标掩盖。</p></div>
|
||||
<div><time>2026-07-30</time><b>正式重放不把 wall time 纳入 exact</b><p>Block / seed-1 的模型、优化器、曲线、历史、诊断和环境八字段 exact;计时受调度影响,单独报告。</p></div>
|
||||
@@ -229,7 +230,7 @@ const workstreams = [
|
||||
<div><time>2026-07-30</time><b>尖峰是定向复查,不是盲发现</b><p>layer 21–25 来自 Round 05;Round 06 先固定目标集合,再检查训练时点、张量位置、reduction 与反向路径。</p></div>
|
||||
<div><time>2026-07-30</time><b>最早可见不等于物理起源</b><p>pre-attention input 是六个采样点中最早可见位置;更早 mixer 与跨层回传已经作用,不能写成尖峰从这里注入。</p></div>
|
||||
<div><time>2026-07-30</time><b>同一前向只识别反向规则敏感性</b><p>三模式的 logits、loss、activations 与 mixer summaries exact;uniform value-backward 的 70.2% contrast 降幅不是训练变体或因果贡献百分比。</p></div>
|
||||
<div><time>2026-07-30</time><b>相关性、全局干预与局部归因分三层</b><p>MLP latest weight 的局部 r≈.69 只提供候选;全局 value-route 干预支持路径敏感性,下一轮才做 group 6 / 7 局部归因矩阵。</p></div>
|
||||
<div><time>2026-07-30</time><b>局部充分性不自动成为双向定位</b><p>groups 6+7 的 sufficiency 6/6 通过,但 restoration peak 0/3 通过;非线性交互让两个方向不同,正式结论保持 localization not established。</p></div>
|
||||
<div><time>2026-07-29</time><b>32-token 对照改为同源 16→24</b><p>TNEWS 只有 105/10,000 条达到 32 tokens,强行统一会落入约 1% 极端长尾;24-token eligibility 仍保留 1,609 条中文候选。</p></div>
|
||||
<div><time>2026-07-29</time><b>长度敏感性必须成对重采样</b><p>16-token 输入严格是 24-token 输入前缀,2,000 次 bootstrap 共用 prompt indices;结果只描述固定 cohort 的长度敏感性。</p></div>
|
||||
<div><time>2026-07-29</time><b>三类 cohort 永久分身份</b><p>自然长度回答本批样本如何路由;matched-16 / 24 回答同一 prompt 多看 8 tokens 后如何变化,不把二者混成内容因果。</p></div>
|
||||
|
||||
Reference in New Issue
Block a user