Compare commits

...

5 Commits

Author SHA1 Message Date
wuyang f7670efcdd feat: add AttnRes gradient scale lab 2026-07-30 10:24:42 +08:00
wuyang 26fc824409 research: audit AttnRes gradient scale study 2026-07-30 09:49:53 +08:00
wuyang 3f7cc1f544 research: align AttnRes residual precision 2026-07-30 07:54:38 +08:00
wuyang 1675ca54f3 research: lock AttnRes gradient runner 2026-07-30 07:53:12 +08:00
wuyang 091a05f0a0 research: preregister AttnRes gradient scale study 2026-07-30 07:46:07 +08:00
41 changed files with 296691 additions and 17 deletions
+14 -2
View File
@@ -41,7 +41,7 @@
- [x] 完成 486 篇关键论文索引,覆盖 16 个标签专题与 Kimi/DeepSeek 聚光主线。
- [x] 完成可检索、可按专题筛选的论文库页面。
- [x] 完成 K3、语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全十七篇首版长文。
- [x] 完成 K3 三轴架构、八联报告实验、四联开放工件实验与五联 AttnRes 独立实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 二十二联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等九十四个原创交互视图。
- [x] 完成 K3 三轴架构、八联报告实验、四联开放工件实验、Round 04 / 05 各五联 AttnRes 独立实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 二十二联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等九十九个原创交互视图。
- [x] 完成长上下文首版:五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。
- [x] 核验 FlashAttention、DeepSeek-V2/V3.2/V4、Kimi Linear/K3 等六份论文原文,并建立长上下文研究账本。
- [x] 核验 Switch、ST-MoE、DeepSeekMoE、Loss-Free、V3、LatentMoE 与 K3 原文,并建立 MoE 研究账本。
@@ -279,10 +279,18 @@
- [x] K3 Round 04 五视图实验室完成:三 seed BPC 曲线、Residual RMS / Block 锯齿、Full / Block depth-weight heatmap、梯度反证、成本/哈希/claim boundary 分开展示;完整 9-run JSON、compact 数据、复现清单、协议、审计、训练与聚合代码进入公开仓库。
- [x] Round 04 本地闸门通过:91 个受检文件零诊断/提示,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;冻结数据、AttnRes 专项与 K3 全量真实 Chrome 回归通过,桌面/390px 移动端零文档级溢出、零运行时异常。
- [x] K3 Round 04 以功能源提交 `4ce780d`、不可变镜像 `20260729T233142Z-4ce780d` 发布;OCI index digest `sha256:6e89f802…25f582`,复用 NAS `12010→8080`、NPM host 31 / cert 41 与门户 `LLM ATLAS / projects / 180`。容器 healthy、0 次重启,21/21 公网页面、HTTPS/2、gzip / immutable assets、AttnRes 专项与 K3 全量生产 Chrome 回归通过;保留 `20260729T221654Z-975ed3d` 回滚。
- [x] K3 Round 05 一手定义审计确认 Figure 5(c) 未公开 gradient tensor、norm、reduction、diagnostic batch、AMP / clipping 时点或统计代码;Round 04 参数梯度与 Round 05 post-MLP output activation gradient 永久分对象记账,不把本站 operationalization 冒充作者实现。
- [x] 在任何 formal 输出前冻结 16 / 32 blocks、Baseline / Block、三个 seed、8,000 steps、六个诊断点、CV + 首尾四分位 imbalance 联合判据、FP32 residual accumulator、20-step 双 smoke 与 depth-32 Block 完整 replay;768,000-window schedule SHA-256 为 `5041e09b…f4e`。
- [x] 12 个 formal 格全部完成,共 786,432,000 target bytes;公共主干与三个 gate input tensor hashes 在 paired 架构间 exact。Block 的验证 BPC 在 6 / 6 配对中更低,depth-16 / 32 mean delta 为 `−.008938 / −.009872`,但不追加事后 BPC support 阈值。
- [x] activation-gradient 结果分裂:首/末四分位 imbalance 在 6 / 6 配对改善,depth-16 / 32 均值为 `+61.0% / +72.0%`;全层 CV 却在 6 / 6 配对恶化,均值相对 reduction 为 `−10.3% / −60.0%`。两个 depth 都按预注册规则判为 `mixed / inconclusive`,总判定 `depth-dependent or inconclusive`。
- [x] 绝对 gradient mean 仅为 Baseline 的 `57.4% / 54.4%`;参数 gradient CV 从 `0.416→0.683 / 0.397→0.772`,继续保留反结果。Output RMS 最后/第一层比则由 Baseline `4.59× / 6.08×` 降至 Block `1.17× / 1.89×`。
- [x] 指定 depth-32 / Block / seed-2026073001 从初始化完整重训 8,000 steps;排除 run-kind / timing 后冻结字段 compare SHA-256 同为 `46300a45…4817`,model / optimizer state hashes exact。正式/compact/reproduction 物理 SHA-256 为 `ad461cbe…a8d / 5377a5e7…68e3 / aedcde6a…dea6`。
- [x] K3 Round 05 五视图实验室完成:论文定义已知/未定义、绝对/归一化深度谱、六 checkpoint 时间轨迹、Output RMS 组节律、activation/parameter/BPC/成本/重放联合账全部可切换;21 个 raw JSON、完整 aggregate、compact、runner、analyzer、协议与审计进入公开树。
- [x] Round 05 本地闸门通过:94 个 Astro 文件零诊断/提示,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;冻结数据、新专项、Round 04 与 K3 全量真实 Chrome 回归通过,桌面/390px 移动端零文档级溢出、零运行时异常。
## 正在进行
- [ ] K3 四轮下一闸门:对齐 AttnRes 论文的 activation / residual-output gradient 定义,增加模型深度与训练预算,检验本轮梯度反结果是否随尺度翻转;真实 K3 forward 继续等待 `A_log [128]→[96]` 官方转换或权重修订。
- [ ] K3 五轮下一闸门:对齐 layer 21–25 的 activation-gradient 尖峰、pre-attention / pre-MLP 位置与 mixer source weights,并做公开 reduction sensitivity;真实 K3 forward 继续等待 `A_log [128]→[96]` 官方转换或权重修订。
- [ ] DeepSeek 八轮下一闸门:推进干预式 mediation、SM90 FlashMLA、FP8 / pipeline traces 与 R1-like RL 小模型复现。
- [ ] 表示、位置与残差二轮:真实 hidden-state / norm traces、长上下文位置外推复现与 mHC / AttnRes 深层稳定性消融。
- [ ] 评测安全二轮:真实 cross-harness / pass@k 复跑、Judge 元评测、动态污染与过拒案例。
@@ -476,6 +484,10 @@
| 2026-07-30 | 主结果与机制反结果同时发布 | Full / Block BPC 方向支持;核心参数 gradient RMS CV 却高于 Baseline,明确写成未复现论文梯度叙述 |
| 2026-07-30 | 独立重放按数值合同而非计时合同验收 | Block / seed-1 的八组冻结字段 2,000 steps exact;wall time 受调度影响,不要求或声称 bit-exact |
| 2026-07-30 | K3 Round 04 缩小 AttnRes 里程碑发布 | 功能源 `4ce780d`、镜像 `20260729T233142Z-4ce780d`、OCI `sha256:6e89f802…25f582`;21/21 公网页面与生产专项/全量 Chrome 通过,保留 Round 08 回滚点 |
| 2026-07-30 | AttnRes 的“梯度”先按公开证据拆对象 | 论文 Figure 5 没有公开唯一 telemetry 合同;参数梯度与 post-MLP activation gradient 不再互相代称 |
| 2026-07-30 | “更均匀”拆成首尾失衡与全层 CV | Block 6/6 改善 first/last,却 6/6 恶化 CV;局部尖峰与系统性早层隆起必须分开解释 |
| 2026-07-30 | 绝对尺度与归一化形状永久同报 | Block activation-gradient mean 约为 Baseline 54%–57%;不能把更接近 1 的首尾比自动解释为各层信号更强 |
| 2026-07-30 | Round 05 完整重放过闸 | depth-32 Block seed-1 从零重训 8,000 steps;全部冻结字段与 model/optimizer state hashes exact,timing 仍单独报告 |
## 未决问题
+14 -1
View File
@@ -19,7 +19,7 @@
当前里程碑包含 17 专题学习地图、486 篇关键论文索引、Kimi K3 完整导读,
语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 技术谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、工具使用与长程 Agent、原生多模态、训练系统、推理服务、数值优化,以及评测与安全深度专题,
以及 94 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、
以及 99 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、
8 个交互实验和 100 个一手/官方节点,完整覆盖架构、预训练、后训练、系统、评测、案例与附录。
第三轮已完成开放工件与首个真实 kernel 里程碑:固定官方模型与 FlashKDA revisions,审计 96 个 checkpoint shards、
497,220 个 tensor entries、真实 KDA / MLA / MoE / MoonViT shapes 与小范围参数统计,并用 4 个新视图
@@ -39,6 +39,19 @@
[K3_ATTNRES_REDUCED_PROTOCOL.md](./research/K3_ATTNRES_REDUCED_PROTOCOL.md)、
[K3_ATTNRES_REDUCED_AUDIT.md](./research/K3_ATTNRES_REDUCED_AUDIT.md) 与
[AttnRes experiment](./experiments/k3/attnres/)。
第五轮先审计 Attention Residuals Figure 5 的公开定义边界,再冻结
`llm-atlas-k3-attnres-gradient-scale-v1`:以 post-MLP block output activation gradient
为公开 operationalization,把深度扩为 16 / 32 blocks、预算扩为每格 8,000 steps,
完成 Baseline / Block × 三 seed 共 12 格、786,432,000 formal target bytes。结果把
“更均匀”拆成两个相反方向:Block 在 6 / 6 配对中把首/末四分位失衡改善 56%–81%,
却因中后段局部尖峰让全层 CV 在 6 / 6 配对中恶化;两个深度都按预注册联合规则判为
mixed / inconclusive。验证 BPC 仍在 6 / 6 配对中更低,但实际 step time 约 2.6×、
peak allocated memory 约 2.2×,不冒充同算力优势。指定 depth-32 / Block / seed-1
从零重训完整 8,000 steps,全部冻结字段以及 model / optimizer state hashes exact。
详见
[K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md](./research/K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md)、
[K3_ATTNRES_GRADIENT_SCALE_AUDIT.md](./research/K3_ATTNRES_GRADIENT_SCALE_AUDIT.md) 与
[gradient experiment](./experiments/k3/attnres_gradient/)。
DeepSeek 八轮专题以 24 张问题账、10 次技术转向、
22 个交互实验和 60 个一手/官方节点,串起 Dense、MoE、MLA、V3 协同、R1 与 V4;
并固定官方 V2-Lite revision,在 RTX 5090 上连续执行 7/27 层,记录 3,240 次真实专家选择、
+57
View File
@@ -0,0 +1,57 @@
# Attention Residuals activation-gradient/depth study
This directory implements preregistered protocol
`llm-atlas-k3-attnres-gradient-scale-v1`:
- `research/K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md`
- `research/K3_ATTNRES_GRADIENT_SCALE_PROTOCOL.md`
It is an independent reduced mechanism experiment. It is not a Kimi K3
checkpoint forward pass and does not claim to recover the paper's unpublished
Figure 5 telemetry definition.
## Frozen environment
```text
Python /home/wuyang/.pyenv/versions/3.10.14/envs/navi-router-cu128/bin/python
PyTorch 2.11.0+cu128
GPU NVIDIA GeForce RTX 5090
CUBLAS_WORKSPACE_CONFIG=:4096:8
```
## Build the manifest
```bash
python experiments/k3/attnres_gradient/build_dataset.py \
--cache-dir /home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1 \
--manifest experiments/k3/attnres_gradient/manifest.json
```
## Run a smoke cell
```bash
CUBLAS_WORKSPACE_CONFIG=:4096:8 \
python experiments/k3/attnres_gradient/train.py \
--run-kind smoke \
--architecture block \
--depth 32 \
--seed 2026073001 \
--cache-dir /home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1 \
--manifest experiments/k3/attnres_gradient/manifest.json \
--output /home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1/smoke-a/depth-32-block.json
```
Smoke is fixed to 20 steps. Formal and replay runs are fixed to 8,000 steps;
the runner rejects alternative budgets. The same command uses
`--run-kind formal` or `--run-kind replay` and omits an explicit `--steps`.
Formal output keys use:
```text
formal/depth-{16|32}-{baseline|block}-seed-{seed}.json
replay/depth-32-block-seed-2026073001.json
```
Raw parquet/binary files and working runs remain in the local cache. The
manifest, runner, complete result JSON, compact website payload, reproduction
hashes, protocol, and audit enter the public repository.
+685
View File
@@ -0,0 +1,685 @@
#!/usr/bin/env python3
"""Validate, aggregate, and publish K3 AttnRes Round 05 experiment data."""
from __future__ import annotations
import argparse
import hashlib
import json
import math
import os
import shutil
import statistics
from pathlib import Path
from typing import Any, Iterable
PROTOCOL_ID = "llm-atlas-k3-attnres-gradient-scale-v1"
ARCHITECTURES = ("baseline", "block")
DEPTHS = (16, 32)
SEEDS = (2026073001, 2026073002, 2026073003)
STEPS = (0, 100, 500, 2000, 4000, 8000)
FORMAL_STEPS = 8000
FORMAL_BATCH = 32
TARGET_BYTES_PER_RUN = 65_536_000
EXPECTED_TOTAL_TARGET_BYTES = 786_432_000
SMOKE_COMPARE_FIELDS = (
"protocol_id",
"run_kind",
"architecture",
"depth",
"seed",
"steps",
"batch_size",
"target_bytes_seen",
"manifest",
"model",
"optimizer",
"hashes",
"evaluations",
"diagnostics",
"training_history",
"gradient_gate",
"environment",
)
REPLAY_COMPARE_FIELDS = tuple(
field for field in SMOKE_COMPARE_FIELDS if field != "run_kind"
)
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser()
parser.add_argument("--formal-dir", type=Path, required=True)
parser.add_argument("--smoke-a-dir", type=Path, required=True)
parser.add_argument("--smoke-b-dir", type=Path, required=True)
parser.add_argument("--replay", type=Path, required=True)
parser.add_argument("--manifest", type=Path, required=True)
parser.add_argument("--raw-output-dir", type=Path, required=True)
parser.add_argument("--output", type=Path, required=True)
parser.add_argument("--compact-output", type=Path, required=True)
parser.add_argument("--reproduction-output", type=Path, required=True)
return parser.parse_args()
def read_json(path: Path) -> dict[str, Any]:
return json.loads(path.read_text())
def file_sha256(path: Path) -> str:
digest = hashlib.sha256()
with path.open("rb") as handle:
for block in iter(lambda: handle.read(1024 * 1024), b""):
digest.update(block)
return digest.hexdigest()
def canonical_sha256(value: Any) -> str:
payload = json.dumps(
value, ensure_ascii=False, sort_keys=True, separators=(",", ":")
).encode()
return hashlib.sha256(payload).hexdigest()
def atomic_json(path: Path, value: dict[str, Any]) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
temporary = path.with_suffix(path.suffix + ".tmp")
temporary.write_text(
json.dumps(value, ensure_ascii=False, indent=2, sort_keys=True) + "\n"
)
os.replace(temporary, path)
def mean(values: Iterable[float]) -> float:
return statistics.fmean(values)
def require_finite(value: Any, path: str = "root") -> None:
if isinstance(value, float):
if not math.isfinite(value):
raise ValueError(f"non-finite float at {path}")
elif isinstance(value, dict):
for key, child in value.items():
require_finite(child, f"{path}.{key}")
elif isinstance(value, list):
for index, child in enumerate(value):
require_finite(child, f"{path}[{index}]")
def selected(run: dict[str, Any], fields: tuple[str, ...]) -> dict[str, Any]:
return {field: run[field] for field in fields}
def final_diagnostic(run: dict[str, Any]) -> dict[str, Any]:
diagnostic = run["diagnostics"][-1]
if diagnostic["step"] != FORMAL_STEPS:
raise ValueError("final diagnostic is not step 8000")
return diagnostic
def validate_run(
run: dict[str, Any],
*,
path: Path,
depth: int,
architecture: str,
seed: int,
manifest: dict[str, Any],
manifest_hash: str,
) -> None:
if run["protocol_id"] != PROTOCOL_ID or run["run_kind"] != "formal":
raise ValueError(f"formal protocol/kind mismatch: {path}")
if (
run["depth"] != depth
or run["architecture"] != architecture
or run["seed"] != seed
):
raise ValueError(f"formal identity mismatch: {path}")
if (
run["steps"] != FORMAL_STEPS
or run["batch_size"] != FORMAL_BATCH
or run["target_bytes_seen"] != TARGET_BYTES_PER_RUN
):
raise ValueError(f"formal budget mismatch: {path}")
if run["manifest"]["file_sha256"] != manifest_hash:
raise ValueError(f"manifest file hash mismatch: {path}")
for key in (
"formal_schedule_sha256",
"validation_tensor_sha256",
"diagnostic_tensor_sha256",
):
if run["manifest"][key] != manifest["windows"][key]:
raise ValueError(f"manifest {key} mismatch: {path}")
if [row["step"] for row in run["evaluations"]] != list(STEPS):
raise ValueError(f"evaluation steps mismatch: {path}")
if [row["step"] for row in run["diagnostics"]] != list(STEPS):
raise ValueError(f"diagnostic steps mismatch: {path}")
if run["model"]["layers"] != depth:
raise ValueError(f"model depth mismatch: {path}")
for diagnostic in run["diagnostics"]:
capture = diagnostic["capture"]
if (
capture["count"] != depth
or capture["shape"] != [16, 256, 192]
or set(capture["dtypes"]) != {"torch.float32"}
or not capture["all_gradients_finite"]
or not capture["all_gradients_present"]
or not capture["storage_unique"]
):
raise ValueError(f"activation capture gate mismatch: {path}")
for field in (
"activation_grad_rms_by_block",
"activation_output_rms_by_block",
"core_parameter_grad_rms_by_block",
):
if len(diagnostic[field]) != depth:
raise ValueError(f"{field} length mismatch: {path}")
for field in (
"layer_input_rms_by_sublayer",
"branch_output_rms_by_sublayer",
"stream_state_rms_by_sublayer",
):
if len(diagnostic[field]) != depth * 2:
raise ValueError(f"{field} length mismatch: {path}")
if architecture == "baseline":
if diagnostic["depth_weights"] or diagnostic["output_weights"] is not None:
raise ValueError(f"unexpected Baseline mixer trace: {path}")
else:
if (
len(diagnostic["depth_weights"]) != depth * 2
or diagnostic["output_weights"]["sources"] != 9
):
raise ValueError(f"Block mixer trace mismatch: {path}")
require_finite(run, path.name)
def relative_reduction(baseline: float, block: float) -> float:
return (baseline - block) / baseline
def depth_verdict(rows: list[dict[str, Any]]) -> dict[str, Any]:
cv_reductions = [row["relative_cv_reduction"] for row in rows]
imbalance_reductions = [
row["relative_imbalance_reduction"] for row in rows
]
mean_cv = mean(cv_reductions)
mean_imbalance = mean(imbalance_reductions)
support = (
all(value > 0 for value in cv_reductions)
and mean_cv >= 0.20
and all(value > 0 for value in imbalance_reductions)
and mean_imbalance >= 0.20
)
concern = (
all(value < 0 for value in cv_reductions)
and -mean_cv >= 0.20
and all(value < 0 for value in imbalance_reductions)
and -mean_imbalance >= 0.20
)
if support:
label = "joint directional support at this depth"
elif concern:
label = "joint directional concern at this depth"
else:
label = "mixed / inconclusive at this depth"
return {
"label": label,
"threshold_relative": 0.20,
"cv_reductions": cv_reductions,
"mean_cv_reduction": mean_cv,
"imbalance_reductions": imbalance_reductions,
"mean_imbalance_reduction": mean_imbalance,
"all_cv_improve": all(value > 0 for value in cv_reductions),
"all_imbalance_improve": all(
value > 0 for value in imbalance_reductions
),
}
def summarize_depth(
depth: int, runs: dict[tuple[int, str, int], dict[str, Any]]
) -> dict[str, Any]:
rows = []
for seed in SEEDS:
baseline = runs[(depth, "baseline", seed)]
block = runs[(depth, "block", seed)]
baseline_diagnostic = final_diagnostic(baseline)
block_diagnostic = final_diagnostic(block)
baseline_activation = baseline_diagnostic[
"activation_grad_statistics"
]
block_activation = block_diagnostic["activation_grad_statistics"]
baseline_bpc = baseline["evaluations"][-1]["bits_per_byte"]
block_bpc = block["evaluations"][-1]["bits_per_byte"]
rows.append(
{
"seed": seed,
"baseline_bpc": baseline_bpc,
"block_bpc": block_bpc,
"block_minus_baseline_bpc": block_bpc - baseline_bpc,
"baseline_activation_grad_mean": baseline_activation["mean"],
"block_activation_grad_mean": block_activation["mean"],
"block_to_baseline_activation_grad_mean": (
block_activation["mean"] / baseline_activation["mean"]
),
"baseline_cv": baseline_activation["population_cv"],
"block_cv": block_activation["population_cv"],
"relative_cv_reduction": relative_reduction(
baseline_activation["population_cv"],
block_activation["population_cv"],
),
"baseline_first_to_last_ratio": baseline_activation[
"first_to_last_ratio"
],
"block_first_to_last_ratio": block_activation[
"first_to_last_ratio"
],
"baseline_imbalance": baseline_activation[
"imbalance_abs_log_ratio"
],
"block_imbalance": block_activation[
"imbalance_abs_log_ratio"
],
"relative_imbalance_reduction": relative_reduction(
baseline_activation["imbalance_abs_log_ratio"],
block_activation["imbalance_abs_log_ratio"],
),
"baseline_parameter_grad_cv": baseline_diagnostic[
"core_parameter_grad_statistics"
]["population_cv"],
"block_parameter_grad_cv": block_diagnostic[
"core_parameter_grad_statistics"
]["population_cv"],
}
)
verdict = depth_verdict(rows)
return {
"depth": depth,
"by_seed": rows,
"means": {
"baseline_bpc": mean(row["baseline_bpc"] for row in rows),
"block_bpc": mean(row["block_bpc"] for row in rows),
"block_minus_baseline_bpc": mean(
row["block_minus_baseline_bpc"] for row in rows
),
"baseline_cv": mean(row["baseline_cv"] for row in rows),
"block_cv": mean(row["block_cv"] for row in rows),
"relative_cv_reduction": mean(
row["relative_cv_reduction"] for row in rows
),
"baseline_imbalance": mean(
row["baseline_imbalance"] for row in rows
),
"block_imbalance": mean(
row["block_imbalance"] for row in rows
),
"relative_imbalance_reduction": mean(
row["relative_imbalance_reduction"] for row in rows
),
"block_to_baseline_activation_grad_mean": mean(
row["block_to_baseline_activation_grad_mean"] for row in rows
),
"baseline_parameter_grad_cv": mean(
row["baseline_parameter_grad_cv"] for row in rows
),
"block_parameter_grad_cv": mean(
row["block_parameter_grad_cv"] for row in rows
),
},
"verdict": verdict,
}
def compact_cell(run: dict[str, Any]) -> dict[str, Any]:
diagnostics = []
for row in run["diagnostics"]:
compact = {
"step": row["step"],
"loss_nats": row["loss_nats"],
"bits_per_byte": row["bits_per_byte"],
"activation_grad_rms_by_block": row[
"activation_grad_rms_by_block"
],
"activation_grad_statistics": row[
"activation_grad_statistics"
],
"activation_output_rms_by_block": row[
"activation_output_rms_by_block"
],
"activation_output_statistics": row[
"activation_output_statistics"
],
"core_parameter_grad_rms_by_block": row[
"core_parameter_grad_rms_by_block"
],
"core_parameter_grad_statistics": row[
"core_parameter_grad_statistics"
],
}
if row["depth_weights"]:
compact["depth_weights"] = row["depth_weights"]
compact["output_weights"] = row["output_weights"]
diagnostics.append(compact)
return {
"architecture": run["architecture"],
"depth": run["depth"],
"seed": run["seed"],
"evaluations": run["evaluations"],
"diagnostics": diagnostics,
"timing": run["timing"],
"parameters": run["model"]["parameters"],
"hashes": run["hashes"],
}
def main() -> None:
args = parse_args()
manifest = read_json(args.manifest)
manifest_hash = file_sha256(args.manifest)
if manifest["protocol_id"] != PROTOCOL_ID:
raise ValueError("manifest protocol mismatch")
if (
manifest["windows"]["formal_schedule_cells"] != 768_000
or manifest["windows"]["formal_steps"] != FORMAL_STEPS
or manifest["windows"]["formal_batch"] != FORMAL_BATCH
):
raise ValueError("manifest schedule budget mismatch")
runs: dict[tuple[int, str, int], dict[str, Any]] = {}
source_paths: dict[str, Path] = {}
formal_hashes: dict[str, str] = {}
for depth in DEPTHS:
for architecture in ARCHITECTURES:
for seed in SEEDS:
name = f"depth-{depth}-{architecture}-seed-{seed}.json"
path = args.formal_dir / name
run = read_json(path)
validate_run(
run,
path=path,
depth=depth,
architecture=architecture,
seed=seed,
manifest=manifest,
manifest_hash=manifest_hash,
)
runs[(depth, architecture, seed)] = run
public_name = f"formal-{name}"
source_paths[public_name] = path
formal_hashes[public_name] = file_sha256(path)
if sum(run["target_bytes_seen"] for run in runs.values()) != (
EXPECTED_TOTAL_TARGET_BYTES
):
raise ValueError("formal total target-byte budget mismatch")
common_initial_exact: dict[str, Any] = {}
input_gate_exact: dict[str, Any] = {}
for depth in DEPTHS:
for seed in SEEDS:
baseline = runs[(depth, "baseline", seed)]
block = runs[(depth, "block", seed)]
public_fields = (
"initial_public_parameter_structure",
"initial_public_parameter_tensors",
"initial_public_parameter_elements",
"initial_public_parameters",
)
exact = all(
baseline["hashes"][field] == block["hashes"][field]
for field in public_fields
)
gate_exact = (
baseline["manifest"]["input_gate_tensor_hashes"]
== block["manifest"]["input_gate_tensor_hashes"]
)
key = f"depth-{depth}-seed-{seed}"
common_initial_exact[key] = {
"exact": exact,
"baseline": {
field: baseline["hashes"][field] for field in public_fields
},
"block": {
field: block["hashes"][field] for field in public_fields
},
}
input_gate_exact[key] = {
"exact": gate_exact,
"hashes": baseline["manifest"]["input_gate_tensor_hashes"],
}
if not exact or not gate_exact:
raise ValueError(f"paired equality gate failed: {key}")
smoke_exact: dict[str, Any] = {}
smoke_hashes: dict[str, str] = {}
for depth in DEPTHS:
for architecture in ARCHITECTURES:
name = f"depth-{depth}-{architecture}.json"
left_path = args.smoke_a_dir / name
right_path = args.smoke_b_dir / name
left = read_json(left_path)
right = read_json(right_path)
left_selected = selected(left, SMOKE_COMPARE_FIELDS)
right_selected = selected(right, SMOKE_COMPARE_FIELDS)
exact = left_selected == right_selected
if (
not exact
or left["run_kind"] != "smoke"
or left["steps"] != 20
or not left["gradient_gate"]["passed"]
):
raise ValueError(f"smoke gate failed: {name}")
key = f"depth-{depth}-{architecture}"
smoke_exact[key] = {
"exact": exact,
"compare_sha256": canonical_sha256(left_selected),
"gradient_gate": left["gradient_gate"],
}
for label, path in (("a", left_path), ("b", right_path)):
public_name = f"smoke-{label}-{name}"
source_paths[public_name] = path
smoke_hashes[public_name] = file_sha256(path)
replay = read_json(args.replay)
replay_formal = runs[(32, "block", 2026073001)]
replay_left = selected(replay_formal, REPLAY_COMPARE_FIELDS)
replay_right = selected(replay, REPLAY_COMPARE_FIELDS)
replay_exact = replay_left == replay_right
if (
replay["run_kind"] != "replay"
or replay["depth"] != 32
or replay["architecture"] != "block"
or replay["seed"] != 2026073001
or not replay_exact
):
raise ValueError("formal replay gate failed")
replay_public_name = "replay-depth-32-block-seed-2026073001.json"
source_paths[replay_public_name] = args.replay
depth_summaries = {
str(depth): summarize_depth(depth, runs) for depth in DEPTHS
}
depth_labels = [
depth_summaries[str(depth)]["verdict"]["label"] for depth in DEPTHS
]
if all(
label == "joint directional support at this depth"
for label in depth_labels
):
overall_verdict = (
"scale-consistent directional support in this operationalization"
)
elif all(
label == "joint directional concern at this depth"
for label in depth_labels
):
overall_verdict = (
"scale-consistent directional concern in this operationalization"
)
else:
overall_verdict = "depth-dependent or inconclusive"
full = {
"schema_version": 1,
"protocol_id": PROTOCOL_ID,
"manifest": manifest,
"study": {
"architectures": list(ARCHITECTURES),
"depths": list(DEPTHS),
"seeds": list(SEEDS),
"diagnostic_steps": list(STEPS),
"formal_runs": len(runs),
"formal_target_bytes": EXPECTED_TOTAL_TARGET_BYTES,
"replay_target_bytes": TARGET_BYTES_PER_RUN,
"gradient_object": (
"RMS of d(mean token CE)/d(post-MLP Transformer-block output) "
"over batch×time×channel"
),
},
"depth_summaries": depth_summaries,
"overall_verdict": overall_verdict,
"gates": {
"common_initial_parameters": common_initial_exact,
"paired_input_tensors": input_gate_exact,
"smoke_exact": smoke_exact,
"replay": {
"exact": replay_exact,
"compare_fields": list(REPLAY_COMPARE_FIELDS),
"formal_compare_sha256": canonical_sha256(replay_left),
"replay_compare_sha256": canonical_sha256(replay_right),
"formal_final_model_state": replay_formal["hashes"][
"final_model_state"
],
"replay_final_model_state": replay["hashes"][
"final_model_state"
],
"formal_final_optimizer_state": replay_formal["hashes"][
"final_optimizer_state"
],
"replay_final_optimizer_state": replay["hashes"][
"final_optimizer_state"
],
},
},
"runs": {
f"depth-{depth}-{architecture}-seed-{seed}": run
for (depth, architecture, seed), run in sorted(runs.items())
},
}
full["canonical_sha256_without_self"] = canonical_sha256(full)
compact = {
"schema_version": 1,
"protocol_id": PROTOCOL_ID,
"study": full["study"],
"manifest_summary": {
"file_sha256": manifest_hash,
"dataset_revision": manifest["dataset"]["revision"],
"train_bytes_sha256": manifest["dataset"]["splits"]["train"][
"concatenated_sha256"
],
"formal_schedule_sha256": manifest["windows"][
"formal_schedule_sha256"
],
"validation_tensor_sha256": manifest["windows"][
"validation_tensor_sha256"
],
"diagnostic_tensor_sha256": manifest["windows"][
"diagnostic_tensor_sha256"
],
},
"depth_summaries": depth_summaries,
"overall_verdict": overall_verdict,
"replay_exact": replay_exact,
"cells": [
compact_cell(runs[(depth, architecture, seed)])
for depth in DEPTHS
for architecture in ARCHITECTURES
for seed in SEEDS
],
}
compact["canonical_sha256_without_self"] = canonical_sha256(compact)
args.raw_output_dir.mkdir(parents=True, exist_ok=True)
for public_name, source_path in sorted(source_paths.items()):
target = args.raw_output_dir / public_name
temporary = target.with_suffix(target.suffix + ".tmp")
shutil.copyfile(source_path, temporary)
os.replace(temporary, target)
reproduction = {
"schema_version": 1,
"protocol_id": PROTOCOL_ID,
"manifest": {
"path": str(args.manifest),
"sha256": manifest_hash,
},
"protocol_sha256": file_sha256(
Path("research/K3_ATTNRES_GRADIENT_SCALE_PROTOCOL.md")
),
"definition_audit_sha256": file_sha256(
Path("research/K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md")
),
"runner_sha256": file_sha256(
Path("experiments/k3/attnres_gradient/train.py")
),
"analyzer_sha256": file_sha256(Path(__file__)),
"formal_raw_sha256": formal_hashes,
"smoke_raw_sha256": smoke_hashes,
"replay_raw_sha256": {
replay_public_name: file_sha256(args.replay)
},
"formal_target_bytes": EXPECTED_TOTAL_TARGET_BYTES,
"replay_target_bytes": TARGET_BYTES_PER_RUN,
"smoke_exact": smoke_exact,
"replay_exact": {
"exact": replay_exact,
"compare_sha256": canonical_sha256(replay_left),
"final_model_state": replay["hashes"]["final_model_state"],
"final_optimizer_state": replay["hashes"][
"final_optimizer_state"
],
},
"aggregate_sha256": full["canonical_sha256_without_self"],
"compact_sha256": compact["canonical_sha256_without_self"],
"overall_verdict": overall_verdict,
}
reproduction["canonical_sha256_without_self"] = canonical_sha256(
reproduction
)
atomic_json(args.output, full)
atomic_json(args.compact_output, compact)
atomic_json(args.reproduction_output, reproduction)
print(
json.dumps(
{
"formal_runs": len(runs),
"formal_target_bytes": EXPECTED_TOTAL_TARGET_BYTES,
"smoke_exact": all(
row["exact"] for row in smoke_exact.values()
),
"replay_exact": replay_exact,
"depth_verdicts": {
depth: depth_summaries[str(depth)]["verdict"]["label"]
for depth in DEPTHS
},
"overall_verdict": overall_verdict,
"aggregate_sha256": full[
"canonical_sha256_without_self"
],
"compact_sha256": compact[
"canonical_sha256_without_self"
],
"reproduction_sha256": reproduction[
"canonical_sha256_without_self"
],
},
ensure_ascii=False,
indent=2,
)
)
if __name__ == "__main__":
main()
@@ -0,0 +1,206 @@
#!/usr/bin/env python3
"""Freeze the byte-level corpus and window schedule for K3 AttnRes Round 05."""
from __future__ import annotations
import argparse
import hashlib
import json
import os
import urllib.request
from pathlib import Path
from typing import Any
import pyarrow.parquet as pq
PROTOCOL_ID = "llm-atlas-k3-attnres-gradient-scale-v1"
DATASET_REPO = "Salesforce/wikitext"
DATASET_REVISION = "b08601e04326c79dfdd32d625aee71d232d685c3"
DATASET_VARIANT = "wikitext-2-raw-v1"
SPLITS = ("train", "validation", "test")
SEEDS = (2026073001, 2026073002, 2026073003)
CONTEXT = 256
FORMAL_STEPS = 8000
FORMAL_BATCH = 32
VALIDATION_WINDOWS = 64
DIAGNOSTIC_WINDOWS = 16
GATE_STEPS = (0, 1, 7999)
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser()
parser.add_argument("--cache-dir", type=Path, required=True)
parser.add_argument("--manifest", type=Path, required=True)
return parser.parse_args()
def file_sha256(path: Path) -> str:
digest = hashlib.sha256()
with path.open("rb") as handle:
for block in iter(lambda: handle.read(1024 * 1024), b""):
digest.update(block)
return digest.hexdigest()
def atomic_json(path: Path, value: dict[str, Any]) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
temporary = path.with_suffix(path.suffix + ".tmp")
temporary.write_text(
json.dumps(value, ensure_ascii=False, indent=2, sort_keys=True) + "\n"
)
os.replace(temporary, path)
def download(url: str, path: Path) -> None:
if path.exists():
return
path.parent.mkdir(parents=True, exist_ok=True)
temporary = path.with_suffix(path.suffix + ".part")
request = urllib.request.Request(
url,
headers={"User-Agent": "llm-atlas-k3-attnres-gradient-scale/1.0"},
)
with urllib.request.urlopen(request, timeout=120) as response:
with temporary.open("wb") as output:
while block := response.read(1024 * 1024):
output.write(block)
os.replace(temporary, path)
def hashed_start(fields: list[str], corpus_length: int) -> int:
value = int.from_bytes(
hashlib.sha256("\0".join(fields).encode()).digest()[:8], "big"
)
return value % (corpus_length - (CONTEXT + 1))
def fixed_window_start(label: str, index: int, corpus_length: int) -> int:
return hashed_start([PROTOCOL_ID, label, str(index)], corpus_length)
def train_window_start(seed: int, step: int, row: int, corpus_length: int) -> int:
return hashed_start(
[PROTOCOL_ID, "train-window", str(seed), str(step), str(row)],
corpus_length,
)
def concatenate_split(parquet_path: Path) -> tuple[bytes, int]:
table = pq.read_table(parquet_path, columns=["text"])
rows = table.column("text").to_pylist()
payload = b"".join(((row or "") + "\n").encode("utf-8") for row in rows)
return payload, len(rows)
def tensor_hash(payload: bytes, starts: list[int]) -> str:
digest = hashlib.sha256()
for start in starts:
digest.update(payload[start : start + CONTEXT + 1])
return digest.hexdigest()
def main() -> None:
args = parse_args()
args.cache_dir.mkdir(parents=True, exist_ok=True)
split_manifest: dict[str, Any] = {}
split_bytes: dict[str, bytes] = {}
for split in SPLITS:
relative = f"{DATASET_VARIANT}/{split}-00000-of-00001.parquet"
url = (
f"https://huggingface.co/datasets/{DATASET_REPO}/resolve/"
f"{DATASET_REVISION}/{relative}"
)
parquet_path = args.cache_dir / f"{split}.parquet"
download(url, parquet_path)
payload, rows = concatenate_split(parquet_path)
binary_path = args.cache_dir / f"{split}.bin"
if not binary_path.exists() or file_sha256(binary_path) != hashlib.sha256(
payload
).hexdigest():
temporary = binary_path.with_suffix(".bin.tmp")
temporary.write_bytes(payload)
os.replace(temporary, binary_path)
split_bytes[split] = payload
split_manifest[split] = {
"source_path": relative,
"source_url": url,
"parquet_bytes": parquet_path.stat().st_size,
"parquet_sha256": file_sha256(parquet_path),
"rows": rows,
"concatenated_bytes": len(payload),
"concatenated_sha256": hashlib.sha256(payload).hexdigest(),
"binary_path": str(binary_path),
"binary_sha256": file_sha256(binary_path),
}
train = split_bytes["train"]
validation = split_bytes["validation"]
schedule_digest = hashlib.sha256()
schedule_cells = 0
for seed in SEEDS:
for step in range(1, FORMAL_STEPS + 1):
for row in range(FORMAL_BATCH):
start = train_window_start(seed, step, row, len(train))
schedule_digest.update(start.to_bytes(8, "big"))
schedule_cells += 1
validation_starts = [
fixed_window_start("validation-window", index, len(validation))
for index in range(VALIDATION_WINDOWS)
]
diagnostic_starts = [
fixed_window_start("diagnostic-window", index, len(validation))
for index in range(DIAGNOSTIC_WINDOWS)
]
gate_tensor_hashes: dict[str, dict[str, str]] = {}
for seed in SEEDS:
gate_tensor_hashes[str(seed)] = {}
for step in GATE_STEPS:
starts = [
train_window_start(seed, step, row, len(train))
for row in range(FORMAL_BATCH)
]
gate_tensor_hashes[str(seed)][str(step)] = tensor_hash(train, starts)
manifest = {
"schema_version": 1,
"protocol_id": PROTOCOL_ID,
"status": "frozen-before-model-output",
"dataset": {
"repository": DATASET_REPO,
"revision": DATASET_REVISION,
"variant": DATASET_VARIANT,
"preprocessing": (
"parquet row order; (text or empty string) + LF; UTF-8; "
"no normalization; vocabulary is raw bytes 0..255"
),
"splits": split_manifest,
},
"windows": {
"context": CONTEXT,
"target_bytes_per_window": CONTEXT,
"seeds": list(SEEDS),
"formal_steps": FORMAL_STEPS,
"formal_batch": FORMAL_BATCH,
"formal_schedule_cells": schedule_cells,
"formal_schedule_sha256": schedule_digest.hexdigest(),
"validation_starts": validation_starts,
"validation_tensor_sha256": tensor_hash(
validation, validation_starts
),
"diagnostic_starts": diagnostic_starts,
"diagnostic_tensor_sha256": tensor_hash(
validation, diagnostic_starts
),
"gate_steps": list(GATE_STEPS),
"gate_training_tensor_sha256": gate_tensor_hashes,
},
}
atomic_json(args.manifest, manifest)
print(json.dumps(manifest, ensure_ascii=False, indent=2))
if __name__ == "__main__":
main()
@@ -0,0 +1,167 @@
{
"dataset": {
"preprocessing": "parquet row order; (text or empty string) + LF; UTF-8; no normalization; vocabulary is raw bytes 0..255",
"repository": "Salesforce/wikitext",
"revision": "b08601e04326c79dfdd32d625aee71d232d685c3",
"splits": {
"test": {
"binary_path": "/home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1/test.bin",
"binary_sha256": "bfe9eb16ab9987fb88bde4ea9a30a00f2a45db01dfc14bad78d05325789c4f12",
"concatenated_bytes": 1292014,
"concatenated_sha256": "bfe9eb16ab9987fb88bde4ea9a30a00f2a45db01dfc14bad78d05325789c4f12",
"parquet_bytes": 732610,
"parquet_sha256": "5f1bea067869d04849c0f975a2b29c4ff47d867f484f5010ea5e861eab246d91",
"rows": 4358,
"source_path": "wikitext-2-raw-v1/test-00000-of-00001.parquet",
"source_url": "https://huggingface.co/datasets/Salesforce/wikitext/resolve/b08601e04326c79dfdd32d625aee71d232d685c3/wikitext-2-raw-v1/test-00000-of-00001.parquet"
},
"train": {
"binary_path": "/home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1/train.bin",
"binary_sha256": "0ca7d3e74dbe44564ea5942b85232f1bbcb525c9cd481cd5d28a87ee90e7e9b4",
"concatenated_bytes": 10951563,
"concatenated_sha256": "0ca7d3e74dbe44564ea5942b85232f1bbcb525c9cd481cd5d28a87ee90e7e9b4",
"parquet_bytes": 6357543,
"parquet_sha256": "e83889baabc497075506f91975be5fac0d45c5290b6b20582c8cd1e853d0c9f7",
"rows": 36718,
"source_path": "wikitext-2-raw-v1/train-00000-of-00001.parquet",
"source_url": "https://huggingface.co/datasets/Salesforce/wikitext/resolve/b08601e04326c79dfdd32d625aee71d232d685c3/wikitext-2-raw-v1/train-00000-of-00001.parquet"
},
"validation": {
"binary_path": "/home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1/validation.bin",
"binary_sha256": "a42356f6a8ff1d25daf25ec9db49e10a537c265581b61c74604bb63231dee719",
"concatenated_bytes": 1148008,
"concatenated_sha256": "a42356f6a8ff1d25daf25ec9db49e10a537c265581b61c74604bb63231dee719",
"parquet_bytes": 657209,
"parquet_sha256": "204929b7ff9d6184953f867dedb860e40aa69c078fc1e54b3baaa8fb28511c4c",
"rows": 3760,
"source_path": "wikitext-2-raw-v1/validation-00000-of-00001.parquet",
"source_url": "https://huggingface.co/datasets/Salesforce/wikitext/resolve/b08601e04326c79dfdd32d625aee71d232d685c3/wikitext-2-raw-v1/validation-00000-of-00001.parquet"
}
},
"variant": "wikitext-2-raw-v1"
},
"protocol_id": "llm-atlas-k3-attnres-gradient-scale-v1",
"schema_version": 1,
"status": "frozen-before-model-output",
"windows": {
"context": 256,
"diagnostic_starts": [
399861,
210983,
449025,
1098024,
323754,
152932,
1091551,
1078021,
415985,
612288,
910624,
530272,
827285,
765798,
1086876,
1035447
],
"diagnostic_tensor_sha256": "21117e31db302b10d67b63f035665dc8f220b879d216ccd12b7d2ba86e7b1716",
"formal_batch": 32,
"formal_schedule_cells": 768000,
"formal_schedule_sha256": "5041e09b167f229248d2462324e8c254b8f5938975f135dcd8192b00a54a4f4e",
"formal_steps": 8000,
"gate_steps": [
0,
1,
7999
],
"gate_training_tensor_sha256": {
"2026073001": {
"0": "52fdd6885cc2bef8e29cedec8c293e1ea71f63fe3cd0b83f627fe640e19a95c5",
"1": "9d0a960595a3f57cd18834d880bde56fa8dcbb0b1cbed26bcc9bbf773a67949c",
"7999": "8f4f04a889d1f8c9eb75e917dc7cd6ddef1196bdea94e87466270db7be1289ee"
},
"2026073002": {
"0": "156813ff7ab93736c8dba340711a9633b6c952d0f54df30f277e254705cc8b82",
"1": "9ee5a434bdf417d658a1485135d48749e98c86f03239eef86dbd81a9d1309da0",
"7999": "bca78ffefa000dc3693a790d65933251646c70facc64b42007d500c19bb90977"
},
"2026073003": {
"0": "8ba13ad55eb54501ec443430446793ef41dea21fd067b2a5b0c84eda202ed0ba",
"1": "326a12f07637fe70fd39fa2758b389e91dacc7f2f2ede708b80e09c9d4d1c39d",
"7999": "7d6b10c3febfb697444909f695d3a45297789022714f8b71dadfec4d6236aea6"
}
},
"seeds": [
2026073001,
2026073002,
2026073003
],
"target_bytes_per_window": 256,
"validation_starts": [
19785,
600807,
1029500,
319878,
652204,
1072662,
679625,
1027264,
500250,
944134,
187313,
834295,
968556,
550645,
239009,
452519,
356354,
134015,
64555,
397632,
203140,
346032,
314812,
10817,
1141274,
807645,
417975,
870687,
377265,
635426,
597238,
805324,
12300,
264343,
84743,
596894,
188690,
992517,
854512,
427504,
94167,
296670,
760313,
912279,
1054297,
81970,
419690,
971472,
1041491,
669963,
735537,
434513,
169153,
6229,
136413,
1098303,
400950,
457810,
659776,
911665,
909832,
532969,
555820,
1019138
],
"validation_tensor_sha256": "f459316f13078a163b47c133511bb7181e05170ab89516e196490113893ce338"
}
}
@@ -0,0 +1,217 @@
{
"aggregate_sha256": "69be133c8f5b11f8a58e831bc03f5432e737ed21efe89bf186784ead7c7f7b51",
"analyzer_sha256": "017d38d92bfd5bc26e10028015f9bf85d28516fb5c5bae06d7bf315b5eca2fc8",
"canonical_sha256_without_self": "addb2e5990af4bbec4b88b4654fc72464f4df5b57fd00e46cad1febbb042f68a",
"compact_sha256": "8cdb71808a429e5f9513c1c6c9426bdc93e30f237001a1b47a2a2fba574a044f",
"definition_audit_sha256": "79221c5648ba280d9177386d2fcfb2fa653b3dd14405185b5997754db2961cc6",
"formal_raw_sha256": {
"formal-depth-16-baseline-seed-2026073001.json": "c9144ddb458010868668ab0bb1272d4e5051300936d9b21dccee75c32a9e46de",
"formal-depth-16-baseline-seed-2026073002.json": "939cad5054960ba8a72b50b4b250b94c3d22a6b46796f669c6e4f033d59fb9de",
"formal-depth-16-baseline-seed-2026073003.json": "afd9a60ff323e97d21b24b413d9b9fa97b8d5f960bfc88d2040502aa5ad96a1e",
"formal-depth-16-block-seed-2026073001.json": "23e3c68e54ec55d981066ad5730316db9084a6637a8064763952ba109ea1f715",
"formal-depth-16-block-seed-2026073002.json": "027f8786f530d6ad5d910f0aad3d0f287887d6457c670bc8b50af7248be1487d",
"formal-depth-16-block-seed-2026073003.json": "254b637f403a4c61e8e8bb2dd083349406b6baa7a3d8a616e37a9ab6f8ebaeb9",
"formal-depth-32-baseline-seed-2026073001.json": "232ed3dfbbbc898c42622c4a9aee8400f76cb773c576bf1c522cdbb19ead998c",
"formal-depth-32-baseline-seed-2026073002.json": "f4fb5fca608a6623ce4c9707714a8a9ebd2c10f6e63f54e243ec3c06d18b1fff",
"formal-depth-32-baseline-seed-2026073003.json": "1b0bd279b165c3a2431685d8c1e6bc36ff72eb55333d5de689dca4f89506bb5e",
"formal-depth-32-block-seed-2026073001.json": "29e1d638b481619c7b32de402122523db8b17cc1fc67a8881fa1ba132a1d5c38",
"formal-depth-32-block-seed-2026073002.json": "21199deb2199395061e51220fd8c7afd04a1135a6381e406da9b5795e3ad5032",
"formal-depth-32-block-seed-2026073003.json": "c0d7f1bcfa9fa7f8f3134ca4571bdf23a951182d03b7de9611a6b4b89667d4ac"
},
"formal_target_bytes": 786432000,
"manifest": {
"path": "experiments/k3/attnres_gradient/manifest.json",
"sha256": "080afb17d1e036c0bba0a799fdb8b98ee4ad652bd42dd1b3b67110dd2ede6371"
},
"overall_verdict": "depth-dependent or inconclusive",
"protocol_id": "llm-atlas-k3-attnres-gradient-scale-v1",
"protocol_sha256": "f772629b3b82975b6756721c3a3bb57cc4171dfa26ba5b1e8043dc91c9dcce22",
"replay_exact": {
"compare_sha256": "46300a452840a9dc6a5180efe7949cf3d81cc4471ed4a942da407e9343064817",
"exact": true,
"final_model_state": "3f0b97ece3a15571ba3d656f589f512ca0bb9e20083c9f58a42ccaee14892f59",
"final_optimizer_state": "ed03e6fbd4a12d8b063dcb22e0437754285f54d585374cd52fbd534f05d24637"
},
"replay_raw_sha256": {
"replay-depth-32-block-seed-2026073001.json": "5cea76d68a642c2b4b4f189a813b2e4af9361a939ddd31a68dfdb0ab090b4aad"
},
"replay_target_bytes": 65536000,
"runner_sha256": "04ae69e10c58972c9193c2d31c7e09d924a0d0e834107afa4ba128c64ac5800f",
"schema_version": 1,
"smoke_exact": {
"depth-16-baseline": {
"compare_sha256": "525cdcd9a79ec61096ecf64fedcbf93144d3168b434e2f8fd384c4374cbebd2a",
"exact": true,
"gradient_gate": {
"first_to_last_ratio_abs_delta": 0.0,
"max_abs_scale_ratio_error": 0.0,
"normalized_spectrum_max_abs_delta": 0.0,
"passed": true,
"per_block_scale_ratios": [
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0
],
"population_cv_abs_delta": 0.0,
"thresholds": {
"scale_ratio_abs": 1e-05,
"shape_abs": 1e-06
}
}
},
"depth-16-block": {
"compare_sha256": "e4330a0d878d10b474b9aeb0be58131f801ebe2e5dd9553d4713c3d16b6590cc",
"exact": true,
"gradient_gate": {
"first_to_last_ratio_abs_delta": 0.0,
"max_abs_scale_ratio_error": 0.0,
"normalized_spectrum_max_abs_delta": 0.0,
"passed": true,
"per_block_scale_ratios": [
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0
],
"population_cv_abs_delta": 0.0,
"thresholds": {
"scale_ratio_abs": 1e-05,
"shape_abs": 1e-06
}
}
},
"depth-32-baseline": {
"compare_sha256": "8ee0dc37e88742b6705968f5f2845a3a1df1250f5e0fd77be6b909236ea28e56",
"exact": true,
"gradient_gate": {
"first_to_last_ratio_abs_delta": 0.0,
"max_abs_scale_ratio_error": 0.0,
"normalized_spectrum_max_abs_delta": 0.0,
"passed": true,
"per_block_scale_ratios": [
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0
],
"population_cv_abs_delta": 0.0,
"thresholds": {
"scale_ratio_abs": 1e-05,
"shape_abs": 1e-06
}
}
},
"depth-32-block": {
"compare_sha256": "289073b1940d4747b7e049f8d767f9c22a49867d37403ff6ef7746a55bd1734c",
"exact": true,
"gradient_gate": {
"first_to_last_ratio_abs_delta": 0.0,
"max_abs_scale_ratio_error": 0.0,
"normalized_spectrum_max_abs_delta": 0.0,
"passed": true,
"per_block_scale_ratios": [
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0
],
"population_cv_abs_delta": 0.0,
"thresholds": {
"scale_ratio_abs": 1e-05,
"shape_abs": 1e-06
}
}
}
},
"smoke_raw_sha256": {
"smoke-a-depth-16-baseline.json": "fe0ff3ed21bf9b86d238d7a22e99b6e07ba99135b58d29eb6b205f2602e64556",
"smoke-a-depth-16-block.json": "faa06a2bd02135858d3a2503e1ecfb34187b4ee3b4ce3428ed473a55b15b6705",
"smoke-a-depth-32-baseline.json": "e327a6f5db87515c1b29b2808205e23b81228b08ce356b98804717e2e55eef1c",
"smoke-a-depth-32-block.json": "2d9a97c2d6272c928d5b5bc01febd142d4b2f2ad3080330e849f2731a8d080c2",
"smoke-b-depth-16-baseline.json": "fe0ff3ed21bf9b86d238d7a22e99b6e07ba99135b58d29eb6b205f2602e64556",
"smoke-b-depth-16-block.json": "faa06a2bd02135858d3a2503e1ecfb34187b4ee3b4ce3428ed473a55b15b6705",
"smoke-b-depth-32-baseline.json": "e327a6f5db87515c1b29b2808205e23b81228b08ce356b98804717e2e55eef1c",
"smoke-b-depth-32-block.json": "2d9a97c2d6272c928d5b5bc01febd142d4b2f2ad3080330e849f2731a8d080c2"
}
}
@@ -0,0 +1,700 @@
{
"architecture": "baseline",
"batch_size": 32,
"canonical_sha256_without_self": "0ae9ad1697ca15ec4c84270ad82b9fa98056f14f20fc7372c4b9886c29845415",
"depth": 16,
"diagnostics": [
{
"activation_grad_rms_by_block": [
0.0002865509013645351,
0.00024785567075014114,
0.00021551651298068464,
0.00019683866412378848,
0.0001823508064262569,
0.0001679604029050097,
0.00015584587526973337,
0.00014737028686795384,
0.00014130656199995428,
0.00013612695329356939,
0.00013129493163432926,
0.00012807638267986476,
0.00012550120300147682,
0.00012324050476308912,
0.00012108208466088399,
0.00011888353037647903
],
"activation_grad_statistics": {
"first_quartile_mean": 0.00023669043730478734,
"first_to_last_ratio": 1.9372775995887173,
"imbalance_abs_log_ratio": 0.6612836883477485,
"last_quartile_mean": 0.00012217683070048224,
"mean": 0.00016411257956860936,
"normalized": [
1.7460629899168627,
1.5102783187106137,
1.313223602646409,
1.1994124072706904,
1.1111324123086055,
1.0234462424910682,
0.9496278449793059,
0.8979828801383491,
0.8610343117596252,
0.8294729974472175,
0.8000296624393728,
0.7804178266926869,
0.7647262832098098,
0.7509509940495869,
0.7377989242455608,
0.7244023016942358
],
"population_cv": 0.29361872380036635
},
"activation_output_rms_by_block": [
0.02866268903017044,
0.029099803417921066,
0.02953805774450302,
0.030088091269135475,
0.030644793063402176,
0.03131929785013199,
0.03213750571012497,
0.03286394104361534,
0.033869802951812744,
0.03499744459986687,
0.03588006645441055,
0.037365153431892395,
0.03789033368229866,
0.039445169270038605,
0.04050002992153168,
0.04121527820825577
],
"activation_output_statistics": {
"first_quartile_mean": 0.0293471603654325,
"first_to_last_ratio": 0.7380574840395957,
"imbalance_abs_log_ratio": 0.3037335657624931,
"last_quartile_mean": 0.03976270277053118,
"mean": 0.034094841103069484,
"normalized": [
0.8406752488894867,
0.8534957922212247,
0.8663497699023965,
0.8824822259232266,
0.8988102619619861,
0.9185934539320196,
0.942591449919727,
0.9638977622528551,
0.9933996421752943,
1.0264733158329964,
1.0523605710888722,
1.0959180985456622,
1.1113216092650298,
1.1569248600043318,
1.1878638706395246,
1.2088420674453664
],
"population_cv": 0.11952987076282337
},
"bits_per_byte": 8.096463027059821,
"branch_output_rms_by_sublayer": [
0.0033055038657039404,
0.0037518907338380814,
0.0033751516602933407,
0.003911525942385197,
0.004229962360113859,
0.0038461871445178986,
0.004243654198944569,
0.0038366704247891903,
0.004372371360659599,
0.003851204412057996,
0.005078981164842844,
0.003818925702944398,
0.0054668826051056385,
0.0038749484810978174,
0.005823117680847645,
0.0037566579412668943,
0.0065947300754487514,
0.003876061411574483,
0.006763988174498081,
0.003795720636844635,
0.006863070651888847,
0.0038706199266016483,
0.007856340147554874,
0.003839300014078617,
0.00830968376249075,
0.003927029203623533,
0.009655521251261234,
0.0038513424806296825,
0.009004893712699413,
0.0039484030567109585,
0.008608299307525158,
0.003791053779423237
],
"capture": {
"all_gradients_finite": true,
"all_gradients_present": true,
"count": 16,
"dtypes": [
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32"
],
"position": "post-MLP Transformer-block output; Block AttnRes is captured before aggregation-partial reset",
"shape": [
16,
256,
192
],
"storage_unique": true
},
"core_parameter_grad_rms_by_block": [
0.00809059897248305,
0.007313812062277777,
0.007090710224412361,
0.0067033738367094624,
0.006535407867454384,
0.006785021795736036,
0.006295993569919972,
0.005509300326230647,
0.006231736048910802,
0.006047390980174007,
0.005629405359532731,
0.005784042737956458,
0.0059591736113607996,
0.0061983173180850636,
0.006571690039252182,
0.005060205437758869
],
"core_parameter_grad_statistics": {
"first_quartile_mean": 0.007299623773970663,
"first_to_last_ratio": 1.227374872012571,
"imbalance_abs_log_ratio": 0.2048776382299483,
"last_quartile_mean": 0.005947346601614228,
"mean": 0.006362886261765913,
"normalized": [
1.271529717747562,
1.1494488132258316,
1.1143858200043408,
1.05351149791715,
1.0271137340180259,
1.066343404015675,
0.9894870520870547,
0.8658492545019393,
0.9793882512652816,
0.9504163254515969,
0.8847251275509024,
0.909028151691524,
0.9365519618304368,
0.9741361173356609,
1.0328158902888074,
0.7952688810682109
],
"population_cv": 0.11384977604868386
},
"depth_weights": [],
"layer_input_rms_by_sublayer": [
0.02823694422841072,
0.028404271230101585,
0.02866268903017044,
0.02886480651795864,
0.029099803417921066,
0.029365191236138344,
0.02953805774450302,
0.029881296679377556,
0.030088091269135475,
0.030441828072071075,
0.030644793063402176,
0.03106885403394699,
0.03131929785013199,
0.031892918050289154,
0.03213750571012497,
0.03266792371869087,
0.03286394104361534,
0.03364725783467293,
0.033869802951812744,
0.03485054895281792,
0.03499744459986687,
0.03562505170702934,
0.03588006645441055,
0.03719272464513779,
0.037365153431892395,
0.03768136352300644,
0.03789033368229866,
0.03927159309387207,
0.039445169270038605,
0.04041972756385803,
0.04050002992153168,
0.04110744968056679
],
"loss_nats": 5.6120405197143555,
"loss_scale": 1.0,
"output_weights": null,
"step": 0,
"stream_state_rms_by_sublayer": [
0.028404271230101585,
0.02866268903017044,
0.02886480651795864,
0.029099803417921066,
0.029365191236138344,
0.02953805774450302,
0.029881296679377556,
0.030088091269135475,
0.030441828072071075,
0.030644793063402176,
0.03106885403394699,
0.03131929785013199,
0.031892918050289154,
0.03213750571012497,
0.03266792371869087,
0.03286394104361534,
0.03364725783467293,
0.033869802951812744,
0.03485054895281792,
0.03499744459986687,
0.03562505170702934,
0.03588006645441055,
0.03719272464513779,
0.037365153431892395,
0.03768136352300644,
0.03789033368229866,
0.03927159309387207,
0.039445169270038605,
0.04041972756385803,
0.04050002992153168,
0.04110744968056679,
0.04121527820825577
]
},
{
"activation_grad_rms_by_block": [
5.69471885683015e-05,
5.527245593839325e-05,
5.393734318204224e-05,
5.292302375892177e-05,
5.23102717124857e-05,
5.178990977583453e-05,
5.1564093155320734e-05,
5.129844430484809e-05,
5.121947833686136e-05,
5.116340616950765e-05,
5.109804988023825e-05,
5.104695082991384e-05,
5.09646451973822e-05,
5.100828275317326e-05,
5.106439857627265e-05,
5.11144389747642e-05
],
"activation_grad_statistics": {
"first_quartile_mean": 5.477000286191469e-05,
"first_to_last_ratio": 1.0731232762518041,
"imbalance_abs_log_ratio": 0.07057334637995329,
"last_quartile_mean": 5.103794137539808e-05,
"mean": 5.2170148819641327e-05,
"normalized": [
1.0915665348238701,
1.0594651767139285,
1.033873669184083,
1.0144311441756324,
1.002685882559561,
0.9927115591500164,
0.988383094968431,
0.9832911246274795,
0.9817775010367221,
0.9807027069519371,
0.9794499543578176,
0.9784704852268963,
0.9768928467805085,
0.9777292936141551,
0.9788049244944386,
0.9797641013345227
],
"population_cv": 0.032815404485150704
},
"activation_output_rms_by_block": [
0.028743742033839226,
0.02942933700978756,
0.030596865341067314,
0.03249936178326607,
0.03542664647102356,
0.03894684836268425,
0.04318666458129883,
0.04936420917510986,
0.05370106175541878,
0.06048284471035004,
0.0664696991443634,
0.07231792062520981,
0.07627613097429276,
0.07997097074985504,
0.08494995534420013,
0.08972473442554474
],
"activation_output_statistics": {
"first_quartile_mean": 0.03031732654199004,
"first_to_last_ratio": 0.3664591129538783,
"imbalance_abs_log_ratio": 1.0038683247140607,
"last_quartile_mean": 0.08273044787347317,
"mean": 0.05450543703045696,
"normalized": [
0.5273555006590916,
0.5399339701348108,
0.5613543713807885,
0.5962590808162097,
0.6499653686150171,
0.7145497859400932,
0.7923368187501488,
0.9056749539963465,
0.9852422929002714,
1.1096662646068314,
1.21950584686113,
1.326801958945847,
1.399422427007981,
1.4672108895332452,
1.5585592919240543,
1.6461611779281335
],
"population_cv": 0.3808757530834043
},
"bits_per_byte": 6.7509657011324835,
"branch_output_rms_by_sublayer": [
0.003436450148001313,
0.003791899885982275,
0.003964710980653763,
0.003938332665711641,
0.0056790695525705814,
0.003919641952961683,
0.006960175931453705,
0.0038887569680809975,
0.008646421134471893,
0.003928068559616804,
0.009785857982933521,
0.0038193853106349707,
0.010625048540532589,
0.003996006678789854,
0.012892307713627815,
0.0039014811627566814,
0.012757784686982632,
0.003861474571749568,
0.014726920053362846,
0.003962590359151363,
0.012820222415030003,
0.004551138263195753,
0.01401793584227562,
0.004161422606557608,
0.01284075528383255,
0.004107494372874498,
0.012963366694748402,
0.0034911984112113714,
0.014392351731657982,
0.004168955609202385,
0.013408888131380081,
0.003909120801836252
],
"capture": {
"all_gradients_finite": true,
"all_gradients_present": true,
"count": 16,
"dtypes": [
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32"
],
"position": "post-MLP Transformer-block output; Block AttnRes is captured before aggregation-partial reset",
"shape": [
16,
256,
192
],
"storage_unique": true
},
"core_parameter_grad_rms_by_block": [
0.0004507690637370199,
0.0004530669153600251,
0.0005174622333717066,
0.0005579942111201043,
0.0006494724560479704,
0.0007477190266084469,
0.0008433132451421535,
0.0008983201732199723,
0.001008715304343127,
0.0011143118429534733,
0.0010428854825857738,
0.0010916116075341876,
0.0011085001103672855,
0.0012404769719650786,
0.0013545740271648265,
0.0014243271893498689
],
"core_parameter_grad_statistics": {
"first_quartile_mean": 0.000494823105897214,
"first_to_last_ratio": 0.385986622193018,
"imbalance_abs_log_ratio": 0.9519525676489338,
"last_quartile_mean": 0.001281969574711765,
"mean": 0.0009064699913044387,
"normalized": [
0.49727963204644987,
0.4998145771025995,
0.5708542349284638,
0.6155683215912455,
0.7164853357289404,
0.824869034585972,
0.930326710461313,
0.99100927977468,
1.112795033503044,
1.2292870736404011,
1.1504909071341995,
1.2042446170372658,
1.2228756836970622,
1.3684699812069823,
1.494339625314625,
1.571289952246756
],
"population_cv": 0.338812252534038
},
"depth_weights": [],
"layer_input_rms_by_sublayer": [
0.028241552412509918,
0.028533460572361946,
0.028743742033839226,
0.029236938804388046,
0.02942933700978756,
0.03035305254161358,
0.030596865341067314,
0.032195623964071274,
0.03249936178326607,
0.03506496921181679,
0.03542664647102356,
0.03852483630180359,
0.03894684836268425,
0.04251723736524582,
0.04318666458129883,
0.04861506074666977,
0.04936420917510986,
0.052915990352630615,
0.05370106175541878,
0.0598941408097744,
0.06048284471035004,
0.06538087129592896,
0.0664696991443634,
0.07142146676778793,
0.07231792062520981,
0.07521561533212662,
0.07627613097429276,
0.07935076206922531,
0.07997097074985504,
0.08368266373872757,
0.08494995534420013,
0.08860929310321808
],
"loss_nats": 4.679412841796875,
"loss_scale": 1.0,
"output_weights": null,
"step": 20,
"stream_state_rms_by_sublayer": [
0.028533460572361946,
0.028743742033839226,
0.029236938804388046,
0.02942933700978756,
0.03035305254161358,
0.030596865341067314,
0.032195623964071274,
0.03249936178326607,
0.03506496921181679,
0.03542664647102356,
0.03852483630180359,
0.03894684836268425,
0.04251723736524582,
0.04318666458129883,
0.04861506074666977,
0.04936420917510986,
0.052915990352630615,
0.05370106175541878,
0.0598941408097744,
0.06048284471035004,
0.06538087129592896,
0.0664696991443634,
0.07142146676778793,
0.07231792062520981,
0.07521561533212662,
0.07627613097429276,
0.07935076206922531,
0.07997097074985504,
0.08368266373872757,
0.08494995534420013,
0.08860929310321808,
0.08972473442554474
]
}
],
"environment": {
"autocast": "cuda-bfloat16-forward-fp32-cross-entropy",
"compile": false,
"compute_capability": [
12,
0
],
"cublas_workspace_config": ":4096:8",
"cuda": "12.8",
"deterministic_algorithms": true,
"gpu": "NVIDIA GeForce RTX 5090",
"python": "3.10.14",
"torch": "2.11.0+cu128"
},
"evaluations": [
{
"bits_per_byte": 8.076786148009306,
"cross_entropy_nats": 5.5984015464782715,
"step": 0
},
{
"bits_per_byte": 6.752401670275862,
"cross_entropy_nats": 4.680408179759979,
"step": 20
}
],
"gradient_gate": {
"first_to_last_ratio_abs_delta": 0.0,
"max_abs_scale_ratio_error": 0.0,
"normalized_spectrum_max_abs_delta": 0.0,
"passed": true,
"per_block_scale_ratios": [
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0
],
"population_cv_abs_delta": 0.0,
"thresholds": {
"scale_ratio_abs": 1e-05,
"shape_abs": 1e-06
}
},
"hashes": {
"final_mixer_parameters": null,
"final_model_state": "d53dab0fa7c76215ab91de676b2aef3f9ef14cb0cc1819b7c4a887915bed97c0",
"final_optimizer_state": "a7bce44db1478ce53933758aa5033bbb1e0aa21296e9615c33146920f30a2057",
"final_public_parameters": "d53dab0fa7c76215ab91de676b2aef3f9ef14cb0cc1819b7c4a887915bed97c0",
"initial_mixer_parameters": null,
"initial_public_parameter_elements": 9541824,
"initial_public_parameter_structure": "e732db766f25e01f6ce1772cc182ced9de2c56c4a2384130117242b9444f0abe",
"initial_public_parameter_tensors": 115,
"initial_public_parameters": "af2724a1c34bcfd61d8a8bef402246430898c815e56b6e5c5257949a4eb0e7b1"
},
"manifest": {
"diagnostic_tensor_sha256": "21117e31db302b10d67b63f035665dc8f220b879d216ccd12b7d2ba86e7b1716",
"file_sha256": "080afb17d1e036c0bba0a799fdb8b98ee4ad652bd42dd1b3b67110dd2ede6371",
"formal_schedule_sha256": "5041e09b167f229248d2462324e8c254b8f5938975f135dcd8192b00a54a4f4e",
"input_gate_tensor_hashes": {
"0": "65136111a29a042e61a7909132560d95cd4bcf0f9b52f64d0fb2e57773856434",
"1": "d995676b4e7dec8f661cd8c2345fe7fc7a513c17f528c02fc946a441a6995a94",
"7999": "2345e7ac3decca2bdaebf13094fdc92fcabef42e3e461a2fefd6e8e7e76baccc"
},
"path": "experiments/k3/attnres_gradient/manifest.json",
"validation_tensor_sha256": "f459316f13078a163b47c133511bb7181e05170ab89516e196490113893ce338"
},
"model": {
"attnres_aggregation_groups": 8,
"context": 256,
"d_ff": 768,
"d_head": 32,
"d_model": 192,
"heads": 6,
"layers": 16,
"parameters": {
"core": 9541824,
"embedding": 98304,
"mixer": 0,
"total": 9541824
},
"sublayers": 32,
"sublayers_per_attnres_group": 4,
"transformer_blocks_per_attnres_group": 2,
"vocabulary": 256
},
"optimizer": {
"betas": [
0.9,
0.95
],
"epsilon": 1e-08,
"grad_clip": 1.0,
"min_lr": 3e-05,
"name": "AdamW",
"peak_lr": 0.0003,
"warmup_steps": 400,
"weight_decay_ndim_ge_2": 0.1
},
"protocol_id": "llm-atlas-k3-attnres-gradient-scale-v1",
"run_kind": "smoke",
"schema_version": 1,
"seed": 2026073001,
"steps": 20,
"target_bytes_seen": 163840,
"timing": {
"mean_ms": null,
"measured_steps": 0,
"median_ms": null,
"p95_ms": null,
"peak_allocated_bytes": 1648265728,
"peak_reserved_bytes": 3282042880,
"warmup_steps_excluded": 20
},
"training_history": [
{
"bits_per_byte": 8.088097790921855,
"learning_rate": 7.499999999999999e-07,
"loss_nats": 5.6062421798706055,
"step": 1,
"unclipped_grad_norm": 19.475919723510742
},
{
"bits_per_byte": 7.474302716882146,
"learning_rate": 7.499999999999999e-06,
"loss_nats": 5.180791854858398,
"step": 10,
"unclipped_grad_norm": 12.938376426696777
},
{
"bits_per_byte": 6.789615018295581,
"learning_rate": 1.4999999999999999e-05,
"loss_nats": 4.706202507019043,
"step": 20,
"unclipped_grad_norm": 4.482712745666504
}
]
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,700 @@
{
"architecture": "baseline",
"batch_size": 32,
"canonical_sha256_without_self": "0ae9ad1697ca15ec4c84270ad82b9fa98056f14f20fc7372c4b9886c29845415",
"depth": 16,
"diagnostics": [
{
"activation_grad_rms_by_block": [
0.0002865509013645351,
0.00024785567075014114,
0.00021551651298068464,
0.00019683866412378848,
0.0001823508064262569,
0.0001679604029050097,
0.00015584587526973337,
0.00014737028686795384,
0.00014130656199995428,
0.00013612695329356939,
0.00013129493163432926,
0.00012807638267986476,
0.00012550120300147682,
0.00012324050476308912,
0.00012108208466088399,
0.00011888353037647903
],
"activation_grad_statistics": {
"first_quartile_mean": 0.00023669043730478734,
"first_to_last_ratio": 1.9372775995887173,
"imbalance_abs_log_ratio": 0.6612836883477485,
"last_quartile_mean": 0.00012217683070048224,
"mean": 0.00016411257956860936,
"normalized": [
1.7460629899168627,
1.5102783187106137,
1.313223602646409,
1.1994124072706904,
1.1111324123086055,
1.0234462424910682,
0.9496278449793059,
0.8979828801383491,
0.8610343117596252,
0.8294729974472175,
0.8000296624393728,
0.7804178266926869,
0.7647262832098098,
0.7509509940495869,
0.7377989242455608,
0.7244023016942358
],
"population_cv": 0.29361872380036635
},
"activation_output_rms_by_block": [
0.02866268903017044,
0.029099803417921066,
0.02953805774450302,
0.030088091269135475,
0.030644793063402176,
0.03131929785013199,
0.03213750571012497,
0.03286394104361534,
0.033869802951812744,
0.03499744459986687,
0.03588006645441055,
0.037365153431892395,
0.03789033368229866,
0.039445169270038605,
0.04050002992153168,
0.04121527820825577
],
"activation_output_statistics": {
"first_quartile_mean": 0.0293471603654325,
"first_to_last_ratio": 0.7380574840395957,
"imbalance_abs_log_ratio": 0.3037335657624931,
"last_quartile_mean": 0.03976270277053118,
"mean": 0.034094841103069484,
"normalized": [
0.8406752488894867,
0.8534957922212247,
0.8663497699023965,
0.8824822259232266,
0.8988102619619861,
0.9185934539320196,
0.942591449919727,
0.9638977622528551,
0.9933996421752943,
1.0264733158329964,
1.0523605710888722,
1.0959180985456622,
1.1113216092650298,
1.1569248600043318,
1.1878638706395246,
1.2088420674453664
],
"population_cv": 0.11952987076282337
},
"bits_per_byte": 8.096463027059821,
"branch_output_rms_by_sublayer": [
0.0033055038657039404,
0.0037518907338380814,
0.0033751516602933407,
0.003911525942385197,
0.004229962360113859,
0.0038461871445178986,
0.004243654198944569,
0.0038366704247891903,
0.004372371360659599,
0.003851204412057996,
0.005078981164842844,
0.003818925702944398,
0.0054668826051056385,
0.0038749484810978174,
0.005823117680847645,
0.0037566579412668943,
0.0065947300754487514,
0.003876061411574483,
0.006763988174498081,
0.003795720636844635,
0.006863070651888847,
0.0038706199266016483,
0.007856340147554874,
0.003839300014078617,
0.00830968376249075,
0.003927029203623533,
0.009655521251261234,
0.0038513424806296825,
0.009004893712699413,
0.0039484030567109585,
0.008608299307525158,
0.003791053779423237
],
"capture": {
"all_gradients_finite": true,
"all_gradients_present": true,
"count": 16,
"dtypes": [
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32"
],
"position": "post-MLP Transformer-block output; Block AttnRes is captured before aggregation-partial reset",
"shape": [
16,
256,
192
],
"storage_unique": true
},
"core_parameter_grad_rms_by_block": [
0.00809059897248305,
0.007313812062277777,
0.007090710224412361,
0.0067033738367094624,
0.006535407867454384,
0.006785021795736036,
0.006295993569919972,
0.005509300326230647,
0.006231736048910802,
0.006047390980174007,
0.005629405359532731,
0.005784042737956458,
0.0059591736113607996,
0.0061983173180850636,
0.006571690039252182,
0.005060205437758869
],
"core_parameter_grad_statistics": {
"first_quartile_mean": 0.007299623773970663,
"first_to_last_ratio": 1.227374872012571,
"imbalance_abs_log_ratio": 0.2048776382299483,
"last_quartile_mean": 0.005947346601614228,
"mean": 0.006362886261765913,
"normalized": [
1.271529717747562,
1.1494488132258316,
1.1143858200043408,
1.05351149791715,
1.0271137340180259,
1.066343404015675,
0.9894870520870547,
0.8658492545019393,
0.9793882512652816,
0.9504163254515969,
0.8847251275509024,
0.909028151691524,
0.9365519618304368,
0.9741361173356609,
1.0328158902888074,
0.7952688810682109
],
"population_cv": 0.11384977604868386
},
"depth_weights": [],
"layer_input_rms_by_sublayer": [
0.02823694422841072,
0.028404271230101585,
0.02866268903017044,
0.02886480651795864,
0.029099803417921066,
0.029365191236138344,
0.02953805774450302,
0.029881296679377556,
0.030088091269135475,
0.030441828072071075,
0.030644793063402176,
0.03106885403394699,
0.03131929785013199,
0.031892918050289154,
0.03213750571012497,
0.03266792371869087,
0.03286394104361534,
0.03364725783467293,
0.033869802951812744,
0.03485054895281792,
0.03499744459986687,
0.03562505170702934,
0.03588006645441055,
0.03719272464513779,
0.037365153431892395,
0.03768136352300644,
0.03789033368229866,
0.03927159309387207,
0.039445169270038605,
0.04041972756385803,
0.04050002992153168,
0.04110744968056679
],
"loss_nats": 5.6120405197143555,
"loss_scale": 1.0,
"output_weights": null,
"step": 0,
"stream_state_rms_by_sublayer": [
0.028404271230101585,
0.02866268903017044,
0.02886480651795864,
0.029099803417921066,
0.029365191236138344,
0.02953805774450302,
0.029881296679377556,
0.030088091269135475,
0.030441828072071075,
0.030644793063402176,
0.03106885403394699,
0.03131929785013199,
0.031892918050289154,
0.03213750571012497,
0.03266792371869087,
0.03286394104361534,
0.03364725783467293,
0.033869802951812744,
0.03485054895281792,
0.03499744459986687,
0.03562505170702934,
0.03588006645441055,
0.03719272464513779,
0.037365153431892395,
0.03768136352300644,
0.03789033368229866,
0.03927159309387207,
0.039445169270038605,
0.04041972756385803,
0.04050002992153168,
0.04110744968056679,
0.04121527820825577
]
},
{
"activation_grad_rms_by_block": [
5.69471885683015e-05,
5.527245593839325e-05,
5.393734318204224e-05,
5.292302375892177e-05,
5.23102717124857e-05,
5.178990977583453e-05,
5.1564093155320734e-05,
5.129844430484809e-05,
5.121947833686136e-05,
5.116340616950765e-05,
5.109804988023825e-05,
5.104695082991384e-05,
5.09646451973822e-05,
5.100828275317326e-05,
5.106439857627265e-05,
5.11144389747642e-05
],
"activation_grad_statistics": {
"first_quartile_mean": 5.477000286191469e-05,
"first_to_last_ratio": 1.0731232762518041,
"imbalance_abs_log_ratio": 0.07057334637995329,
"last_quartile_mean": 5.103794137539808e-05,
"mean": 5.2170148819641327e-05,
"normalized": [
1.0915665348238701,
1.0594651767139285,
1.033873669184083,
1.0144311441756324,
1.002685882559561,
0.9927115591500164,
0.988383094968431,
0.9832911246274795,
0.9817775010367221,
0.9807027069519371,
0.9794499543578176,
0.9784704852268963,
0.9768928467805085,
0.9777292936141551,
0.9788049244944386,
0.9797641013345227
],
"population_cv": 0.032815404485150704
},
"activation_output_rms_by_block": [
0.028743742033839226,
0.02942933700978756,
0.030596865341067314,
0.03249936178326607,
0.03542664647102356,
0.03894684836268425,
0.04318666458129883,
0.04936420917510986,
0.05370106175541878,
0.06048284471035004,
0.0664696991443634,
0.07231792062520981,
0.07627613097429276,
0.07997097074985504,
0.08494995534420013,
0.08972473442554474
],
"activation_output_statistics": {
"first_quartile_mean": 0.03031732654199004,
"first_to_last_ratio": 0.3664591129538783,
"imbalance_abs_log_ratio": 1.0038683247140607,
"last_quartile_mean": 0.08273044787347317,
"mean": 0.05450543703045696,
"normalized": [
0.5273555006590916,
0.5399339701348108,
0.5613543713807885,
0.5962590808162097,
0.6499653686150171,
0.7145497859400932,
0.7923368187501488,
0.9056749539963465,
0.9852422929002714,
1.1096662646068314,
1.21950584686113,
1.326801958945847,
1.399422427007981,
1.4672108895332452,
1.5585592919240543,
1.6461611779281335
],
"population_cv": 0.3808757530834043
},
"bits_per_byte": 6.7509657011324835,
"branch_output_rms_by_sublayer": [
0.003436450148001313,
0.003791899885982275,
0.003964710980653763,
0.003938332665711641,
0.0056790695525705814,
0.003919641952961683,
0.006960175931453705,
0.0038887569680809975,
0.008646421134471893,
0.003928068559616804,
0.009785857982933521,
0.0038193853106349707,
0.010625048540532589,
0.003996006678789854,
0.012892307713627815,
0.0039014811627566814,
0.012757784686982632,
0.003861474571749568,
0.014726920053362846,
0.003962590359151363,
0.012820222415030003,
0.004551138263195753,
0.01401793584227562,
0.004161422606557608,
0.01284075528383255,
0.004107494372874498,
0.012963366694748402,
0.0034911984112113714,
0.014392351731657982,
0.004168955609202385,
0.013408888131380081,
0.003909120801836252
],
"capture": {
"all_gradients_finite": true,
"all_gradients_present": true,
"count": 16,
"dtypes": [
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32",
"torch.float32"
],
"position": "post-MLP Transformer-block output; Block AttnRes is captured before aggregation-partial reset",
"shape": [
16,
256,
192
],
"storage_unique": true
},
"core_parameter_grad_rms_by_block": [
0.0004507690637370199,
0.0004530669153600251,
0.0005174622333717066,
0.0005579942111201043,
0.0006494724560479704,
0.0007477190266084469,
0.0008433132451421535,
0.0008983201732199723,
0.001008715304343127,
0.0011143118429534733,
0.0010428854825857738,
0.0010916116075341876,
0.0011085001103672855,
0.0012404769719650786,
0.0013545740271648265,
0.0014243271893498689
],
"core_parameter_grad_statistics": {
"first_quartile_mean": 0.000494823105897214,
"first_to_last_ratio": 0.385986622193018,
"imbalance_abs_log_ratio": 0.9519525676489338,
"last_quartile_mean": 0.001281969574711765,
"mean": 0.0009064699913044387,
"normalized": [
0.49727963204644987,
0.4998145771025995,
0.5708542349284638,
0.6155683215912455,
0.7164853357289404,
0.824869034585972,
0.930326710461313,
0.99100927977468,
1.112795033503044,
1.2292870736404011,
1.1504909071341995,
1.2042446170372658,
1.2228756836970622,
1.3684699812069823,
1.494339625314625,
1.571289952246756
],
"population_cv": 0.338812252534038
},
"depth_weights": [],
"layer_input_rms_by_sublayer": [
0.028241552412509918,
0.028533460572361946,
0.028743742033839226,
0.029236938804388046,
0.02942933700978756,
0.03035305254161358,
0.030596865341067314,
0.032195623964071274,
0.03249936178326607,
0.03506496921181679,
0.03542664647102356,
0.03852483630180359,
0.03894684836268425,
0.04251723736524582,
0.04318666458129883,
0.04861506074666977,
0.04936420917510986,
0.052915990352630615,
0.05370106175541878,
0.0598941408097744,
0.06048284471035004,
0.06538087129592896,
0.0664696991443634,
0.07142146676778793,
0.07231792062520981,
0.07521561533212662,
0.07627613097429276,
0.07935076206922531,
0.07997097074985504,
0.08368266373872757,
0.08494995534420013,
0.08860929310321808
],
"loss_nats": 4.679412841796875,
"loss_scale": 1.0,
"output_weights": null,
"step": 20,
"stream_state_rms_by_sublayer": [
0.028533460572361946,
0.028743742033839226,
0.029236938804388046,
0.02942933700978756,
0.03035305254161358,
0.030596865341067314,
0.032195623964071274,
0.03249936178326607,
0.03506496921181679,
0.03542664647102356,
0.03852483630180359,
0.03894684836268425,
0.04251723736524582,
0.04318666458129883,
0.04861506074666977,
0.04936420917510986,
0.052915990352630615,
0.05370106175541878,
0.0598941408097744,
0.06048284471035004,
0.06538087129592896,
0.0664696991443634,
0.07142146676778793,
0.07231792062520981,
0.07521561533212662,
0.07627613097429276,
0.07935076206922531,
0.07997097074985504,
0.08368266373872757,
0.08494995534420013,
0.08860929310321808,
0.08972473442554474
]
}
],
"environment": {
"autocast": "cuda-bfloat16-forward-fp32-cross-entropy",
"compile": false,
"compute_capability": [
12,
0
],
"cublas_workspace_config": ":4096:8",
"cuda": "12.8",
"deterministic_algorithms": true,
"gpu": "NVIDIA GeForce RTX 5090",
"python": "3.10.14",
"torch": "2.11.0+cu128"
},
"evaluations": [
{
"bits_per_byte": 8.076786148009306,
"cross_entropy_nats": 5.5984015464782715,
"step": 0
},
{
"bits_per_byte": 6.752401670275862,
"cross_entropy_nats": 4.680408179759979,
"step": 20
}
],
"gradient_gate": {
"first_to_last_ratio_abs_delta": 0.0,
"max_abs_scale_ratio_error": 0.0,
"normalized_spectrum_max_abs_delta": 0.0,
"passed": true,
"per_block_scale_ratios": [
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0,
2.0
],
"population_cv_abs_delta": 0.0,
"thresholds": {
"scale_ratio_abs": 1e-05,
"shape_abs": 1e-06
}
},
"hashes": {
"final_mixer_parameters": null,
"final_model_state": "d53dab0fa7c76215ab91de676b2aef3f9ef14cb0cc1819b7c4a887915bed97c0",
"final_optimizer_state": "a7bce44db1478ce53933758aa5033bbb1e0aa21296e9615c33146920f30a2057",
"final_public_parameters": "d53dab0fa7c76215ab91de676b2aef3f9ef14cb0cc1819b7c4a887915bed97c0",
"initial_mixer_parameters": null,
"initial_public_parameter_elements": 9541824,
"initial_public_parameter_structure": "e732db766f25e01f6ce1772cc182ced9de2c56c4a2384130117242b9444f0abe",
"initial_public_parameter_tensors": 115,
"initial_public_parameters": "af2724a1c34bcfd61d8a8bef402246430898c815e56b6e5c5257949a4eb0e7b1"
},
"manifest": {
"diagnostic_tensor_sha256": "21117e31db302b10d67b63f035665dc8f220b879d216ccd12b7d2ba86e7b1716",
"file_sha256": "080afb17d1e036c0bba0a799fdb8b98ee4ad652bd42dd1b3b67110dd2ede6371",
"formal_schedule_sha256": "5041e09b167f229248d2462324e8c254b8f5938975f135dcd8192b00a54a4f4e",
"input_gate_tensor_hashes": {
"0": "65136111a29a042e61a7909132560d95cd4bcf0f9b52f64d0fb2e57773856434",
"1": "d995676b4e7dec8f661cd8c2345fe7fc7a513c17f528c02fc946a441a6995a94",
"7999": "2345e7ac3decca2bdaebf13094fdc92fcabef42e3e461a2fefd6e8e7e76baccc"
},
"path": "experiments/k3/attnres_gradient/manifest.json",
"validation_tensor_sha256": "f459316f13078a163b47c133511bb7181e05170ab89516e196490113893ce338"
},
"model": {
"attnres_aggregation_groups": 8,
"context": 256,
"d_ff": 768,
"d_head": 32,
"d_model": 192,
"heads": 6,
"layers": 16,
"parameters": {
"core": 9541824,
"embedding": 98304,
"mixer": 0,
"total": 9541824
},
"sublayers": 32,
"sublayers_per_attnres_group": 4,
"transformer_blocks_per_attnres_group": 2,
"vocabulary": 256
},
"optimizer": {
"betas": [
0.9,
0.95
],
"epsilon": 1e-08,
"grad_clip": 1.0,
"min_lr": 3e-05,
"name": "AdamW",
"peak_lr": 0.0003,
"warmup_steps": 400,
"weight_decay_ndim_ge_2": 0.1
},
"protocol_id": "llm-atlas-k3-attnres-gradient-scale-v1",
"run_kind": "smoke",
"schema_version": 1,
"seed": 2026073001,
"steps": 20,
"target_bytes_seen": 163840,
"timing": {
"mean_ms": null,
"measured_steps": 0,
"median_ms": null,
"p95_ms": null,
"peak_allocated_bytes": 1648265728,
"peak_reserved_bytes": 3282042880,
"warmup_steps_excluded": 20
},
"training_history": [
{
"bits_per_byte": 8.088097790921855,
"learning_rate": 7.499999999999999e-07,
"loss_nats": 5.6062421798706055,
"step": 1,
"unclipped_grad_norm": 19.475919723510742
},
{
"bits_per_byte": 7.474302716882146,
"learning_rate": 7.499999999999999e-06,
"loss_nats": 5.180791854858398,
"step": 10,
"unclipped_grad_norm": 12.938376426696777
},
{
"bits_per_byte": 6.789615018295581,
"learning_rate": 1.4999999999999999e-05,
"loss_nats": 4.706202507019043,
"step": 20,
"unclipped_grad_norm": 4.482712745666504
}
]
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+827
View File
@@ -0,0 +1,827 @@
#!/usr/bin/env python3
"""Run one preregistered AttnRes activation-gradient/depth experiment cell."""
from __future__ import annotations
import argparse
import hashlib
import importlib.util
import json
import math
import os
import platform
import statistics
import sys
import time
from dataclasses import dataclass
from pathlib import Path
from typing import Any, Iterable
import numpy as np
import torch
import torch.nn as nn
import torch.nn.functional as F
PROTOCOL_ID = "llm-atlas-k3-attnres-gradient-scale-v1"
ARCHITECTURES = ("baseline", "block")
DEPTHS = (16, 32)
EXPECTED_SEEDS = (2026073001, 2026073002, 2026073003)
DIAGNOSTIC_STEPS = (0, 100, 500, 2000, 4000, 8000)
FORMAL_STEPS = 8000
SMOKE_STEPS = 20
CONTEXT = 256
VOCABULARY = 256
BLOCK_GROUPS = 8
PEAK_LR = 3e-4
MIN_LR = 3e-5
WARMUP_STEPS = 400
WEIGHT_DECAY = 0.1
BETAS = (0.9, 0.95)
ADAM_EPS = 1e-8
GRAD_CLIP = 1.0
def load_round04_module() -> Any:
path = Path(__file__).resolve().parents[1] / "attnres" / "train.py"
spec = importlib.util.spec_from_file_location("k3_attnres_round04_train", path)
if spec is None or spec.loader is None:
raise RuntimeError(f"cannot import Round 04 runner from {path}")
module = importlib.util.module_from_spec(spec)
sys.modules[spec.name] = module
spec.loader.exec_module(module)
return module
round04 = load_round04_module()
def configure_round04_globals(depth: int) -> None:
round04.PROTOCOL_ID = PROTOCOL_ID
round04.LAYERS = depth
round04.SUBLAYERS = depth * 2
round04.BLOCKS = BLOCK_GROUPS
round04.SUBLAYERS_PER_BLOCK = (depth * 2) // BLOCK_GROUPS
round04.WARMUP_STEPS = WARMUP_STEPS
round04.EVAL_STEPS = DIAGNOSTIC_STEPS
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser()
parser.add_argument("--architecture", choices=ARCHITECTURES, required=True)
parser.add_argument("--depth", type=int, choices=DEPTHS, required=True)
parser.add_argument("--seed", type=int, required=True)
parser.add_argument("--steps", type=int)
parser.add_argument("--batch-size", type=int, default=32)
parser.add_argument("--cache-dir", type=Path, required=True)
parser.add_argument("--manifest", type=Path, required=True)
parser.add_argument("--output", type=Path, required=True)
parser.add_argument("--validation-windows", type=int, default=64)
parser.add_argument("--diagnostic-windows", type=int, default=16)
parser.add_argument("--eval-batch-size", type=int, default=8)
parser.add_argument("--timing-warmup", type=int, default=20)
parser.add_argument(
"--run-kind", choices=("smoke", "formal", "replay"), default="formal"
)
args = parser.parse_args()
expected_steps = SMOKE_STEPS if args.run_kind == "smoke" else FORMAL_STEPS
if args.steps is None:
args.steps = expected_steps
if args.steps != expected_steps:
raise ValueError(
f"{args.run_kind} must run exactly {expected_steps} steps, got {args.steps}"
)
if args.batch_size != 32:
raise ValueError("the frozen protocol requires batch size 32")
if args.validation_windows != 64 or args.diagnostic_windows != 16:
raise ValueError("the frozen protocol requires 64 validation / 16 diagnostic windows")
return args
def configure_determinism(seed: int) -> None:
if os.environ.get("CUBLAS_WORKSPACE_CONFIG") != ":4096:8":
raise RuntimeError("CUBLAS_WORKSPACE_CONFIG must be :4096:8 before Python starts")
torch.manual_seed(seed)
torch.cuda.manual_seed_all(seed)
torch.use_deterministic_algorithms(True)
torch.backends.cudnn.benchmark = False
torch.backends.cudnn.deterministic = True
torch.backends.cuda.matmul.allow_tf32 = False
torch.backends.cudnn.allow_tf32 = False
torch.set_float32_matmul_precision("highest")
def canonical_sha256(value: Any) -> str:
payload = json.dumps(
value, ensure_ascii=False, sort_keys=True, separators=(",", ":")
).encode()
return hashlib.sha256(payload).hexdigest()
def file_sha256(path: Path) -> str:
digest = hashlib.sha256()
with path.open("rb") as handle:
for block in iter(lambda: handle.read(1024 * 1024), b""):
digest.update(block)
return digest.hexdigest()
def tensor_bytes(tensor: torch.Tensor) -> bytes:
value = tensor.detach().cpu().contiguous()
return (
f"{value.dtype}|{tuple(value.shape)}|".encode()
+ value.reshape(-1).view(torch.uint8).numpy().tobytes()
)
def named_state_hash(
model: nn.Module, *, include_mixers: bool | None
) -> str:
digest = hashlib.sha256()
for name, tensor in sorted(model.state_dict().items()):
is_mixer = name.startswith("mixers.") or name.startswith("output_mixer.")
if include_mixers is not None and is_mixer != include_mixers:
continue
digest.update(name.encode())
digest.update(b"\0")
digest.update(tensor_bytes(tensor))
return digest.hexdigest()
def state_structure_hash(
model: nn.Module, *, include_mixers: bool | None
) -> tuple[str, int, int]:
digest = hashlib.sha256()
tensor_count = 0
element_count = 0
for name, tensor in sorted(model.state_dict().items()):
is_mixer = name.startswith("mixers.") or name.startswith("output_mixer.")
if include_mixers is not None and is_mixer != include_mixers:
continue
digest.update(
f"{name}|{tuple(tensor.shape)}|{tensor.dtype}|{tensor.numel()}\n".encode()
)
tensor_count += 1
element_count += tensor.numel()
return digest.hexdigest(), tensor_count, element_count
def recursive_state_hash(value: Any) -> str:
digest = hashlib.sha256()
def visit(path: str, item: Any) -> None:
if torch.is_tensor(item):
digest.update(f"{path}|tensor|".encode())
digest.update(tensor_bytes(item))
elif isinstance(item, dict):
digest.update(f"{path}|dict|{len(item)}\n".encode())
for key in sorted(item, key=lambda candidate: str(candidate)):
visit(f"{path}/{key}", item[key])
elif isinstance(item, (list, tuple)):
digest.update(f"{path}|sequence|{len(item)}\n".encode())
for index, child in enumerate(item):
visit(f"{path}/{index}", child)
else:
digest.update(f"{path}|scalar|{repr(item)}\n".encode())
visit("root", value)
return digest.hexdigest()
@dataclass
class ActivationTrace:
block_outputs: list[torch.Tensor]
layer_input_rms: list[float]
branch_output_rms: list[float]
stream_state_rms: list[float]
depth_weights: list[dict[str, Any]]
output_weights: dict[str, Any] | None = None
def rms(value: torch.Tensor) -> float:
return value.float().square().mean().sqrt().detach().cpu().item()
class GradientLanguageModel(round04.ReducedLanguageModel):
"""Round 04 trunk with aligned post-MLP activation capture."""
def forward(
self, input_ids: torch.Tensor, capture: bool = False
) -> tuple[torch.Tensor, ActivationTrace | None]:
embedded = self.embed(input_ids)
trace = ActivationTrace([], [], [], [], []) if capture else None
if self.architecture == "baseline":
hidden = embedded
for block in self.blocks:
attention_input = hidden
attention_output = block.attention(block.attention_norm(attention_input))
hidden = hidden + attention_output
if trace is not None:
trace.layer_input_rms.append(rms(attention_input))
trace.branch_output_rms.append(rms(attention_output))
trace.stream_state_rms.append(rms(hidden))
mlp_input = hidden
mlp_output = block.mlp(block.mlp_norm(mlp_input))
hidden = hidden + mlp_output
if trace is not None:
hidden.retain_grad()
trace.block_outputs.append(hidden)
trace.layer_input_rms.append(rms(mlp_input))
trace.branch_output_rms.append(rms(mlp_output))
trace.stream_state_rms.append(rms(hidden))
else:
completed = [embedded]
partial: torch.Tensor | None = None
mixer_index = 0
for block in self.blocks:
for branch_index in range(2):
sources = completed + ([] if partial is None else [partial])
branch_input, weights = self.mixers[mixer_index](
sources, capture
)
mixer_index += 1
if branch_index == 0:
branch_output = block.attention(
block.attention_norm(branch_input)
)
else:
branch_output = block.mlp(block.mlp_norm(branch_input))
branch_for_residual = branch_output.float()
partial = (
branch_for_residual
if partial is None
else partial + branch_for_residual
)
if trace is not None:
trace.layer_input_rms.append(rms(branch_input))
trace.branch_output_rms.append(rms(branch_output))
trace.stream_state_rms.append(rms(partial))
trace.depth_weights.append(weights or {})
if branch_index == 1:
partial.retain_grad()
trace.block_outputs.append(partial)
if mixer_index % round04.SUBLAYERS_PER_BLOCK == 0:
completed.append(partial)
partial = None
if partial is not None or len(completed) != BLOCK_GROUPS + 1:
raise RuntimeError("Block AttnRes aggregation contract failed")
if self.output_mixer is None:
raise RuntimeError("Block AttnRes output mixer missing")
hidden, output_weights = self.output_mixer(completed, capture)
if trace is not None:
trace.output_weights = output_weights
normalized = self.final_norm(hidden)
logits = F.linear(normalized, self.token_embedding.weight)
return logits, trace
def cross_entropy(logits: torch.Tensor, targets: torch.Tensor) -> torch.Tensor:
return F.cross_entropy(
logits.float().reshape(-1, VOCABULARY), targets.reshape(-1)
)
def learning_rate(step: int, total_steps: int) -> float:
if step <= WARMUP_STEPS:
return PEAK_LR * step / WARMUP_STEPS
progress = (step - WARMUP_STEPS) / max(1, total_steps - WARMUP_STEPS)
cosine = 0.5 * (1 + math.cos(math.pi * progress))
return MIN_LR + (PEAK_LR - MIN_LR) * cosine
@torch.no_grad()
def evaluate(
model: GradientLanguageModel,
corpus: Any,
window_count: int,
eval_batch_size: int,
) -> dict[str, float]:
model.eval()
loss_sum = 0.0
target_count = 0
for begin in range(0, window_count, eval_batch_size):
end = min(begin + eval_batch_size, window_count)
inputs, targets = corpus.fixed_batch(
corpus.validation_starts, begin, end
)
with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
logits, _ = model(inputs)
loss = F.cross_entropy(
logits.float().reshape(-1, VOCABULARY),
targets.reshape(-1),
reduction="sum",
)
loss_sum += loss.detach().cpu().item()
target_count += targets.numel()
nats = loss_sum / target_count
return {"cross_entropy_nats": nats, "bits_per_byte": nats / math.log(2)}
def mean(values: Iterable[float]) -> float:
return statistics.fmean(values)
def depth_statistics(values: list[float]) -> dict[str, Any]:
average = mean(values)
variance = mean((value - average) ** 2 for value in values)
quartile = len(values) // 4
first = mean(values[:quartile])
last = mean(values[-quartile:])
ratio = first / last
return {
"mean": average,
"population_cv": math.sqrt(variance) / average,
"normalized": [value / average for value in values],
"first_quartile_mean": first,
"last_quartile_mean": last,
"first_to_last_ratio": ratio,
"imbalance_abs_log_ratio": abs(math.log(ratio)),
}
def core_parameter_gradient_rms(model: GradientLanguageModel) -> list[float]:
values = []
for block in model.blocks:
sum_square = 0.0
count = 0
for parameter in block.parameters():
if parameter.grad is None:
raise RuntimeError("missing core parameter gradient")
gradient = parameter.grad.detach().float()
if not torch.isfinite(gradient).all():
raise RuntimeError("non-finite core parameter gradient")
sum_square += gradient.square().sum().detach().cpu().item()
count += gradient.numel()
values.append(math.sqrt(sum_square / count))
return values
def activation_storage_unique(outputs: list[torch.Tensor]) -> bool:
pointers = [output.untyped_storage().data_ptr() for output in outputs]
return len(pointers) == len(set(pointers))
def diagnostic(
model: GradientLanguageModel,
corpus: Any,
window_count: int,
*,
loss_scale: float = 1.0,
) -> dict[str, Any]:
model.eval()
model.zero_grad(set_to_none=True)
inputs, targets = corpus.fixed_batch(
corpus.diagnostic_starts, 0, window_count
)
with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
logits, trace = model(inputs, capture=True)
unscaled_loss = cross_entropy(logits, targets)
loss = unscaled_loss * loss_scale
if trace is None or len(trace.block_outputs) != len(model.blocks):
raise RuntimeError("aligned activation capture count mismatch")
expected_shape = (window_count, CONTEXT, round04.D_MODEL)
if any(tuple(output.shape) != expected_shape for output in trace.block_outputs):
raise RuntimeError("aligned activation capture shape mismatch")
if any(output.dtype != torch.float32 for output in trace.block_outputs):
raise RuntimeError("aligned activation capture must use FP32 residual state")
if not activation_storage_unique(trace.block_outputs):
raise RuntimeError("captured block outputs alias storage")
loss.backward()
activation_grad_rms = []
activation_output_rms = []
activation_dtypes = []
for output in trace.block_outputs:
if output.grad is None:
raise RuntimeError("captured activation gradient is None")
gradient = output.grad.detach().float()
if not torch.isfinite(gradient).all():
raise RuntimeError("captured activation gradient is non-finite")
activation_grad_rms.append(
gradient.square().mean().sqrt().detach().cpu().item()
)
activation_output_rms.append(rms(output))
activation_dtypes.append(str(output.dtype))
parameter_grad_rms = core_parameter_gradient_rms(model)
return {
"loss_nats": unscaled_loss.detach().cpu().item(),
"bits_per_byte": unscaled_loss.detach().cpu().item() / math.log(2),
"loss_scale": loss_scale,
"capture": {
"count": len(trace.block_outputs),
"shape": list(expected_shape),
"dtypes": activation_dtypes,
"all_gradients_finite": True,
"all_gradients_present": True,
"storage_unique": True,
"position": (
"post-MLP Transformer-block output; Block AttnRes is captured "
"before aggregation-partial reset"
),
},
"activation_grad_rms_by_block": activation_grad_rms,
"activation_grad_statistics": depth_statistics(activation_grad_rms),
"activation_output_rms_by_block": activation_output_rms,
"activation_output_statistics": depth_statistics(activation_output_rms),
"core_parameter_grad_rms_by_block": parameter_grad_rms,
"core_parameter_grad_statistics": depth_statistics(parameter_grad_rms),
"layer_input_rms_by_sublayer": trace.layer_input_rms,
"branch_output_rms_by_sublayer": trace.branch_output_rms,
"stream_state_rms_by_sublayer": trace.stream_state_rms,
"depth_weights": trace.depth_weights,
"output_weights": trace.output_weights,
}
def loss_scale_gate(
model: GradientLanguageModel, corpus: Any, window_count: int
) -> tuple[dict[str, Any], dict[str, Any]]:
base = diagnostic(model, corpus, window_count, loss_scale=1.0)
doubled = diagnostic(model, corpus, window_count, loss_scale=2.0)
base_values = base["activation_grad_rms_by_block"]
doubled_values = doubled["activation_grad_rms_by_block"]
ratios = [
doubled_value / base_value
for base_value, doubled_value in zip(base_values, doubled_values)
]
base_stats = base["activation_grad_statistics"]
doubled_stats = doubled["activation_grad_statistics"]
cv_delta = abs(
doubled_stats["population_cv"] - base_stats["population_cv"]
)
ratio_delta = abs(
doubled_stats["first_to_last_ratio"]
- base_stats["first_to_last_ratio"]
)
normalized_max_delta = max(
abs(left - right)
for left, right in zip(
base_stats["normalized"], doubled_stats["normalized"]
)
)
passed = (
all(abs(ratio - 2.0) <= 1e-5 for ratio in ratios)
and cv_delta <= 1e-6
and ratio_delta <= 1e-6
and normalized_max_delta <= 1e-6
)
gate = {
"passed": passed,
"per_block_scale_ratios": ratios,
"max_abs_scale_ratio_error": max(abs(ratio - 2.0) for ratio in ratios),
"population_cv_abs_delta": cv_delta,
"first_to_last_ratio_abs_delta": ratio_delta,
"normalized_spectrum_max_abs_delta": normalized_max_delta,
"thresholds": {
"scale_ratio_abs": 1e-5,
"shape_abs": 1e-6,
},
}
if not passed:
raise RuntimeError(f"loss-scale diagnostic gate failed: {gate}")
return base, gate
def percentile(values: list[float], quantile: float) -> float:
return float(np.quantile(np.asarray(values, dtype=np.float64), quantile))
def parameter_inventory(model: GradientLanguageModel) -> dict[str, int]:
total = sum(parameter.numel() for parameter in model.parameters())
mixer = sum(
parameter.numel()
for name, parameter in model.named_parameters()
if name.startswith("mixers.") or name.startswith("output_mixer.")
)
return {
"total": total,
"core": total - mixer,
"mixer": mixer,
"embedding": (
model.token_embedding.weight.numel()
+ model.position_embedding.weight.numel()
),
}
def model_input_gate_hashes(
corpus: Any, manifest: dict[str, Any], seed: int, batch_size: int
) -> dict[str, str]:
values: dict[str, str] = {}
for step in manifest["windows"]["gate_steps"]:
raw_digest = hashlib.sha256()
for row in range(batch_size):
start = round04.window_start(seed, step, row, len(corpus.train))
raw_digest.update(
np.asarray(
corpus.train[start : start + CONTEXT + 1], dtype=np.uint8
).tobytes()
)
expected_raw_hash = manifest["windows"][
"gate_training_tensor_sha256"
][str(seed)][str(step)]
if raw_digest.hexdigest() != expected_raw_hash:
raise RuntimeError(f"manifest gate tensor mismatch at step {step}")
inputs, targets = corpus.training_batch(seed, step, batch_size)
digest = hashlib.sha256()
digest.update(tensor_bytes(inputs))
digest.update(tensor_bytes(targets))
values[str(step)] = digest.hexdigest()
return values
def main() -> None:
args = parse_args()
if not torch.cuda.is_available():
raise RuntimeError("CUDA is required by the frozen protocol")
if args.seed not in EXPECTED_SEEDS:
raise ValueError(f"seed is not preregistered: {args.seed}")
configure_round04_globals(args.depth)
configure_determinism(args.seed)
device = torch.device("cuda")
manifest = json.loads(args.manifest.read_text())
if manifest["protocol_id"] != PROTOCOL_ID:
raise ValueError("manifest protocol mismatch")
if manifest["windows"]["formal_steps"] != FORMAL_STEPS:
raise ValueError("manifest formal-step mismatch")
corpus = round04.ByteCorpus(args.cache_dir, manifest, device)
model = GradientLanguageModel(args.architecture).to(device)
public_structure_hash, public_tensors, public_elements = state_structure_hash(
model, include_mixers=False
)
initial_public_hash = named_state_hash(model, include_mixers=False)
initial_mixer_hash = (
named_state_hash(model, include_mixers=True)
if args.architecture == "block"
else None
)
input_gate_hashes = model_input_gate_hashes(
corpus, manifest, args.seed, args.batch_size
)
decay_parameters: list[nn.Parameter] = []
no_decay_parameters: list[nn.Parameter] = []
for parameter in model.parameters():
(decay_parameters if parameter.ndim >= 2 else no_decay_parameters).append(
parameter
)
optimizer = torch.optim.AdamW(
[
{"params": decay_parameters, "weight_decay": WEIGHT_DECAY},
{"params": no_decay_parameters, "weight_decay": 0.0},
],
lr=PEAK_LR,
betas=BETAS,
eps=ADAM_EPS,
)
evaluation_steps = sorted(
set(step for step in DIAGNOSTIC_STEPS if step <= args.steps)
| {0, args.steps}
)
evaluations = [
{
"step": 0,
**evaluate(
model, corpus, args.validation_windows, args.eval_batch_size
),
}
]
if args.run_kind == "smoke":
initial_diagnostic, gradient_gate = loss_scale_gate(
model, corpus, args.diagnostic_windows
)
else:
initial_diagnostic = diagnostic(
model, corpus, args.diagnostic_windows
)
gradient_gate = None
diagnostics = [{"step": 0, **initial_diagnostic}]
print(
json.dumps(
{
"event": "diagnostic",
"step": 0,
"architecture": args.architecture,
"depth": args.depth,
"validation_bpc": evaluations[0]["bits_per_byte"],
"activation_gradient_cv": initial_diagnostic[
"activation_grad_statistics"
]["population_cv"],
},
sort_keys=True,
),
flush=True,
)
model.zero_grad(set_to_none=True)
training_history: list[dict[str, float | int]] = []
step_times: list[float] = []
model.train()
for step in range(1, args.steps + 1):
lr = learning_rate(step, args.steps)
for group in optimizer.param_groups:
group["lr"] = lr
inputs, targets = corpus.training_batch(args.seed, step, args.batch_size)
optimizer.zero_grad(set_to_none=True)
torch.cuda.synchronize()
started = time.perf_counter()
with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
logits, _ = model(inputs)
loss = cross_entropy(logits, targets)
if not torch.isfinite(loss):
raise RuntimeError(f"non-finite loss at step {step}: {loss}")
loss.backward()
unclipped_norm = torch.nn.utils.clip_grad_norm_(
model.parameters(), GRAD_CLIP
)
optimizer.step()
torch.cuda.synchronize()
elapsed_ms = (time.perf_counter() - started) * 1000
if step == args.timing_warmup:
torch.cuda.reset_peak_memory_stats()
elif step > args.timing_warmup:
step_times.append(elapsed_ms)
if step == 1 or step % 10 == 0 or step == args.steps:
training_history.append(
{
"step": step,
"loss_nats": loss.detach().cpu().item(),
"bits_per_byte": loss.detach().cpu().item() / math.log(2),
"learning_rate": lr,
"unclipped_grad_norm": float(unclipped_norm.detach().cpu()),
}
)
if step in evaluation_steps and step != 0:
evaluations.append(
{
"step": step,
**evaluate(
model,
corpus,
args.validation_windows,
args.eval_batch_size,
),
}
)
diagnostics.append(
{
"step": step,
**diagnostic(
model, corpus, args.diagnostic_windows
),
}
)
print(
json.dumps(
{
"event": "diagnostic",
"step": step,
"architecture": args.architecture,
"depth": args.depth,
"validation_bpc": evaluations[-1]["bits_per_byte"],
"activation_gradient_cv": diagnostics[-1][
"activation_grad_statistics"
]["population_cv"],
},
sort_keys=True,
),
flush=True,
)
model.zero_grad(set_to_none=True)
model.train()
training_peak_allocated = torch.cuda.max_memory_allocated()
training_peak_reserved = torch.cuda.max_memory_reserved()
final_public_hash = named_state_hash(model, include_mixers=False)
final_mixer_hash = (
named_state_hash(model, include_mixers=True)
if args.architecture == "block"
else None
)
final_full_hash = named_state_hash(model, include_mixers=None)
optimizer_hash = recursive_state_hash(optimizer.state_dict())
timing = {
"warmup_steps_excluded": args.timing_warmup,
"measured_steps": len(step_times),
"mean_ms": mean(step_times) if step_times else None,
"median_ms": statistics.median(step_times) if step_times else None,
"p95_ms": percentile(step_times, 0.95) if step_times else None,
"peak_allocated_bytes": training_peak_allocated,
"peak_reserved_bytes": training_peak_reserved,
}
result = {
"schema_version": 1,
"protocol_id": PROTOCOL_ID,
"run_kind": args.run_kind,
"architecture": args.architecture,
"depth": args.depth,
"seed": args.seed,
"steps": args.steps,
"batch_size": args.batch_size,
"target_bytes_seen": args.steps * args.batch_size * CONTEXT,
"manifest": {
"path": str(args.manifest),
"file_sha256": file_sha256(args.manifest),
"formal_schedule_sha256": manifest["windows"][
"formal_schedule_sha256"
],
"validation_tensor_sha256": manifest["windows"][
"validation_tensor_sha256"
],
"diagnostic_tensor_sha256": manifest["windows"][
"diagnostic_tensor_sha256"
],
"input_gate_tensor_hashes": input_gate_hashes,
},
"model": {
"layers": args.depth,
"sublayers": args.depth * 2,
"attnres_aggregation_groups": BLOCK_GROUPS,
"sublayers_per_attnres_group": args.depth * 2 // BLOCK_GROUPS,
"transformer_blocks_per_attnres_group": args.depth // BLOCK_GROUPS,
"d_model": round04.D_MODEL,
"heads": round04.HEADS,
"d_head": round04.D_HEAD,
"d_ff": round04.D_FF,
"context": CONTEXT,
"vocabulary": VOCABULARY,
"parameters": parameter_inventory(model),
},
"optimizer": {
"name": "AdamW",
"betas": list(BETAS),
"epsilon": ADAM_EPS,
"weight_decay_ndim_ge_2": WEIGHT_DECAY,
"peak_lr": PEAK_LR,
"min_lr": MIN_LR,
"warmup_steps": WARMUP_STEPS,
"grad_clip": GRAD_CLIP,
},
"hashes": {
"initial_public_parameter_structure": public_structure_hash,
"initial_public_parameter_tensors": public_tensors,
"initial_public_parameter_elements": public_elements,
"initial_public_parameters": initial_public_hash,
"initial_mixer_parameters": initial_mixer_hash,
"final_public_parameters": final_public_hash,
"final_mixer_parameters": final_mixer_hash,
"final_model_state": final_full_hash,
"final_optimizer_state": optimizer_hash,
},
"evaluations": evaluations,
"diagnostics": diagnostics,
"training_history": training_history,
"gradient_gate": gradient_gate,
"timing": timing,
"environment": {
"python": platform.python_version(),
"torch": torch.__version__,
"cuda": torch.version.cuda,
"gpu": torch.cuda.get_device_name(0),
"compute_capability": list(torch.cuda.get_device_capability(0)),
"cublas_workspace_config": os.environ["CUBLAS_WORKSPACE_CONFIG"],
"deterministic_algorithms": torch.are_deterministic_algorithms_enabled(),
"autocast": "cuda-bfloat16-forward-fp32-cross-entropy",
"compile": False,
},
}
result["canonical_sha256_without_self"] = canonical_sha256(result)
args.output.parent.mkdir(parents=True, exist_ok=True)
temporary = args.output.with_suffix(args.output.suffix + ".tmp")
temporary.write_text(
json.dumps(result, ensure_ascii=False, indent=2, sort_keys=True) + "\n"
)
os.replace(temporary, args.output)
print(
json.dumps(
{
"output": str(args.output),
"run_kind": args.run_kind,
"architecture": args.architecture,
"depth": args.depth,
"seed": args.seed,
"steps": args.steps,
"final_bpc": evaluations[-1]["bits_per_byte"],
"final_activation_gradient_cv": diagnostics[-1][
"activation_grad_statistics"
]["population_cv"],
"canonical_sha256": result["canonical_sha256_without_self"],
"timing": timing,
},
ensure_ascii=False,
indent=2,
)
)
if __name__ == "__main__":
main()
+2
View File
@@ -20,6 +20,7 @@
"build:data:deepseek-chat-task-bootstrap": "node scripts/build-deepseek-chat-task-bootstrap-crn-compact.mjs",
"check:data:deepseek-chat-task-bootstrap": "node scripts/check-deepseek-chat-task-bootstrap-crn-data.mjs",
"check:data:k3-attnres": "node scripts/check-k3-attnres-data.mjs",
"check:data:k3-attnres-gradient": "node scripts/check-k3-attnres-gradient-data.mjs",
"check:site": "node scripts/check-site.mjs",
"check:moe-browser": "node scripts/check-moe-browser.mjs",
"check:reasoning-browser": "node scripts/check-reasoning-browser.mjs",
@@ -40,6 +41,7 @@
"check:deepseek-cross-source-sampling-browser": "node scripts/check-deepseek-cross-source-sampling-browser.mjs",
"check:deepseek-task-bootstrap-browser": "node scripts/check-deepseek-task-bootstrap-browser.mjs",
"check:k3-attnres-browser": "node scripts/check-k3-attnres-browser.mjs",
"check:k3-attnres-gradient-browser": "node scripts/check-k3-attnres-gradient-browser.mjs",
"check:k3-browser": "node scripts/check-k3-browser.mjs"
},
"dependencies": {
@@ -0,0 +1,198 @@
# Kimi K3 第五轮前置审计:Attention Residuals 的“梯度更均匀”到底指什么
> 审计日期:2026-07-30(Asia/Shanghai)
> 官方仓库:`MoonshotAI/Attention-Residuals@85e22310fe5ee860b4a023de312d791de8a5a5e6`
> 官方 PDF SHA-256:`e5831b0db1347606453b5176b0142115a18887b6a9c2e1d05a266d4805a26b2f`
> 结论性质:一手工件审计,不是新的实验结果
## 0. 先说结论
K3 Round 04 的反结果——Block AttnRes 的**核心参数梯度 RMS 跨层 CV 更高**——不能直接
反驳 Attention Residuals 论文 Figure 5(c) 所说的“梯度分布更均匀”,因为两边很可能测的
不是同一个对象:
```text
Round 04:
每个 Transformer block 内所有核心参数梯度拼接后的 RMS
∇θL,θ = attention + MLP + 两个输入 norm 的参数
论文 Figure 5(c):
图题只写 “Each transformer block's gradient magnitude”
结合 Figure 5(b) 的 block output magnitude 和正文,最自然的操作化是
每个 block 输出 activation 的梯度 ∂L/∂h_l
```
但“最自然”不等于“官方已经明确定义”。论文和当前官方仓库都没有给出足以唯一重建
Figure 5(c) 的测量合同,也没有发布训练代码。因此,下一轮不会把自己的 activation-gradient
定义冒充成论文原始实现,而会把它命名为:
> **与 Figure 5 叙述对齐的一种公开、冻结、可复现的 operationalization**
这一区分很重要:参数梯度回答“这一层的权重此刻收到多大更新信号”,activation 梯度回答
“损失对这一深度的表征有多敏感”。二者相关,但不会因为链式法则而自动同方向。
---
## 1. 官方工件实际提供了什么
官方仓库在固定 revision 下只包含:
- `README.md`;
- `Attention_Residuals.pdf`;
- 论文图片资产;
- citation 与外部入口。
仓库**不包含**:
- 模型或 residual mixer 的可执行实现;
- Figure 5 的统计脚本;
- 训练配置、日志或 checkpoint;
- gradient hook、norm 和 reduction 定义;
- 用于复画 Figure 5 的原始数组。
因此,本审计能固定论文的文字、公式、图与模型尺度,不能从官方代码恢复一个不存在的
隐藏测量合同。
## 2. Figure 5 能确认的事实
官方 `training_dynamics.png` 和 PDF Figure 5 有三个并列面板:
| 面板 | 图题 | 横轴 |
|---|---|---|
| (a) | Validation Loss | training step |
| (b) | Output Magnitude | Layer / Transformer Block Index |
| (c) | Gradient Magnitude ×10⁻⁵ | Layer / Transformer Block Index |
Figure 5 caption 对 (b) 与 (c) 的完整对象描述分别是:
```text
Each transformer block's output magnitude at the end of training.
Each transformer block's gradient magnitude.
```
正文明确表达了两个方向:
1. Baseline 的 hidden-state magnitude 随深度单调增长;Block AttnRes 把增长约束在局部
residual block 内,形成周期性的深度图案。
2. Baseline 的早期层梯度“不成比例地大”;Block AttnRes 的可学习 softmax 权重产生了
“明显更均匀”的梯度分布。
图中最终大模型约有 27 个 Transformer blocks。论文的最终模型描述也是 27 个 Transformer
blocks / 54 个 residual layers;Block AttnRes 每 6 个 residual layers 聚合一次,共 9 个
聚合块,外加 embedding 形成 10 个跨块来源。
## 3. Figure 5 不能确认的事项
下面每一项都会改变曲线,却没有在论文或官方仓库中被唯一指定:
| 未定义项 | 至少两种合理解释 |
|---|---|
| 梯度对象 | block 输出 activation 梯度;block 参数梯度;分支输出梯度 |
| block 输出位置 | attention+MLP 后;只在 MLP 后;进入下一层 norm 前;聚合块边界后 |
| norm | L2 norm;RMS;mean absolute value;每 token norm 后再平均 |
| reduction | batch/token/channel 联合;先按 token 再按 batch;只取末 token |
| loss | token mean;sample mean;未归一化 sum;带或不带 mask |
| 采样 | 一个 batch;多 batch 平均;训练流中的 moving average |
| 时间点 | “训练结束”单点;末段平均;某个 checkpoint |
| 数值阶段 | AMP 缩放前/后;gradient clipping 前/后;BF16 或 FP32 |
| 运行模式 | train 或 eval;dropout 是否开启 |
| 归一化 | 绝对值;再除全层均值;再除 Baseline |
Figure 5(c) 的纵轴是绝对 magnitude 标度,并不等价于 CV。只报告 CV 还会丢失两个信息:
- 全部层梯度是否一起缩小或放大;
- 不均匀来自“早层系统性偏大”,还是某一个中间/末端尖峰。
所以 Round 05 必须同时公开绝对曲线、按层均值归一化曲线、CV 与前后深度分位比。
## 4. 为什么 activation gradient 是合理推断,但仍只是推断
把 Figure 5(c) 操作化为 `∂L/∂h_l` 有三条证据:
1. 它与 Figure 5(b) 的 “transformer block output magnitude” 在横轴和叙述上成对;
2. “早期层的梯度”在表示传播语境中通常可由对 block output 保留梯度直接比较;
3. 参数张量的大小和类型在 attention 与 MLP 间差异很大,若把参数拼接,论文通常需要说明
聚合口径,否则 “each transformer block” 不是天然的单一标量。
但也有无法排除的替代解释:
- 论文作者可能测 block 参数梯度;
- 可能测 residual branch output 而非完整 block output;
- 可能先对每个 token 做 L2 norm,再跨 token 平均;
- 可能在内部训练系统中有未公开的统一 telemetry 定义。
因此,网站和审计只说“与论文叙述对齐的公开定义”,不说“论文就是这样算的”,也不把
数值和 Figure 5 纵轴直接对齐。
## 5. Round 04 与 Round 05 的对象对照
| 维度 | Round 04 已测对象 | Round 05 主对象 |
|---|---|---|
| 数学对象 | `∇θ_l L` | `∂L/∂h_l` |
| `l` 的单位 | Transformer block | Transformer block |
| 张量内容 | block 的核心参数 | block 的 post-MLP output activation |
| 聚合 | 参数元素联合 RMS | batch×time×channel 联合 RMS |
| 是否含 AttnRes 参数 | 否 | 不适用;梯度穿过 mixer |
| 时间 | final diagnostic batch | 全部预注册 diagnostic steps |
| 目的 | 权重更新信号是否均匀 | 表征深度的反向信号是否均匀 |
Round 04 的参数梯度结果不会被改名、删去或用新指标覆盖。它仍是一个有效反结果,只是不能
代表论文未定义清楚的 Figure 5(c)。
## 6. 两种结构怎样取得真正对齐的 16 / 32 个位置
Grok Headless 被用作一次对抗式方法审阅,不作为事实来源。它正确指出了 non-leaf tensor、
alias、AMP、clip 时点和只看 CV 的风险;但它也提出了一个不适用于本实现的担忧:
“Block AttnRes 只有约 8 个 block 输出,无法与 Baseline 的 16 / 32 层对齐”。
这里要区分两种 block:
```text
Transformer block:
attention + MLP;depth=16 时始终有 16 个,depth=32 时始终有 32 个
AttnRes aggregation group:
把若干 residual sublayers 的 partial sum 保存为一个跨组 source;
两种深度都约为 8 组
```
Round 05 在**每个 Transformer block 的 MLP 分支完成后**取 `h_l`。所以 Baseline 与 Block
都有完全相同的 `l=1..depth`:
| 架构 | Round 05 的 `h_l` | shape |
|---|---|---|
| Baseline | 第 `l` 个 attention residual 与 MLP residual 都完成后的 hidden state | `[B,T,C]` |
| Block | 第 `l` 个 MLP branch 加入后、可能保存并清空 aggregation partial **之前**的 partial output | `[B,T,C]` |
Block 的 `h_l` 是局部 residual partial,而不是跨组 source 列表;这正对应论文 Figure 5(b)
所描述的“增长被限制在每个 block 内”的周期性图案。组边界前取值也避免把 reset 后的零张量
错误当成 Transformer block output。
## 7. 实验实现必须通过的梯度测量闸门
正式训练前,四个结构格(2 个深度 × 2 个 residual graph)都必须证明:
1. 所有 `h_l` 都是不同的、非别名的捕获对象,数量严格等于 Transformer depth;
2. `retain_grad()` 或 hook 后所有梯度非 `None`、finite,shape 与 activation 完全一致;
3. diagnostic loss 使用固定输入、固定 token-mean CE、`eval()`、无 optimizer step;
4. backward 发生在 parameter gradient clip 之前,且不经过 `GradScaler`;
5. 将同一 diagnostic loss 精确乘 2 后,每层 activation-gradient RMS 也乘 2;
6. 乘 2 前后的 CV、归一化曲线与深度分位比在数值容差内不变;
7. 同配置全新进程重复运行,冻结诊断字段 exact。
若任一项失败,正式 8,000-step grid 不得开始。
## 8. 本审计带来的研究决策
下一轮不再把一个宽度 192、深度 16、训练 2,000 step 的参数梯度 CV 与论文最终模型图强行
放在同一条结论线上,而是:
- 增加 depth 32;
- 把训练预算扩为 8,000 step;
- Baseline / Block 使用相同的 Transformer block index;
- 在 6 个固定时点测 post-MLP activation gradient;
- 绝对标度、归一化形状、CV、前后四分位失衡一起公开;
- 参数梯度作为次要指标保留;
- 预注册“支持 / 混合 / 不支持”规则,并公开全部反结果。
精确协议见 `research/K3_ATTNRES_GRADIENT_SCALE_PROTOCOL.md`。
+545
View File
@@ -0,0 +1,545 @@
# Kimi K3 第五轮:Attention Residuals 梯度定义与深度扩展实验审计
> 协议:`llm-atlas-k3-attnres-gradient-scale-v1`
> 前置定义审计:`research/K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md`
> 预注册协议:`research/K3_ATTNRES_GRADIENT_SCALE_PROTOCOL.md`
> 数据清单:`experiments/k3/attnres_gradient/manifest.json`
> 执行日期:2026-07-30
> 设备:NVIDIA GeForce RTX 5090;PyTorch `2.11.0+cu128`
## 0. 先说结论
这一轮本来想澄清一个看似矛盾的问题:
```text
论文 Figure 5:
Block AttnRes 的梯度沿深度“明显更均匀”
本站 Round 04:
Block AttnRes 的核心参数梯度 RMS 跨层 CV 反而更高
```
一手工件审计先确认:论文没有公开 Figure 5(c) 的确切 gradient tensor、norm、reduction、
diagnostic batch、AMP / clip 时点或统计代码。因此,Round 05 没有假装恢复作者的隐藏实现,
而是冻结一个与 Figure 5 的 output/gradient 并列叙述对齐、可复现的定义:
```text
h_l:
第 l 个 Transformer block 完成 attention + MLP 后的 FP32 residual output
m_l:
sqrt(mean((∂L / ∂h_l)² over batch × time × channel))
L:
固定 16 × 256 targets 的 token-mean cross entropy
```
结果不是一句“是”或“不是”,而是一个更有信息量的分解:
1. **Block 确实大幅缓解了早层整体偏大。**
Baseline 首四分位梯度平均是末四分位的 `3.01–3.90×`;Block 把它改到
`0.55–1.40×`。预注册的首尾失衡指标在 6 / 6 个 depth×seed 配对中都改善,
depth-16 平均改善 `61.0%`,depth-32 平均改善 `72.0%`。
2. **但 Block 没有让整条深度谱更平。**
它在中后段形成了局部尖峰,所以 population CV 在 6 / 6 个配对中都恶化:
depth-16 平均相对恶化 `10.3%`,depth-32 平均相对恶化 `60.0%`。
3. **Block 的绝对 activation-gradient 平均尺度更小。**
它只有 Baseline 的 `57.4%`(depth-16)和 `54.4%`(depth-32)。这意味着“首尾更接近”
不能自动解释为所有层都获得更强更新信号。
4. **参数梯度仍与 Round 04 同方向。**
核心参数梯度 CV 从 `0.416→0.683`(depth-16),从 `0.397→0.772`
(depth-32);Block 更不均匀。
5. **验证 BPC 在 6 / 6 配对中都更低,但不是同算力优势。**
平均改善 `0.00894 BPC`(depth-16)与 `0.00987 BPC`(depth-32);Block 实际 step
time 是 Baseline 的约 `2.55–2.60×`,peak allocated memory 约 `2.15–2.16×`。
按看结果前冻结的联合判据,CV 与首尾失衡必须同时改善才算 support;必须同时恶化才算
concern。这里二者方向相反,所以:
| depth | 预注册判定 |
|---:|---|
| 16 | **mixed / inconclusive at this depth** |
| 32 | **mixed / inconclusive at this depth** |
| 总判定 | **depth-dependent or inconclusive** |
最准确的中文总结是:
> 在这套公开 operationalization 中,Block AttnRes 把“早层系统性偏大”变成了“首尾更接近、
> 但中后段有局部尖峰”的另一种梯度分布。它修正了一类失衡,却没有降低全层离散度。
这不复现论文 Figure 5 的数值,也不反驳一个没有公开测量合同的隐藏实现。
---
## 1. 为什么要另开一轮,而不是改写 Round 04
Round 04 的梯度对象是:
```text
每个 Transformer block 的 attention、MLP 与两个 input norm
所有核心参数梯度拼接后的 RMS
```
这回答“该层权重收到多大更新信号”。Round 05 的对象是 `∂L/∂h_l`,回答“损失对该深度
表征有多敏感”。链式法则把两者联系起来,但不会保证跨层形状同方向。
因此:
- Round 04 参数梯度反结果继续有效;
- Round 05 不把它改名为 activation gradient;
- 两种对象在网站并排展示;
- 任何方向冲突都保留,而不是选择更像论文的一种。
官方工件边界见前置定义审计。固定 revision 为:
```text
MoonshotAI/Attention-Residuals
85e22310fe5ee860b4a023de312d791de8a5a5e6
Attention_Residuals.pdf SHA-256
e5831b0db1347606453b5176b0142115a18887b6a9c2e1d05a266d4805a26b2f
```
官方仓库没有可执行训练代码或 Figure 5 原始数组。
## 2. 实验规模与配对合同
正式网格:
```text
2 depths
× 2 residual graphs
× 3 seeds
× 8,000 steps
× 32 windows
× 256 target bytes
= 12 independent runs
= 786,432,000 formal target bytes
```
| 项 | depth-16 | depth-32 |
|---|---:|---:|
| Transformer blocks | 16 | 32 |
| residual sublayers | 32 | 64 |
| AttnRes aggregation groups | 8 | 8 |
| sublayers / group | 4 | 8 |
| Transformer blocks / group | 2 | 4 |
| width / heads / FFN | 192 / 6 / 768 | 192 / 6 / 768 |
每个 depth / seed 的 Baseline 与 Block:
- 使用逐 tensor exact 的公共主干初始化;
- 使用逐 step / row exact 的 byte windows;
- 使用相同 optimizer、LR schedule、batch、context 与 target-byte budget;
- 分支线性计算走 BF16 autocast;
- residual accumulator 与被测 `h_l` 都为 FP32;
- 每个结构在全新进程中从零训练。
二者**不匹配**:
- mixer 参数;
- mixer FLOPs;
- step wall time;
- activation memory。
所以 BPC 只能叫“同 token / step 预算对比”,不能叫“同算力优势”。
## 3. 数据与日程
| 对象 | 固定值 |
|---|---|
| dataset | `Salesforce/wikitext` |
| revision | `b08601e04326c79dfdd32d625aee71d232d685c3` |
| variant | `wikitext-2-raw-v1` |
| train bytes | 10,951,563 |
| train SHA-256 | `0ca7d3e7…e9b4` |
| validation bytes | 1,148,008 |
| validation SHA-256 | `a42356f6…e719` |
| formal schedule cells | 768,000 |
| schedule SHA-256 | `5041e09b…f4e` |
| validation tensor SHA-256 | `f459316f…338` |
| diagnostic tensor SHA-256 | `21117e31…716` |
固定诊断时点:
```text
0, 100, 500, 2,000, 4,000, 8,000
```
六个时点的全部 activation gradient、output RMS、parameter gradient、mixer 权重与验证 BPC
都进入 raw JSON;没有只挑“最好看”的 checkpoint。
## 4. 正式训练前的故障与修订
### 4.1 第一次 smoke 的标量序列化错误
首个 depth-16 Baseline 20-step smoke 已完成数值计算,但在写 JSON 前失败:
```text
RuntimeError:
self.dim() cannot be 0 to view Float as Byte
```
原因是 optimizer 的 step 是 0 维 tensor,hash helper 直接把它 `view(torch.uint8)`。修复为:
```text
tensor.reshape(-1).view(torch.uint8)
```
当时:
- 没有 formal 运行;
- 没有输出 JSON;
- 没有可供选择的正式结果。
### 4.2 smoke 发现被测 residual dtype 不一致
第一版 smoke 通过了 finite / loss×2 / replay 闸门,但检查 capture metadata 时发现:
```text
Baseline post-MLP h_l:FP32
Block aggregation partial:BF16
```
原因:
- Baseline 把 BF16 branch 加到 FP32 embedding/residual stream;
- Block 每组第一个 partial 直接引用 BF16 branch output。
这会把数值精度差异混进结构比较。正式训练前,协议与实现补充为:
```text
BF16 branch output → 显式转 FP32 → residual partial 累加
```
随后 4 个 depth×architecture 格的 smoke 全部从头重跑两次。正式输出是在这次修订之后才开始。
### 4.3 一次未启动模型的 zsh 调度错误
首个正式 Baseline 完成后,批处理脚本用 Bash 式标量切分处理 zsh 字符串,第一行就退出:
```text
argument --architecture: invalid choice: ''
```
runner 没有启动,也没有创建新结果文件。调度改为显式 `:` 分隔数组后继续。这是 orchestration
故障,不是模型运行失败,但仍在审计时间线中保留。
## 5. 运行前与复现闸门
### 5.1 activation-gradient 测量闸门
4 / 4 个 depth×architecture 格都通过:
- 捕获数量严格等于 16 / 32;
- shape 严格为 `[16,256,192]`;
- dtype 全部为 FP32;
- gradient 全部 present、finite;
- capture storage 全部不别名;
- diagnostic loss×2 后每层 gradient RMS 精确×2;
- CV、归一化谱、首尾比不变。
### 5.2 两次独立 smoke
每个格都在两个全新进程中训练 20 steps。排除 timing 后的冻结字段:
| 格 | compare SHA-256 |
|---|---|
| depth-16 Baseline | `525cdcd9…bd2a` |
| depth-16 Block | `e4330a0d…90cc` |
| depth-32 Baseline | `8ee0dc37…8e56` |
| depth-32 Block | `289073b1…34c` |
4 / 4 exact。
### 5.3 完整 formal replay
预注册格:
```text
depth-32 / Block / seed-2026073001
```
从初始化重新训练完整 8,000 steps,不加载 formal checkpoint。排除 `run_kind`、timing、
memory 与进程元数据后的全部冻结字段:
```text
formal compare SHA-256
46300a452840a9dc6a5180efe7949cf3d81cc4471ed4a942da407e9343064817
replay compare SHA-256
46300a452840a9dc6a5180efe7949cf3d81cc4471ed4a942da407e9343064817
```
最终状态:
| 对象 | formal | replay |
|---|---|---|
| model state | `3f0b97ec…2f59` | `3f0b97ec…2f59` |
| optimizer state | `ed03e6fb…4637` | `ed03e6fb…4637` |
字段级 exact。
## 6. 主结果:CV 与首尾失衡为什么方向相反
### 6.1 depth-16
| seed | Base CV | Block CV | 相对 CV 改善 | Base imbalance | Block imbalance | 相对 imbalance 改善 |
|---:|---:|---:|---:|---:|---:|---:|
| 2026073001 | 0.41719 | 0.46768 | −12.1% | 1.29458 | 0.38821 | +70.0% |
| 2026073002 | 0.43744 | 0.45175 | −3.3% | 1.35991 | 0.58958 | +56.6% |
| 2026073003 | 0.41927 | 0.48391 | −15.4% | 1.29826 | 0.56693 | +56.3% |
| 均值 | 0.42463 | 0.46778 | **−10.3%** | 1.31759 | 0.51491 | **+61.0%** |
这里“相对 CV 改善”为负,表示恶化。
原始首/末四分位比:
```text
Baseline:3.65×, 3.90×, 3.66×
Block: 0.68×, 0.55×, 0.57×
```
Block 不只是把早层优势降到 1;它在三个 seed 中都略微“过冲”,变成末四分位平均更大。
但 `abs(log(first/last))` 仍比 Baseline 更接近 0,所以失衡改善。
三 seed 平均 normalized activation-gradient 的最高点:
```text
Baseline:layer 3 = 1.50× mean;layer 4 = 1.42×
Block: layer 11 = 2.09× mean;layer 13 = 1.95×
```
Baseline 是宽而平滑的早层隆起;Block 是更局部的中后段尖峰。CV 对尖峰敏感,所以升高。
### 6.2 depth-32
| seed | Base CV | Block CV | 相对 CV 改善 | Base imbalance | Block imbalance | 相对 imbalance 改善 |
|---:|---:|---:|---:|---:|---:|---:|
| 2026073001 | 0.36109 | 0.64203 | −77.8% | 1.10030 | 0.20781 | +81.1% |
| 2026073002 | 0.37162 | 0.72997 | −96.4% | 1.14087 | 0.43810 | +61.6% |
| 2026073003 | 0.40320 | 0.42685 | −5.9% | 1.25169 | 0.33417 | +73.3% |
| 均值 | 0.37864 | 0.59962 | **−60.0%** | 1.16429 | 0.32669 | **+72.0%** |
原始首/末四分位比:
```text
Baseline:3.01×, 3.13×, 3.50×
Block: 0.81×, 0.65×, 1.40×
```
三 seed 平均 normalized spectrum 的 Block 峰值:
```text
layer 21 = 3.04× mean
layer 22 = 2.41×
layer 23 = 1.91×
layer 25 = 1.77×
```
depth-32 每个 AttnRes aggregation group 含 4 个 Transformer blocks。21–24 是第 6 组,
25–28 是第 7 组。尖峰集中在这两个中后段组附近,是数据中直接可见的结构;但仅凭本实验
不能断言 pseudo-query、某个 source 或组边界是唯一因果。
### 6.3 seed-3 的中期反例
depth-32 seed-3 在 step 2,000:
```text
Baseline CV 0.39965
Block CV 0.34914
```
此时 Block 更平;到 step 8,000 才变成:
```text
Baseline CV 0.40320
Block CV 0.42685
```
前两个 seed 在 step 2,000 已明显恶化,seed-3 没有。网站必须保留 seed switch,不能用最终
均值倒写成“三个 seed 从头到尾都一样”。
## 7. 绝对梯度尺度:更平不等于更强
最终 activation-gradient mean:
| depth | Baseline | Block | Block / Baseline |
|---:|---:|---:|---:|
| 16 | 约 `2.02×10⁻⁴` | 约 `1.16×10⁻⁴` | **0.574×** |
| 32 | 约 `1.44×10⁻⁴` | 约 `0.78×10⁻⁴` | **0.544×** |
所以 Block 的首尾比更接近 1,并不是因为它把晚层全部抬高到 Baseline 早层的强度。更接近的
描述是:
> 整体尺度下降,早层系统性高值被削弱,同时某些中后段位置相对全层均值形成尖峰。
这也是只看 normalized curve 或只看 CV 都不够的原因。
## 8. 参数梯度没有翻转 Round 04
最终核心参数梯度 CV:
| depth | Baseline mean | Block mean | Block / Baseline |
|---:|---:|---:|---:|
| 16 | 0.41596 | 0.68287 | 1.64× |
| 32 | 0.39661 | 0.77176 | 1.95× |
6 / 6 个配对中 Block 都更高。Round 04 的反结果不是在把梯度对象改成 activation 后自动消失;
两种梯度对象在本轮 final endpoint 都显示更高的跨层 CV。
但 activation gradient 又显示首尾失衡大幅改善,这说明“均匀”至少要拆成:
```text
首尾是否平衡
全层是否有尖峰
绝对尺度多大
参数更新信号是否平衡
```
一个标量不能代替全部。
## 9. Output RMS:最接近论文叙述的正向结果
三 seed 最终 post-MLP output RMS 的最后/第一层比:
| depth | Baseline | Block |
|---:|---:|---:|
| 16 | 4.59× | 1.17× |
| 32 | 6.08× | 1.89× |
Baseline output magnitude 随深度明显累积;Block 把增长限制在 aggregation group 内并产生
周期性 reset。这个缩小实验的 output-RMS 方向与论文 Figure 5(b) 的叙述一致。
仍不能把数值直接叠到论文图上:
- 模型宽度、深度和数据不同;
- 训练 Token 相差巨大;
- 论文图的确切 output norm / reduction 也未完整公开;
- 本实验的 Block group 是 8 组固定设计。
## 10. 验证 BPC 与真实成本
最终 BPC:
### depth-16
| seed | Baseline | Block | Block − Base |
|---:|---:|---:|---:|
| 2026073001 | 1.74880 | 1.73739 | −0.01141 |
| 2026073002 | 1.73683 | 1.73118 | −0.00565 |
| 2026073003 | 1.73637 | 1.72661 | −0.00976 |
| 均值 | 1.74067 | 1.73173 | **−0.00894** |
### depth-32
| seed | Baseline | Block | Block − Base |
|---:|---:|---:|---:|
| 2026073001 | 1.71790 | 1.71235 | −0.00554 |
| 2026073002 | 1.72698 | 1.70932 | −0.01765 |
| 2026073003 | 1.70951 | 1.70310 | −0.00642 |
| 均值 | 1.71813 | 1.70826 | **−0.00987** |
6 / 6 为负。这是有价值的次要方向,但 Round 05 没有为 BPC 再预注册一个新的 support 阈值,
因此不追加事后显著性结论。
真实成本:
| depth | Base mean ms | Block mean ms | time ratio | Base peak alloc | Block peak alloc | memory ratio |
|---:|---:|---:|---:|---:|---:|---:|
| 16 | 21.47 | 54.71 | 2.55× | 3.04 GB | 6.54 GB | 2.15× |
| 32 | 42.11 | 109.38 | 2.60× | 5.94 GB | 12.82 GB | 2.16× |
这个简单 eager 实现没有论文训练系统的 kernel、并行或工程优化;成本数值不应外推到 K3。
但它足以说明本站的 BPC 对比不是同 wall time / FLOPs。
## 11. 预注册判定为何是 mixed
支持需要:
```text
CV:3 / 3 seeds 改善,平均相对改善 ≥20%
AND
imbalance:3 / 3 seeds 改善,平均相对改善 ≥20%
```
concern 需要两项都以相同规则恶化。
实际:
```text
depth-16:
CV 3 / 3 恶化,平均 10.3%
imbalance 3 / 3 改善,平均 61.0%
depth-32:
CV 3 / 3 恶化,平均 60.0%
imbalance 3 / 3 改善,平均 72.0%
```
两个指标相反,所以两个 depth 都是 `mixed / inconclusive at this depth`。这不是“数据没规律”,
而是预注册的“更均匀”概念被实验拆成了两个方向相反的组成部分。
## 12. 开放工件与校验哈希
| 工件 | SHA-256 |
|---|---|
| manifest | `080afb17…6371` |
| protocol | `f772629b…ce22` |
| definition audit | `79221c56…cc6` |
| runner | `04ae69e1…800f` |
| analyzer | `017d38d9…2fc8` |
| aggregate canonical | `69be133c…7b51` |
| compact canonical | `8cdb7180…044f` |
| reproduction canonical | `addb2e59…f68a` |
公开目录包含:
- 12 个完整 formal raw JSON;
- 8 个两套 smoke raw JSON;
- 1 个完整 replay raw JSON;
- 完整 aggregate;
- 网站 compact payload;
- manifest、runner、analyzer 与 reproduction 清单;
- 前置定义审计、本协议和本结果审计。
`experiments/k3/attnres_gradient/reproduction.json` 记录每个 raw 文件 SHA-256。
## 13. 允许和禁止的结论
允许:
> 在本轮公开定义下,Block AttnRes 一致缓解了首/末深度四分位失衡,但在中后段形成局部
> 梯度尖峰,导致全层 CV 一致升高;因此“梯度更均匀”必须拆成多个指标解释。
> 同 token / step 预算下,Block 的最终验证 BPC 在 6 / 6 个配对中更低,但实际运行成本
> 约为 Baseline 的 2.6× step time 与 2.2× peak allocated memory。
禁止:
- “复现了论文 Figure 5(c)”;
- “论文的梯度结论是错的”;
- “已测到 Kimi K3 checkpoint 的真实梯度”;
- “Block 解决了梯度消失 / 爆炸”;
- “CV 更高就代表训练一定更不稳定”;
- “BPC 改善是同 FLOPs / wall time 优势”;
- 从 3 seeds 推导总体显著性;
- 从 depth 16 / 32 外推到 48B、1T+400B Token 或 K3 2.8T 参数。
## 14. 下一步最值得问什么
这轮已经把“梯度”从一个模糊词拆成了可复查对象。下一个有价值的问题不是再换一个漂亮
汇总指标,而是追踪局部尖峰从哪里来:
1. 分开捕获 pre-attention 与 pre-MLP residual positions;
2. 把 layer 21–25 的 activation-gradient 与 mixer source weights 同步对齐;
3. 比较 aggregation-group boundary 前后;
4. 在不改变 formal 结果的前提下,对相同 raw gradient tensor 做多种公开 reduction
sensitivity analysis;
5. 若官方之后发布 Figure 5 telemetry 代码,再按其定义单独开新 protocol。
这些属于后续轮次,不能倒写进本轮预注册结论。
@@ -0,0 +1,459 @@
# Kimi K3 第五轮:Attention Residuals 梯度定义与深度扩展实验协议
> 协议 ID:`llm-atlas-k3-attnres-gradient-scale-v1`
> 冻结日期:2026-07-30(Asia/Shanghai)
> 状态:正式输出前预注册
> 前置定义审计:`research/K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md`
## 0. 目标与一句话研究问题
Round 04 在缩小模型中得到两个同时成立的结果:
- Full / Block AttnRes 的 2,000-step 验证 BPC 都优于 Baseline;
- 以“每个 Transformer block 的核心**参数**梯度 RMS”定义时,跨深度 CV 反而更高。
论文 Figure 5(c) 没有公开足以唯一恢复的梯度测量合同。Round 05 不猜作者的隐藏代码,而是
冻结一个可复现、与 Figure 5 的 output/gradient 并列叙述对齐的 activation-gradient 定义,
再问:
> 当深度从 16 增至 32、训练预算从 2,000 增至 8,000 step 时,Block AttnRes 是否比
> PreNorm Baseline 更一致地降低 post-MLP block-output activation gradient 的跨深度失衡?
这不是 K3 checkpoint forward,也不是论文 Figure 5 数值复画。
## 1. 一手来源与不可补写的空白
| 工件 | 固定 revision / checksum | 用途 |
|---|---|---|
| Attention Residuals GitHub | `85e22310fe5ee860b4a023de312d791de8a5a5e6` | 公式、Figure 5 / 8、模型尺度 |
| `Attention_Residuals.pdf` | SHA-256 `e5831b0d…a26b2f` | 论文一手图文 |
| WikiText-2 raw | `Salesforce/wikitext@b08601e04326c79dfdd32d625aee71d232d685c3` | 固定公开训练语料 |
| Round 04 protocol | `llm-atlas-k3-attnres-reduced-v1` | 公共主干与数据合同来源 |
官方仓库没有模型实现、训练脚本、checkpoint 或 Figure 5 原始数据。以下字段不能归因给论文:
- Figure 5 的确切 gradient tensor;
- norm / reduction;
- diagnostic batch;
- AMP / clipping 时点;
- 单点还是时间平均。
Grok Headless 只进行一次对抗式方法检查;其建议和错误都在前置审计中公开,不是事实来源。
## 2. 设计总览
```text
2 个深度:16 / 32 Transformer blocks
× 2 个 residual graph:PreNorm Baseline / Block AttnRes
× 3 个冻结 seed
× 8,000 training steps
× 32 windows/step
× 256 target bytes/window
= 12 个正式训练格
= 786,432,000 target bytes
```
只比较 Baseline 与 Block,因为论文 Figure 5 的训练动力学面板也是这两个结构的直接对照。
Round 04 的 Full AttnRes 结果保持公开,但本轮不增加一个与主问题无关的 6-run 分支。
## 3. 数据合同
继承 Round 04 的语料和预处理:
1. 按固定 parquet 行序读取 `text`;
2. 每行追加一个 `\n`;
3. UTF-8 编码,无 normalization、strip、去空行或大小写改写;
4. byte vocabulary `0..255`;
5. 每个窗口连续取 257 bytes,前 256 预测后 256。
固定拼接后 split:
| split | bytes | SHA-256 |
|---|---:|---|
| train | 10,951,563 | `0ca7d3e74dbe44564ea5942b85232f1bbcb525c9cd481cd5d28a87ee90e7e9b4` |
| validation | 1,148,008 | `a42356f6a8ff1d25daf25ec9db49e10a537c265581b61c74604bb63231dee719` |
| test | 1,292,014 | `bfe9eb16ab9987fb88bde4ea9a30a00f2a45db01dfc14bad78d05325789c4f12` |
训练第 `step`、第 `row` 的 window 起点:
```text
z = first 8 bytes of SHA256(
protocol_id + "\0train-window\0" + seed + "\0" + step + "\0" + row
)
start = uint64_be(z) mod (len(train_bytes) - 257)
```
同一 seed 的 4 个结构格逐 step / row 使用完全相同的 token tensor。validation 64 windows、
diagnostic 16 windows,分别由标签 `validation-window` / `diagnostic-window` 与固定 index
生成,对全部结构与 seed 相同。
正式运行前 manifest 必须记录 parquet hash、split bytes/hash、全部 `3×8,000×32=768,000`
唯一训练窗口起点的 schedule hash、validation tensor hash 与 diagnostic tensor hash。
## 4. 模型合同
### 4.1 两个深度共享的结构
| 项 | 固定值 |
|---|---:|
| vocabulary | 256 bytes |
| context | 256 |
| `d_model` | 192 |
| heads / head dimension | 6 / 32 |
| `d_ff` | 768 |
| dropout | 0 |
| positional embedding | learned absolute,256 × 192 |
| norm | RMSNorm,`eps=1e-6` |
| attention | causal MHA;score 以 FP32 softmax |
| MLP | bias-free SwiGLU,`192→768`, `192→768`, `768→192` |
| embedding / readout | tied;final RMSNorm 后乘 token embedding |
全部 bias-free linear 与 embedding 初始化为 `N(0,0.02)`。attention output projection 与 MLP
down projection 的标准差为 `0.02 / sqrt(2×depth)`;普通 RMSNorm 为 1。
数值精度进一步固定为:attention / MLP 线性分支受 BF16 autocast;embedding、Baseline hidden
residual stream 与 Block aggregation partial 都以 FP32 累加。也就是说,Block 每个 BF16
branch output 在进入 `partial` 前显式转为 FP32。这样两种结构被捕获的 `h_l` 都是 FP32,
不会把 residual accumulator 精度差异混进梯度形状对比。
### 4.2 深度与 Block AttnRes 聚合
| Transformer depth | residual sublayers | aggregation groups | sublayers/group | Transformer blocks/group |
|---:|---:|---:|---:|---:|
| 16 | 32 | 8 | 4 | 2 |
| 32 | 64 | 8 | 8 | 4 |
Baseline 子层为:
```text
h ← h + f(RMSNorm(h))
```
Block AttnRes:
- embedding 永远是 source 0;
- 对已经完成的 aggregation-group sums 做跨组 softmax mixture;
- 组内 attention / MLP branch output 累加到 `partial`;
- 达到组边界时,保存完整 `partial` 为新 source,再开始下一组;
- output mixer 聚合 embedding + 8 个完整 group sums。
每个子层的 pseudo-query 为 `d_model` 向量,严格 zero-init;每个 source 的 key RMSNorm weight
严格 one-init。Baseline / Block 不要求总参数量或 residual-mixer FLOPs 相等,但同 seed /
depth 的 token/position embedding、attention、MLP、input norm、final norm 和 tied readout
必须逐 tensor SHA-256 exact。
## 5. 训练合同
| 项 | 固定值 |
|---|---:|
| seeds | `2026073001, 2026073002, 2026073003` |
| formal steps | 8,000 |
| batch | 32 |
| context | 256 |
| target bytes / run | 65,536,000 |
| optimizer | AdamW |
| betas / epsilon | `(0.9,0.95)` / `1e-8` |
| peak / min LR | `3e-4` / `3e-5` |
| warmup | 400 steps,linear |
| decay | cosine,step 400→8,000 |
| weight decay | `0.1` if `ndim>=2`,否则 `0` |
| parameter grad clip | global norm `1.0` |
| compute | BF16 autocast;FP32 optimizer state |
| compile | off / eager |
| device | one RTX 5090 |
| deterministic | deterministic algorithms;`CUBLAS_WORKSPACE_CONFIG=:4096:8` |
验证和诊断都发生在:
```text
step 0, 100, 500, 2,000, 4,000, 8,000
```
验证固定 64 windows,8 windows/eval batch,报告 token-mean CE nats 与
`bits_per_byte = CE / ln(2)`。诊断固定 16 windows,一次性输入,不改变 optimizer state。
计时合同:
- 前 20 个 training step 不计入;
- step 21–8,000 每步前后 CUDA synchronize;
- step 20 后 reset peak memory;
- 报告 mean / median / p95 step ms、peak allocated / reserved;
- timing、wall clock、hostname、GPU temperature 不进入 exact replay 字段。
## 6. 主梯度对象:精确到代码位置
### 6.1 `h_l` 定义
对 `l=1..depth`,统一在第 `l` 个 Transformer block 的 MLP branch 完成后捕获:
| 架构 | 精确定义 |
|---|---|
| Baseline | attention residual 与 MLP residual 都完成后的 hidden state |
| Block | MLP branch 已加入、本 aggregation partial 可能保存/reset **之前**的 partial |
所有 `h_l.shape = [16,256,192]`、dtype 为 FP32。实现必须为 diagnostic forward 返回独立
引用列表,不允许捕获 reset 后的零张量,不允许把 8 个 aggregation sources 当成 16 / 32 个
Transformer outputs。
### 6.2 diagnostic loss 与 gradient magnitude
```text
model.eval()
logits, h[1..depth] = forward(fixed_diagnostic_x, capture=true)
L = mean(cross_entropy(logits.float(), fixed_diagnostic_y))
backward(L) # 不使用 GradScaler,不执行 optimizer.step
m_l = sqrt(mean(float32(h_l.grad)² over batch×time×channel))
```
规则:
- forward 仍使用与训练一致的 BF16 autocast;
- logits 在 FP32 中计算 CE;
- loss 对全部 `16×256` targets 做算术平均,无 mask、无 label smoothing;
- backward 前 model / optimizer gradients 清零;
- activation gradient 在任何 parameter clipping 之前读取;
- diagnostic 不消耗训练数据,不进入 optimizer,不改变学习率或模型状态。
### 6.3 同位置 output magnitude
同一批 `h_l` 计算:
```text
o_l = sqrt(mean(float32(h_l)² over batch×time×channel))
```
它用于显示 Baseline 单调累积与 Block 组内周期,而不作为主 confirmatory endpoint。
## 7. 主指标与预注册判据
每个 diagnostic step 保存完整 `m_1..m_depth`。以下统计由未四舍五入的 float64 数组计算。
### 7.1 绝对尺度
```text
mean_grad = mean_l(m_l)
```
它防止“曲线更平只是全部梯度趋近于零”被 CV 隐藏。绝对尺度不设置优劣阈值,只公开。
### 7.2 归一化谱与 CV
```text
n_l = m_l / mean_grad
CV = population_std_l(m_l) / mean_grad
```
使用 population standard deviation(`ddof=0`)。网站必须同时显示 `m_l` 与 `n_l`,不能只显示
CV 排名。
### 7.3 前后四分位失衡
```text
q = depth / 4
Q_first = mean(m_1 .. m_q)
Q_last = mean(m_(depth-q+1) .. m_depth)
imbalance = abs(ln(Q_first / Q_last))
```
depth 16 时各取 4 层;depth 32 时各取 8 层。`imbalance=0` 才表示首尾一致;这个定义不会把
“早层偏大”和“晚层偏大”错误地都解释为越小越好。原始有符号比 `Q_first/Q_last` 仍公开。
### 7.4 seed 内配对对比
只在最终 step 8,000 做 confirmatory verdict:
```text
relative_CV_reduction
= (CV_baseline - CV_block) / CV_baseline
relative_imbalance_reduction
= (imbalance_baseline - imbalance_block) / imbalance_baseline
```
若 Baseline imbalance 精确为 0,则该 seed 的 relative imbalance reduction 定义为不可计算,
该 depth 自动不能得到“联合支持”;仍公开绝对差。
对每个 depth 分别判定:
- **joint directional support at this depth**:三个 seed 的 CV reduction 都 `>0`,其均值
`>=20%`;同时三个 seed 的 imbalance reduction 都 `>0`,其均值 `>=20%`。
- **joint directional concern at this depth**:三个 seed 的两项 reduction 都 `<0`,且两项
平均相对恶化都 `>=20%`。
- 其他:**mixed / inconclusive at this depth**。
总判定:
- 两个 depth 都 support:**scale-consistent directional support in this operationalization**;
- 两个 depth 都 concern:**scale-consistent directional concern in this operationalization**;
- 其他:**depth-dependent or inconclusive**。
不计算 p-value、population CI,不把 3 seeds 称为统计证明。
## 8. 必须公开的次要指标
### 8.1 全时间轨迹
六个预注册时点的以下数据必须全部公开,不能选择“最好看”的 checkpoint:
- validation BPC;
- absolute activation-gradient spectrum;
- normalized activation-gradient spectrum;
- activation CV;
- first/last quartile ratio 与 imbalance;
- output RMS spectrum;
- mean activation-gradient scale。
### 8.2 参数梯度
延续 Round 04 定义:每个 Transformer block 的 attention、MLP 与两个 input norm 的参数梯度
拼接后计算:
```text
parameter_grad_rms[l] = sqrt(sum(g²) / total_parameter_elements)
```
不含 embedding、final norm、LM head 与 AttnRes mixer 参数;clip 前读取。报告完整谱、CV 和
前后四分位失衡,但它们不进入 Round 05 主判定。
### 8.3 mixer / 成本
Block 同报:
- 各子层 softmax mixture 的 source-depth 分布;
- output mixer 分布;
- entropy 与 embedding / latest-complete-group mass;
- step time、peak allocated/reserved、参数量。
这些用于解释机制与成本,不改变 confirmatory verdict。
## 9. 运行前闸门
### 9.1 数据闸门
- split bytes/hash 与 Round 04 exact;
- 新 protocol 的 768,000-window schedule hash 落盘;
- validation / diagnostic tensor hash 落盘;
- 同 seed 四结构的至少 step 0 / 1 / 7,999 输入 tensor hash exact。
### 9.2 公共权重闸门
每个 depth / seed 的 Baseline 与 Block 公共参数逐 tensor exact;输出:
- 公共参数 tensor 数;
- 公共参数 element 数;
- name / shape / dtype / bytes 联合 hash;
- value bytes 联合 hash。
### 9.3 activation-gradient 闸门
四个 depth×architecture 格都必须通过:
1. 捕获数量严格等于 depth,shape 均为 `[16,256,192]`;
2. 全部 gradient 非 `None`、finite、storage 不别名;
3. 同输入把 loss 乘 2 后,每层 `m_l` 比值在 `2±1e-5`;
4. loss×2 前后 CV、normalized spectrum、quartile ratio 在 `1e-6` 绝对容差内;
5. 全新进程重复 smoke 的冻结字段 exact。
### 9.4 smoke
四个格都运行 20 training steps;冻结字段包括:
- protocol、architecture、depth、seed、device、dtype;
- data/schedule/tensor hashes;
- public-weight hashes;
- step 0 / 20 loss 与 validation;
- 全部预注册 diagnostic 数组;
- finite / alias / loss-scale checks。
smoke 不能写入 formal 目录。
## 10. 正式执行与独立 replay
12 个 formal grid 必须各自在全新进程中执行。目录键为:
```text
depth-{16|32}/{baseline|block}/seed-{2026073001|2026073002|2026073003}
```
完成后预注册 replay:
```text
depth-32 / block / seed-2026073001
```
replay 再从初始化训练完整 8,000 steps,不加载 formal checkpoint。比较时排除:
- wall time / step-time samples;
- peak memory;
- process ID / hostname;
- GPU 温度与驱动层瞬时字段;
- 文件路径和生成时间。
必须 exact 的字段:
- 数据与公共权重 hashes;
- 全部 validation CE/BPC;
- 全部 activation/output/parameter gradient 数组;
- mixer 数组与 summary;
- final model-state tensor hash;
- optimizer-state tensor hash;
- training-loss checkpoint 数组。
若 replay 不 exact,停止聚合并公开失败,不挑选另一 seed 替代。
## 11. 公开工件
正式结果完成后仓库必须包含:
```text
experiments/k3/attnres_gradient/
README.md
build_dataset.py
manifest.json
train.py
analyze.py
results/raw/*.json
results/compact.json
reproduction.json
research/
K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md
K3_ATTNRES_GRADIENT_SCALE_PROTOCOL.md
K3_ATTNRES_GRADIENT_SCALE_AUDIT.md
```
网站至少提供五个互相联动的视图:
1. 论文 Figure 5 的“已知 / 未定义”拆解;
2. activation gradient 绝对谱与 normalized spectrum;
3. depth 16 / 32、三 seed、六时间点对比;
4. output RMS 周期与 Block aggregation boundary;
5. activation gradient / parameter gradient 并排,以及判定、成本、哈希和声明边界。
图中必须能切换到所有负结果;不能只放均值、只放 final 或隐藏某个 seed。
## 12. 允许与禁止的结论
若达到支持条件,允许写:
> 在这个公开定义、两种缩小深度和 8,000-step byte-LM 合同中,Block AttnRes 方向一致地
> 降低了 post-MLP output activation gradient 的跨深度 CV 与首尾四分位失衡。
无论结果怎样,都禁止写:
- 复现了论文 Figure 5 的数值;
- 证明了论文未公开实现采用相同梯度定义;
- 证明了 Kimi K3 的真实梯度更健康;
- 证明 AttnRes 解决梯度消失、梯度爆炸或训练稳定性的全部问题;
- 从 3 seeds 推导总体显著性;
- 从 depth 16 / 32 外推至 48B、1T+400B Token 或 K3 2.8T 参数;
- 隐藏 activation 与 parameter gradient 方向不一致的结果。
## 13. 变更纪律
本文件提交后:
- 允许修复使实现符合本协议的 bug;
- 允许补充日志、注释、可视化和不改变数值的导出;
- 不允许看过 formal 结果后修改主指标、阈值、diagnostic step、seed、训练预算或 replay 格;
- 任何不得不改变实验合同的事项必须先停止、写入审计、升级 protocol ID,再重新执行全部 grid。
@@ -0,0 +1,231 @@
import { writeFileSync } from "node:fs";
const cdpPort = process.env.CDP_PORT ?? "9228";
const baseUrl = process.env.SITE_URL ?? "http://127.0.0.1:4328";
const pages = await fetch(`http://127.0.0.1:${cdpPort}/json/list`).then((response) => response.json());
const page = pages.find((entry) => entry.type === "page");
if (!page) throw new Error(`CDP ${cdpPort} 没有可用页面`);
const socket = new WebSocket(page.webSocketDebuggerUrl);
await new Promise((resolve, reject) => {
socket.addEventListener("open", resolve, { once: true });
socket.addEventListener("error", reject, { once: true });
});
let nextId = 0;
const pending = new Map();
const exceptions = [];
socket.addEventListener("message", (event) => {
const message = JSON.parse(event.data);
if (message.id && pending.has(message.id)) {
const { resolve, reject } = pending.get(message.id);
pending.delete(message.id);
if (message.error) reject(new Error(message.error.message));
else resolve(message.result);
}
if (message.method === "Runtime.exceptionThrown") {
exceptions.push(message.params.exceptionDetails.exception?.description ?? message.params.exceptionDetails.text);
}
});
const command = (method, params = {}) => new Promise((resolve, reject) => {
const id = ++nextId;
pending.set(id, { resolve, reject });
socket.send(JSON.stringify({ id, method, params }));
});
const pause = (milliseconds) => new Promise((resolve) => setTimeout(resolve, milliseconds));
const evaluate = async (expression) => {
const result = await command("Runtime.evaluate", { expression, returnByValue: true, awaitPromise: true });
if (result.exceptionDetails) throw new Error(result.exceptionDetails.exception?.description ?? result.exceptionDetails.text);
return result.result.value;
};
const navigate = async (path) => {
await command("Page.navigate", { url: `${baseUrl}${path}` });
for (let attempt = 0; attempt < 100; attempt += 1) {
await pause(100);
if (await evaluate("document.readyState === 'complete'")) return;
}
throw new Error(`${path} 加载超时`);
};
const screenshot = async (path) => {
const result = await command("Page.captureScreenshot", { format: "png", captureBeyondViewport: false });
writeFileSync(path, Buffer.from(result.data, "base64"));
};
await command("Page.enable");
await command("Runtime.enable");
await command("Emulation.setDeviceMetricsOverride", {
width: 1440,
height: 1100,
deviceScaleFactor: 1,
mobile: false,
});
await navigate("/k3/");
const desktop = await evaluate(`(() => {
const root = document.querySelector("[data-gradient-lab]");
root.scrollIntoView({ block: "start", behavior: "instant" });
window.scrollBy(0, -78);
const text = (selector) => root.querySelector(selector)?.textContent.trim();
const panel = () => root.querySelector("[data-gradient-panel]:not([hidden])")?.dataset.gradientPanel;
const points = (selector) => root.querySelector(selector)?.getAttribute("points");
const setSelect = (selector, value) => {
const node = root.querySelector(selector);
node.value = value;
node.dispatchEvent(new Event("change", { bubbles: true }));
};
const initial = {
panel: panel(),
tabs: root.querySelectorAll("[data-gradient-tab]").length,
panels: root.querySelectorAll("[data-gradient-panel]").length,
ledger: root.querySelectorAll(".gradient-ledger article").length,
definition: root.textContent.includes("论文作者就是这样算的") &&
root.textContent.includes("activation gradient") &&
root.textContent.includes("parameter gradient"),
};
root.querySelector('[data-gradient-tab="spectrum"]').click();
const spectrumInitial = {
panel: panel(),
baseCv: text("[data-spectrum-base-cv]"),
blockCv: text("[data-spectrum-block-cv]"),
baseRatio: text("[data-spectrum-base-ratio]"),
blockRatio: text("[data-spectrum-block-ratio]"),
baseLine: points('[data-chart-line="baseline"]'),
blockLine: points('[data-chart-line="block"]'),
boundaries: root.querySelectorAll("[data-chart-groups] .group-boundary").length,
};
root.querySelector('[data-spectrum-scale="normalized"]').click();
const normalized = {
title: text("[data-spectrum-title]"),
baseLine: points('[data-chart-line="baseline"]'),
};
root.querySelector('[data-spectrum-depth="16"]').click();
const depth16 = {
baseCv: text("[data-spectrum-base-cv]"),
blockCv: text("[data-spectrum-block-cv]"),
pointCount: points('[data-chart-line="baseline"]').split(" ").length,
boundaries: root.querySelectorAll("[data-chart-groups] .group-boundary").length,
};
setSelect("[data-spectrum-seed]", "2026073001");
setSelect("[data-spectrum-step]", "2000");
const seedStep = {
state: text("[data-spectrum-state]"),
baseCv: text("[data-spectrum-base-cv]"),
blockCv: text("[data-spectrum-block-cv]"),
};
root.querySelector('[data-gradient-tab="timeline"]').click();
const timelineInitial = {
panel: panel(),
title: text("[data-time-title]"),
line: points('[data-time-line="block"]'),
pointCount: root.querySelectorAll('[data-time-points="block"] circle').length,
};
root.querySelector('[data-time-metric="imbalance"]').click();
const timelineChanged = {
title: text("[data-time-title]"),
line: points('[data-time-line="block"]'),
};
root.querySelector('[data-gradient-tab="output"]').click();
const outputInitial = {
panel: panel(),
pointCount: points('[data-output-line="block"]').split(" ").length,
bars: root.querySelectorAll("[data-output-bars] i").length,
boundaries: root.querySelectorAll("[data-output-groups] .group-boundary").length,
copy: text("[data-output-copy]"),
};
root.querySelector('[data-output-depth="16"]').click();
const outputDepth16 = {
pointCount: points('[data-output-line="block"]').split(" ").length,
bars: root.querySelectorAll("[data-output-bars] i").length,
copy: text("[data-output-copy]"),
};
root.querySelector('[data-gradient-tab="verdict"]').click();
const verdict = {
panel: panel(),
rows: root.querySelectorAll(".verdict-table tbody tr").length,
metrics: root.querySelectorAll(".metric-pairs article").length,
costs: root.querySelectorAll(".cost-compare article").length,
hashes: root.querySelectorAll(".hash-ledger code").length,
exact: root.textContent.includes("model + optimizer exact") &&
root.textContent.includes("all frozen fields exact"),
mixed: root.textContent.includes("depth-dependent or inconclusive"),
};
const first = root.querySelector('[data-gradient-tab="definition"]');
first.focus();
first.dispatchEvent(new KeyboardEvent("keydown", { key: "ArrowRight", bubbles: true }));
const keyboard = {
selected: root.querySelector('[data-gradient-tab][aria-selected="true"]').dataset.gradientTab,
panel: panel(),
};
return {
initial, spectrumInitial, normalized, depth16, seedStep,
timelineInitial, timelineChanged, outputInitial, outputDepth16,
verdict, keyboard,
documentOverflow: document.documentElement.scrollWidth - document.documentElement.clientWidth,
rootOverflow: root.scrollWidth - root.clientWidth,
};
})()`);
await pause(180);
await screenshot("/tmp/llm-atlas-k3-attnres-gradient-desktop.png");
await command("Emulation.setDeviceMetricsOverride", {
width: 390,
height: 844,
deviceScaleFactor: 1,
mobile: true,
});
await navigate("/k3/");
const mobile = await evaluate(`(() => {
const root = document.querySelector("[data-gradient-lab]");
root.scrollIntoView({ block: "start", behavior: "instant" });
window.scrollBy(0, -64);
root.querySelector('[data-gradient-tab="spectrum"]').click();
root.querySelector('[data-spectrum-depth="16"]').click();
root.querySelector('[data-gradient-tab="output"]').click();
root.querySelector('[data-output-depth="16"]').click();
return {
tabs: root.querySelectorAll("[data-gradient-tab]").length,
ledger: root.querySelectorAll(".gradient-ledger article").length,
outputBars: root.querySelectorAll("[data-output-bars] i").length,
visiblePanel: root.querySelector("[data-gradient-panel]:not([hidden])")?.dataset.gradientPanel,
documentOverflow: document.documentElement.scrollWidth - document.documentElement.clientWidth,
rootOverflow: root.scrollWidth - root.clientWidth,
};
})()`);
await pause(180);
await screenshot("/tmp/llm-atlas-k3-attnres-gradient-mobile.png");
const report = { desktop, mobile, exceptions };
console.log(JSON.stringify(report, null, 2));
const numeric = (value) => Number.parseFloat(value.replace("−", "-"));
const failures = [];
if (desktop.initial.panel !== "definition" || desktop.initial.tabs !== 5 || desktop.initial.panels !== 5 || desktop.initial.ledger !== 6) failures.push("五视图初始结构异常");
if (!desktop.initial.definition) failures.push("定义与对象边界缺失");
if (desktop.spectrumInitial.panel !== "spectrum" || Math.abs(numeric(desktop.spectrumInitial.baseCv) - 0.3786) > 1e-4 || Math.abs(numeric(desktop.spectrumInitial.blockCv) - 0.5996) > 1e-4) failures.push("depth-32 final spectrum 读数异常");
if (desktop.spectrumInitial.boundaries !== 7 || desktop.spectrumInitial.baseLine === desktop.normalized.baseLine || !desktop.normalized.title.includes("NORMALIZED")) failures.push("绝对/归一化谱或组边界异常");
if (desktop.depth16.pointCount !== 16 || desktop.depth16.boundaries !== 7 || Math.abs(numeric(desktop.depth16.baseCv) - 0.4246) > 1e-4 || Math.abs(numeric(desktop.depth16.blockCv) - 0.4678) > 1e-4) failures.push("depth-16 spectrum 切换异常");
if (!desktop.seedStep.state.includes("2026073001") || !desktop.seedStep.state.includes("2,000") || Math.abs(numeric(desktop.seedStep.baseCv) - 0.4101) > 1e-4 || Math.abs(numeric(desktop.seedStep.blockCv) - 0.3656) > 1e-4) failures.push("seed/checkpoint spectrum 切换异常");
if (desktop.timelineInitial.panel !== "timeline" || desktop.timelineInitial.pointCount !== 6 || desktop.timelineInitial.line === desktop.timelineChanged.line || !desktop.timelineChanged.title.includes("FIRST / LAST")) failures.push("六时点轨迹指标切换异常");
if (desktop.outputInitial.panel !== "output" || desktop.outputInitial.pointCount !== 32 || desktop.outputInitial.bars !== 32 || desktop.outputInitial.boundaries !== 7 || desktop.outputDepth16.pointCount !== 16 || desktop.outputDepth16.bars !== 16 || !desktop.outputDepth16.copy.includes("DEPTH 16")) failures.push("Output RMS 深度/组节律切换异常");
if (desktop.verdict.panel !== "verdict" || desktop.verdict.rows !== 2 || desktop.verdict.metrics !== 4 || desktop.verdict.costs !== 3 || desktop.verdict.hashes !== 4 || !desktop.verdict.exact || !desktop.verdict.mixed) failures.push("联合判定、成本或重放视图异常");
if (desktop.keyboard.selected !== "spectrum" || desktop.keyboard.panel !== "spectrum") failures.push("键盘 tab 导航异常");
if (desktop.documentOverflow > 1 || desktop.rootOverflow > 1 || mobile.documentOverflow > 1 || mobile.rootOverflow > 1) failures.push("桌面或移动端出现文档级横向溢出");
if (mobile.tabs !== 5 || mobile.ledger !== 6 || mobile.outputBars !== 16 || mobile.visiblePanel !== "output") failures.push("移动端交互结构异常");
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
if (failures.length) {
console.error(`\nFAIL\n- ${failures.join("\n- ")}`);
process.exitCode = 1;
} else {
console.log("\nPASS K3 AttnRes gradient browser regression");
}
socket.close();
@@ -0,0 +1,68 @@
import { createHash } from "node:crypto";
import { readdirSync, readFileSync } from "node:fs";
const read = (path) => {
const bytes = readFileSync(new URL(path, import.meta.url));
return {
bytes,
json: JSON.parse(bytes),
sha256: createHash("sha256").update(bytes).digest("hex"),
};
};
const raw = read("../src/data/k3-attnres-gradient.json");
const compact = read("../src/data/k3-attnres-gradient-compact.json");
const reproduction = read("../experiments/k3/attnres_gradient/reproduction.json");
const rawDirectory = new URL("../experiments/k3/attnres_gradient/results/raw/", import.meta.url);
const failures = [];
const expect = (condition, message) => {
if (!condition) failures.push(message);
};
expect(raw.sha256 === "ad461cbe74fc671f356288d37c9b628618814915116e4fbe03aaaa48367e6a8d", "raw aggregate SHA-256 changed");
expect(compact.sha256 === "5377a5e731db2fdc85a0327f05d43f1cd067d3fc34633619d3004f3fc31968e3", "compact payload SHA-256 changed");
expect(reproduction.sha256 === "aedcde6accc6eb1c24b122ffb55706ef620062e848d9fe4c5f622b084a4fdea6", "reproduction payload SHA-256 changed");
expect(compact.json.protocol_id === "llm-atlas-k3-attnres-gradient-scale-v1", "protocol identity mismatch");
expect(compact.json.study.formal_runs === 12, "formal grid is not 12 cells");
expect(compact.json.study.formal_target_bytes === 786_432_000, "formal target-byte budget mismatch");
expect(compact.json.study.replay_target_bytes === 65_536_000, "replay byte budget mismatch");
expect(compact.json.cells.length === 12, "compact cell count mismatch");
expect(Object.keys(raw.json.runs).length === 12, "raw formal cell count mismatch");
expect(readdirSync(rawDirectory).filter((name) => name.endsWith(".json")).length === 21, "public raw output count is not 21");
expect(reproduction.json.replay_exact.exact, "full formal replay is not exact");
expect(Object.values(reproduction.json.smoke_exact).every((row) => row.exact && row.gradient_gate.passed), "paired smoke or loss-scale gate failed");
for (const depth of ["16", "32"]) {
const summary = compact.json.depth_summaries[depth];
expect(summary.verdict.label === "mixed / inconclusive at this depth", `depth ${depth} verdict changed`);
expect(summary.by_seed.length === 3, `depth ${depth} seed count changed`);
expect(summary.by_seed.every((row) => row.block_minus_baseline_bpc < 0), `depth ${depth} BPC pairing changed`);
expect(summary.by_seed.every((row) => row.relative_cv_reduction < 0), `depth ${depth} CV counterevidence changed`);
expect(summary.by_seed.every((row) => row.relative_imbalance_reduction > 0), `depth ${depth} first/last improvement changed`);
}
expect(Math.abs(compact.json.depth_summaries["16"].means.relative_cv_reduction - (-0.10264135379685868)) < 1e-15, "depth-16 CV contrast changed");
expect(Math.abs(compact.json.depth_summaries["32"].means.relative_cv_reduction - (-0.600330100169428)) < 1e-15, "depth-32 CV contrast changed");
expect(Math.abs(compact.json.depth_summaries["16"].means.relative_imbalance_reduction - 0.6099652484299795) < 1e-15, "depth-16 imbalance contrast changed");
expect(Math.abs(compact.json.depth_summaries["32"].means.relative_imbalance_reduction - 0.7200517407719272) < 1e-15, "depth-32 imbalance contrast changed");
expect(compact.json.overall_verdict === "depth-dependent or inconclusive", "overall preregistered verdict changed");
if (failures.length) {
console.error(`FAIL K3 AttnRes gradient data\n- ${failures.join("\n- ")}`);
process.exit(1);
}
console.log(JSON.stringify({
protocol: compact.json.protocol_id,
formalRuns: compact.json.study.formal_runs,
formalTargetBytes: compact.json.study.formal_target_bytes,
depth16: compact.json.depth_summaries["16"].verdict.label,
depth32: compact.json.depth_summaries["32"].verdict.label,
replayExact: reproduction.json.replay_exact.exact,
hashes: {
raw: raw.sha256,
compact: compact.sha256,
reproduction: reproduction.sha256,
},
}, null, 2));
console.log("PASS K3 AttnRes gradient frozen data");
+7 -3
View File
@@ -81,6 +81,8 @@ const overview = await evaluate(`(() => ({
artifactMismatch: document.querySelector("#artifacts")?.textContent.includes("A_log [128] ≠ expected [96]"),
attnresTabs: document.querySelectorAll("[data-attnres-tab]").length,
attnresPanels: document.querySelectorAll("[data-attnres-panel]").length,
gradientTabs: document.querySelectorAll("[data-gradient-tab]").length,
gradientPanels: document.querySelectorAll("[data-gradient-panel]").length,
nativeVisionCorrected: document.body.textContent.includes("MoonViT‑V2 从头训练") &&
document.body.textContent.includes("同一个 next-token prediction objective"),
staleVisionClaim: document.body.textContent.includes("先固定语言模型训练视觉组件"),
@@ -290,8 +292,9 @@ const mobile = await evaluate(`(() => {
artifactTabs: document.querySelectorAll("[data-artifact-tab]").length,
artifactLayers: document.querySelectorAll("[data-layer-cell]").length,
attnresTabs: document.querySelectorAll("[data-attnres-tab]").length,
gradientTabs: document.querySelectorAll("[data-gradient-tab]").length,
offenders: [...document.querySelectorAll("body *")]
.filter((node) => !node.closest(".paper-chain, .spec-table-wrap, .cache-strip, .architecture-explorer, [data-k3-lab], [data-k3-artifact-lab], [data-attnres-lab]"))
.filter((node) => !node.closest(".paper-chain, .spec-table-wrap, .cache-strip, .architecture-explorer, [data-k3-lab], [data-k3-artifact-lab], [data-attnres-lab], [data-gradient-lab]"))
.filter((node) => node.getBoundingClientRect().right > document.documentElement.clientWidth + 1)
.slice(0, 15)
.map((node) => ({
@@ -322,12 +325,13 @@ console.log(JSON.stringify(report, null, 2));
const numeric = (text) => Number.parseFloat(text.replaceAll(",", "").replace("−", "-"));
const failures = [];
if (!overview.title.includes("因果环节")) failures.push("K3 二轮标题异常");
if (overview.sections !== 33 || overview.tocLinks !== 33) failures.push("32 个编号专题加阅读链的目录结构异常");
if (overview.sections !== 34 || overview.tocLinks !== 34) failures.push("33 个编号专题加阅读链的目录结构异常");
if (overview.ledgers !== 32 || overview.reportMap !== 9) failures.push("32 张问题账或报告地图异常");
if (overview.figureAtlas !== 21 || overview.paperLinks !== 100 || overview.paperGroups < 12) failures.push("图表审计或 100 节点阅读链异常");
if (overview.labTabs !== 8 || overview.labPanels !== 8) failures.push("八联实验结构异常");
if (overview.artifactTabs !== 4 || overview.artifactPanels !== 4 || overview.artifactLayers !== 93 || !overview.artifactMismatch) failures.push("开放工件四视图、93 层条带或形状冲突异常");
if (overview.attnresTabs !== 5 || overview.attnresPanels !== 5) failures.push("AttnRes 独立实验五视图异常");
if (overview.gradientTabs !== 5 || overview.gradientPanels !== 5) failures.push("AttnRes 梯度定义扩展五视图异常");
if (!overview.nativeVisionCorrected || overview.staleVisionClaim) failures.push("原生多模态纠错未生效或旧错误残留");
if (overview.documentOverflow > 1 || mobile.documentOverflow > 1) failures.push("桌面或移动端存在文档级横向溢出");
if (labs.memoryInitial.panel !== "memory" || numeric(labs.memoryInitial.additiveError) <= numeric(labs.memoryInitial.deltaError)) failures.push("Delta memory 初始递推异常");
@@ -352,7 +356,7 @@ if (artifacts.parameterChanged.shape !== "[96,128] F32" || !artifacts.parameterC
if (artifacts.reproductionInitial.panel !== "reproduction" || numeric(artifacts.reproductionInitial.speedup) !== 1.85 || numeric(artifacts.reproductionInitial.localMean) < 2.6 || !artifacts.reproductionInitial.exactSuite || numeric(artifacts.reproductionInitial.cv) < 2) failures.push("FlashKDA H20、本机 exact suite 或 router 初始探针异常");
if (numeric(artifacts.reproductionChanged.speedup) !== 3.27 || numeric(artifacts.reproductionChanged.flash) !== 0.7064 || numeric(artifacts.reproductionChanged.localMean) >= numeric(artifacts.reproductionInitial.localMean) || !artifacts.reproductionChanged.localMode.includes("FP32 state") || numeric(artifacts.reproductionChanged.cv) <= numeric(artifacts.reproductionInitial.cv) || numeric(artifacts.reproductionChanged.zero) <= numeric(artifacts.reproductionInitial.zero)) failures.push("GB200 benchmark、本机 varlen/state 或 synthetic router counterexample 未更新");
if (artifacts.keyboardSelected !== "tensors" || artifacts.keyboardVisible !== "tensors") failures.push("开放工件键盘 tab 导航异常");
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 8 || mobile.artifactTabs !== 4 || mobile.artifactLayers !== 93 || mobile.attnresTabs !== 5) failures.push("移动端导航或实验异常");
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 8 || mobile.artifactTabs !== 4 || mobile.artifactLayers !== 93 || mobile.attnresTabs !== 5 || mobile.gradientTabs !== 5) failures.push("移动端导航或实验异常");
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
+813
View File
@@ -0,0 +1,813 @@
---
import rawLab from "@/data/k3-attnres-gradient-compact.json";
const lab = rawLab as any;
const json = JSON.stringify(lab).replaceAll("<", "\\u003c");
const depth16 = lab.depth_summaries["16"];
const depth32 = lab.depth_summaries["32"];
const formatSigned = (value: number, digits = 4) =>
`${value < 0 ? "−" : value > 0 ? "+" : ""}${Math.abs(value).toFixed(digits)}`;
const pct = (value: number, digits = 1) => `${(value * 100).toFixed(digits)}%`;
const shortHash = (value: string) => `${value.slice(0, 10)}…${value.slice(-8)}`;
---
<figure class="gradient-lab" data-gradient-lab>
<header class="gradient-head">
<div>
<p>ROUND 05 / GRADIENT DEFINITION × DEPTH SCALE</p>
<h3>“早层不再过大”与“整条谱更均匀”不是同一件事</h3>
</div>
<p>
16 / 32 Transformer blocks · Baseline / Block · 3 seeds · 8,000 steps。
被测对象是 post-MLP output activation gradient,不冒充论文未公开的 Figure 5 实现。
</p>
</header>
<div class="gradient-ledger">
<article><span>FORMAL GRID</span><b>2 × 2 × 3</b><p>12 个独立训练格</p></article>
<article><span>TARGET BYTES</span><b>786.432M</b><p>每格 65,536,000</p></article>
<article><span>DEPTH</span><b>16 → 32</b><p>宽度固定 192</p></article>
<article class="split"><span>CV</span><b>6 / 6 更差</b><p>中后段局部尖峰</p></article>
<article class="pass"><span>FIRST ↔ LAST</span><b>6 / 6 更近</b><p>失衡改善 56%–81%</p></article>
<article class="pass"><span>FULL REPLAY</span><b>exact</b><p>model + optimizer state</p></article>
</div>
<div class="gradient-tabs" role="tablist" aria-label="选择 AttnRes 梯度定义与深度实验视图">
<button type="button" role="tab" data-gradient-tab="definition" aria-selected="true">
<span>01</span><b>论文到底定义了什么</b><small>known · underdefined · operationalization</small>
</button>
<button type="button" role="tab" data-gradient-tab="spectrum" aria-selected="false">
<span>02</span><b>绝对谱与归一化谱</b><small>depth · seed · checkpoint</small>
</button>
<button type="button" role="tab" data-gradient-tab="timeline" aria-selected="false">
<span>03</span><b>六个时点怎样演化</b><small>CV · imbalance · mean scale</small>
</button>
<button type="button" role="tab" data-gradient-tab="output" aria-selected="false">
<span>04</span><b>Output RMS 与块节律</b><small>growth · reset · group boundary</small>
</button>
<button type="button" role="tab" data-gradient-tab="verdict" aria-selected="false">
<span>05</span><b>联合判定、成本与重放</b><small>activation ≠ parameter</small>
</button>
</div>
<section class="gradient-panel" data-gradient-panel="definition">
<div class="panel-lead">
<div><span>I / DEFINITION AUDIT</span><h4>论文画出一条“梯度曲线”,却没有给出足以唯一重算的测量合同</h4></div>
<p>
Figure 5(c) 只写 “Each transformer block’s gradient magnitude”。
官方仓库没有训练代码、checkpoint、统计脚本或原始数组。
</p>
</div>
<div class="known-grid">
<article class="known">
<span>OFFICIAL / KNOWN</span>
<h5>图与文字能确认</h5>
<ul>
<li>横轴是 Transformer block index。</li>
<li>Figure 5(b) 同时画 block output magnitude。</li>
<li>正文说 Baseline 早层梯度过大,Block 更均匀。</li>
<li>最终模型约 27 blocks / 54 residual layers。</li>
</ul>
</article>
<article class="unknown">
<span>OFFICIAL / UNDERDEFINED</span>
<h5>无法从公开工件唯一恢复</h5>
<ul>
<li>activation、branch 还是 parameter gradient?</li>
<li>L2、RMS、mean absolute 还是别的 norm?</li>
<li>batch / token / channel 怎样 reduction?</li>
<li>哪个 checkpoint、AMP / clip 前还是后?</li>
</ul>
</article>
</div>
<div class="object-chain" aria-label="Round 05 activation gradient 测量对象">
<div><span>FIXED INPUT</span><b>16 × 256 bytes</b><p>同一 diagnostic tensor</p></div>
<i>→</i>
<div><span>POST-MLP OUTPUT</span><b>h₁ … h<sub>L</sub></b><p>每个 Transformer block 一个</p></div>
<i>→</i>
<div><span>TOKEN-MEAN CE</span><b>ℒ</b><p>FP32 cross entropy</p></div>
<i>→</i>
<div class="accent"><span>MEASURE</span><b>RMS(∂ℒ/∂h<sub>l</sub>)</b><p>B × T × C 联合 RMS</p></div>
</div>
<div class="object-compare">
<article><span>ROUND 04</span><b>∇<sub>θl</sub>ℒ</b><p>核心参数梯度:权重收到多大更新信号。</p></article>
<i>≠</i>
<article><span>ROUND 05</span><b>∂ℒ/∂h<sub>l</sub></b><p>activation gradient:损失对这一深度表示有多敏感。</p></article>
<p>两种对象都公开;新指标不会覆盖上一轮反结果。</p>
</div>
<div class="definition-boundary">
<b>能说</b><p>“这是与 Figure 5 叙述对齐的一种公开 operationalization。”</p>
<b>不能说</b><p>“论文作者就是这样算的,或本站复画了 Figure 5(c)。”</p>
</div>
</section>
<section class="gradient-panel" data-gradient-panel="spectrum" hidden>
<div class="panel-lead">
<div><span>II / ALIGNED DEPTH SPECTRUM</span><h4>同一条梯度谱,绝对值与归一化形状要一起看</h4></div>
<p>竖线是 8 个 AttnRes aggregation groups 的边界;Baseline 也画同位置,方便逐层配对。</p>
</div>
<div class="lab-controls">
<div role="group" aria-label="选择深度">
<button type="button" data-spectrum-depth="16" aria-pressed="false">DEPTH 16</button>
<button type="button" data-spectrum-depth="32" aria-pressed="true">DEPTH 32</button>
</div>
<label>SEED
<select data-spectrum-seed>
<option value="mean">3-SEED MEAN</option>
{lab.study.seeds.map((seed: number) => <option value={String(seed)}>{seed}</option>)}
</select>
</label>
<label>CHECKPOINT
<select data-spectrum-step>
{lab.study.diagnostic_steps.map((step: number) => <option value={String(step)} selected={step === 8000}>STEP {step.toLocaleString("en-US")}</option>)}
</select>
</label>
<div role="group" aria-label="选择绝对或归一化梯度谱">
<button type="button" data-spectrum-scale="absolute" aria-pressed="true">ABSOLUTE RMS</button>
<button type="button" data-spectrum-scale="normalized" aria-pressed="false">÷ LAYER MEAN</button>
</div>
</div>
<div class="spectrum-layout">
<div class="chart-shell">
<header><span data-spectrum-title>ACTIVATION GRADIENT RMS</span><b>POST-MLP BLOCK OUTPUT</b></header>
<svg data-spectrum-chart viewBox="0 0 920 360" role="img" aria-label="Baseline 与 Block 的逐层 activation gradient 谱">
<g data-chart-groups></g>
<g data-chart-grid></g>
<polyline data-chart-line="baseline" class="series baseline"></polyline>
<polyline data-chart-line="block" class="series block"></polyline>
<g data-chart-points="baseline"></g>
<g data-chart-points="block"></g>
<text x="460" y="350" class="axis-title">TRANSFORMER BLOCK INDEX</text>
</svg>
<div class="chart-legend">
<span><i class="baseline"></i>Baseline</span>
<span><i class="block"></i>Block AttnRes</span>
<span><i class="boundary"></i>aggregation boundary</span>
</div>
</div>
<div class="spectrum-readout">
<span data-spectrum-state>DEPTH 32 · 3-SEED MEAN · STEP 8,000</span>
<article><b>POPULATION CV</b><div><span>BASE</span><strong data-spectrum-base-cv>0.3786</strong></div><div><span>BLOCK</span><strong data-spectrum-block-cv>0.5996</strong></div></article>
<article><b>FIRST / LAST QUARTILE</b><div><span>BASE</span><strong data-spectrum-base-ratio>3.21×</strong></div><div><span>BLOCK</span><strong data-spectrum-block-ratio>0.95×</strong></div></article>
<article><b>MEAN ABSOLUTE SCALE</b><div><span>BASE</span><strong data-spectrum-base-mean>1.44e−4</strong></div><div><span>BLOCK</span><strong data-spectrum-block-mean>0.78e−4</strong></div></article>
</div>
</div>
<div class="split-result">
<article><span>FIRST ↔ LAST</span><b>更接近</b><p>Baseline 的早层整体隆起被削弱。</p></article>
<i>但</i>
<article><span>ALL-LAYER CV</span><b>反而更高</b><p>中后段少数位置形成更尖的峰。</p></article>
<i>所以</i>
<article class="accent"><span>PRE-REGISTERED</span><b>mixed</b><p>“更均匀”必须拆成至少两个指标。</p></article>
</div>
</section>
<section class="gradient-panel" data-gradient-panel="timeline" hidden>
<div class="panel-lead">
<div><span>III / SIX FROZEN CHECKPOINTS</span><h4>Block 的中期优势会反转;不能挑一个 checkpoint 讲故事</h4></div>
<p>横轴按六个预注册诊断时点等距排列;标签保留真实 step,不暗示实际时间等距。</p>
</div>
<div class="lab-controls">
<div role="group" aria-label="选择时间轨迹深度">
<button type="button" data-time-depth="16" aria-pressed="false">DEPTH 16</button>
<button type="button" data-time-depth="32" aria-pressed="true">DEPTH 32</button>
</div>
<label>SEED
<select data-time-seed>
<option value="mean">3-SEED MEAN</option>
{lab.study.seeds.map((seed: number) => <option value={String(seed)}>{seed}</option>)}
</select>
</label>
<div role="group" aria-label="选择时间轨迹指标">
<button type="button" data-time-metric="cv" aria-pressed="true">CV</button>
<button type="button" data-time-metric="imbalance" aria-pressed="false">FIRST/LAST IMBALANCE</button>
<button type="button" data-time-metric="mean" aria-pressed="false">ABSOLUTE MEAN</button>
</div>
</div>
<div class="chart-shell timeline-chart">
<header><span data-time-title>POPULATION CV</span><b data-time-copy>DEPTH 32 · 3-SEED MEAN</b></header>
<svg data-time-chart viewBox="0 0 920 360" role="img" aria-label="六个固定训练时点的 activation gradient 指标轨迹">
<g data-time-grid></g>
<polyline data-time-line="baseline" class="series baseline"></polyline>
<polyline data-time-line="block" class="series block"></polyline>
<g data-time-points="baseline"></g>
<g data-time-points="block"></g>
<text x="460" y="350" class="axis-title">PREREGISTERED DIAGNOSTIC STEP</text>
</svg>
<div class="chart-legend"><span><i class="baseline"></i>Baseline</span><span><i class="block"></i>Block AttnRes</span></div>
</div>
<div class="timeline-notes">
<article><span>SEED 01 / DEPTH 32</span><b>0.344 → 0.538</b><p>step 2,000:Block 已显著更尖。</p></article>
<article><span>SEED 02 / DEPTH 32</span><b>0.352 → 0.618</b><p>同一时点复现恶化方向。</p></article>
<article class="counter"><span>SEED 03 / DEPTH 32</span><b>0.400 → 0.349</b><p>中期反例:Block 此时反而更平。</p></article>
<article><span>SEED 03 / FINAL</span><b>0.403 → 0.427</b><p>到 8,000 step 才轻微反转。</p></article>
</div>
</section>
<section class="gradient-panel" data-gradient-panel="output" hidden>
<div class="panel-lead">
<div><span>IV / OUTPUT MAGNITUDE</span><h4>梯度结论 mixed,不代表论文所有训练动力学叙述都没有出现</h4></div>
<p>同一个 post-MLP 位置计算 output RMS;这里不做 backward,也不改变主判定。</p>
</div>
<div class="lab-controls">
<div role="group" aria-label="选择 output RMS 深度">
<button type="button" data-output-depth="16" aria-pressed="false">DEPTH 16</button>
<button type="button" data-output-depth="32" aria-pressed="true">DEPTH 32</button>
</div>
<label>SEED
<select data-output-seed>
<option value="mean">3-SEED MEAN</option>
{lab.study.seeds.map((seed: number) => <option value={String(seed)}>{seed}</option>)}
</select>
</label>
<label>CHECKPOINT
<select data-output-step>
{lab.study.diagnostic_steps.map((step: number) => <option value={String(step)} selected={step === 8000}>STEP {step.toLocaleString("en-US")}</option>)}
</select>
</label>
</div>
<div class="chart-shell">
<header><span>POST-MLP OUTPUT RMS</span><b data-output-copy>DEPTH 32 · 3-SEED MEAN · STEP 8,000</b></header>
<svg data-output-chart viewBox="0 0 920 360" role="img" aria-label="Baseline 与 Block 的逐层 output RMS">
<g data-output-groups></g>
<g data-output-grid></g>
<polyline data-output-line="baseline" class="series baseline"></polyline>
<polyline data-output-line="block" class="series block"></polyline>
<g data-output-points="baseline"></g>
<g data-output-points="block"></g>
<text x="460" y="350" class="axis-title">TRANSFORMER BLOCK INDEX</text>
</svg>
<div class="chart-legend"><span><i class="baseline"></i>Baseline</span><span><i class="block"></i>Block AttnRes</span><span><i class="boundary"></i>aggregation boundary</span></div>
</div>
<div class="output-ratios">
<article><span>DEPTH 16 / LAST ÷ FIRST</span><div><b>Baseline</b><strong>4.59×</strong></div><div><b>Block</b><strong>1.17×</strong></div></article>
<article><span>DEPTH 32 / LAST ÷ FIRST</span><div><b>Baseline</b><strong>6.08×</strong></div><div><b>Block</b><strong>1.89×</strong></div></article>
<article class="accent"><span>WHAT THE CURVE SAYS</span><b>全局累积 → 组内锯齿</b><p>Block 限制 output magnitude 持续跨深度增长;这一方向与 Figure 5(b) 叙述一致。</p></article>
</div>
<div class="group-rhythm">
<header><span data-rhythm-title>BLOCK ATTNRES · DEPTH 32</span><b>8 AGGREGATION GROUPS</b></header>
<div data-output-bars aria-label="Block AttnRes output RMS 组内节律"></div>
<p>粗分隔线是 group boundary;柱高来自当前选择的 checkpoint / seed,不是示意动画。</p>
</div>
</section>
<section class="gradient-panel" data-gradient-panel="verdict" hidden>
<div class="panel-lead">
<div><span>V / JOINT VERDICT</span><h4>一个正结果、两个反结果和一张成本账,要同时摆在桌面上</h4></div>
<p>Round 05 的主判定只读 activation CV + imbalance;BPC、参数梯度与成本是必须公开的次要结果。</p>
</div>
<div class="verdict-table-wrap">
<table class="verdict-table">
<thead><tr><th>Depth</th><th>Δ BPC</th><th>Activation CV</th><th>First/last imbalance</th><th>Mean grad scale</th><th>Parameter CV</th><th>Verdict</th></tr></thead>
<tbody>
<tr>
<th>16</th>
<td>{formatSigned(depth16.means.block_minus_baseline_bpc, 5)}</td>
<td class="bad">{pct(depth16.means.relative_cv_reduction)} reduction</td>
<td class="good">+{pct(depth16.means.relative_imbalance_reduction)}</td>
<td>{pct(depth16.means.block_to_baseline_activation_grad_mean)} of Base</td>
<td>{depth16.means.baseline_parameter_grad_cv.toFixed(3)} → {depth16.means.block_parameter_grad_cv.toFixed(3)}</td>
<td>mixed</td>
</tr>
<tr>
<th>32</th>
<td>{formatSigned(depth32.means.block_minus_baseline_bpc, 5)}</td>
<td class="bad">{pct(depth32.means.relative_cv_reduction)} reduction</td>
<td class="good">+{pct(depth32.means.relative_imbalance_reduction)}</td>
<td>{pct(depth32.means.block_to_baseline_activation_grad_mean)} of Base</td>
<td>{depth32.means.baseline_parameter_grad_cv.toFixed(3)} → {depth32.means.block_parameter_grad_cv.toFixed(3)}</td>
<td>mixed</td>
</tr>
</tbody>
</table>
</div>
<div class="metric-pairs">
<article><span>ACTIVATION / FIRST-LAST</span><b>Block 改善</b><p>6 / 6 配对更接近 1。</p></article>
<article class="warn"><span>ACTIVATION / ALL-LAYER CV</span><b>Block 恶化</b><p>6 / 6 配对 CV 更高。</p></article>
<article class="warn"><span>PARAMETER / ALL-LAYER CV</span><b>Block 恶化</b><p>0.416→0.683;0.397→0.772。</p></article>
<article class="pass"><span>VALIDATION BPC</span><b>Block 更低</b><p>6 / 6 配对为负;非同算力。</p></article>
</div>
<div class="cost-compare">
<article><span>DEPTH 16</span><div><b>STEP TIME</b><strong>21.47 → 54.71 ms</strong></div><div><b>PEAK ALLOC</b><strong>3.04 → 6.54 GB</strong></div><p>约 2.55× time · 2.15× memory</p></article>
<article><span>DEPTH 32</span><div><b>STEP TIME</b><strong>42.11 → 109.38 ms</strong></div><div><b>PEAK ALLOC</b><strong>5.94 → 12.82 GB</strong></div><p>约 2.60× time · 2.16× memory</p></article>
<article class="boundary"><span>CLAIM BOUNDARY</span><b>同 token / step</b><p>不能写成同 FLOPs、同 wall time,或 K3 生产成本。</p></article>
</div>
<div class="replay-ledger">
<article><span>FORMAL</span><b>depth-32 · Block · seed-01</b><p>8,000 steps from initialization</p></article>
<i>≡</i>
<article><span>FRESH REPLAY</span><b>all frozen fields exact</b><p>额外 65,536,000 target bytes</p></article>
<i>→</i>
<article class="pass"><span>STATE HASH</span><b>model + optimizer exact</b><p>{shortHash(lab.cells.find((cell: any) => cell.depth === 32 && cell.architecture === "block" && cell.seed === 2026073001).hashes.final_model_state)}</p></article>
</div>
<div class="hash-ledger">
<article><span>MANIFEST</span><code>{lab.manifest_summary.file_sha256}</code></article>
<article><span>SCHEDULE</span><code>{lab.manifest_summary.formal_schedule_sha256}</code></article>
<article><span>COMPACT CANONICAL</span><code>{lab.canonical_sha256_without_self}</code></article>
<article><span>OVERALL</span><code>{lab.overall_verdict}</code></article>
</div>
<div class="claim-grid">
<article class="yes"><span>THIS STUDY SUPPORTS</span><ul><li>在公开定义下,Block 改写了梯度失衡的形态。</li><li>早/晚深度更接近,但中后段局部峰更尖。</li><li>Output RMS 的跨深度增长显著受限。</li></ul></article>
<article class="no"><span>THIS STUDY DOES NOT SUPPORT</span><ul><li>复现论文 Figure 5(c) 的数值或隐藏实现。</li><li>测到 K3 checkpoint 的真实梯度。</li><li>证明 AttnRes 一般更稳定或更省算力。</li></ul></article>
</div>
</section>
<script is:inline type="application/json" data-gradient-payload set:html={json}></script>
</figure>
<script>
const initializeGradientLab = (root: HTMLElement) => {
if (root.dataset.gradientReady === "true") return;
root.dataset.gradientReady = "true";
const payload = root.querySelector<HTMLScriptElement>("[data-gradient-payload]");
if (!payload?.textContent) return;
const data = JSON.parse(payload.textContent);
const architectures = ["baseline", "block"];
const steps = data.study.diagnostic_steps;
const seeds = data.study.seeds;
const ns = "http://www.w3.org/2000/svg";
const colors: Record<string, string> = { baseline: "#77746b", block: "#ba603b" };
const tabs = [...root.querySelectorAll<HTMLButtonElement>("[data-gradient-tab]")];
const panels = [...root.querySelectorAll<HTMLElement>("[data-gradient-panel]")];
const selectTab = (id: string) => {
tabs.forEach((tab) => tab.setAttribute("aria-selected", String(tab.dataset.gradientTab === id)));
panels.forEach((panel) => { panel.hidden = panel.dataset.gradientPanel !== id; });
};
tabs.forEach((tab, index) => {
tab.addEventListener("click", () => selectTab(tab.dataset.gradientTab || "definition"));
tab.addEventListener("keydown", (event) => {
if (!["ArrowLeft", "ArrowRight", "Home", "End"].includes(event.key)) return;
event.preventDefault();
const nextIndex = event.key === "Home" ? 0
: event.key === "End" ? tabs.length - 1
: (index + (event.key === "ArrowRight" ? 1 : -1) + tabs.length) % tabs.length;
tabs[nextIndex].focus();
selectTab(tabs[nextIndex].dataset.gradientTab || "definition");
});
});
const findCell = (depth: number, architecture: string, seed: number) =>
data.cells.find((cell: any) => cell.depth === depth && cell.architecture === architecture && cell.seed === seed);
const findDiagnostic = (cell: any, step: number) =>
cell.diagnostics.find((row: any) => row.step === step);
const seedList = (seedKey: string) => seedKey === "mean" ? seeds : [Number(seedKey)];
const average = (values: number[]) => values.reduce((sum, value) => sum + value, 0) / values.length;
const arrayFor = (depth: number, architecture: string, seedKey: string, step: number, field: string) => {
const rows = seedList(seedKey).map((seed: number) => findDiagnostic(findCell(depth, architecture, seed), step)[field]);
return rows[0].map((_: number, index: number) => average(rows.map((row: number[]) => row[index])));
};
const statFor = (depth: number, architecture: string, seedKey: string, step: number, group: string, field: string) =>
average(seedList(seedKey).map((seed: number) => findDiagnostic(findCell(depth, architecture, seed), step)[group][field]));
const formatStep = (value: number) => value.toLocaleString("en-US");
const formatScientific = (value: number) => value.toExponential(2).replace("e-", "e−");
const svgNode = (name: string, attrs: Record<string, string | number>) => {
const node = document.createElementNS(ns, name);
Object.entries(attrs).forEach(([key, value]) => node.setAttribute(key, String(value)));
return node;
};
const drawGroups = (group: SVGGElement, depth: number) => {
group.replaceChildren();
for (let index = 1; index < 8; index += 1) {
const block = depth / 8 * index;
const x = 58 + block / (depth - 1) * 822;
group.append(svgNode("line", { x1: x, x2: x, y1: 24, y2: 304, class: "group-boundary" }));
}
};
const drawChart = (
svg: SVGSVGElement,
groups: SVGGElement | null,
grid: SVGGElement,
lines: Record<string, SVGPolylineElement | null>,
pointGroups: Record<string, SVGGElement | null>,
series: Record<string, number[]>,
labels: string[],
options: { depth?: number; formatter?: (value: number) => string; zero?: boolean } = {},
) => {
grid.replaceChildren();
const values = Object.values(series).flat();
const rawMin = options.zero === false ? Math.min(...values) : 0;
const rawMax = Math.max(...values);
const padding = rawMax === rawMin ? 1 : (rawMax - rawMin) * 0.08;
const minimum = Math.max(0, rawMin - padding);
const maximum = rawMax + padding;
const xAt = (index: number) => 58 + index / Math.max(1, labels.length - 1) * 822;
const yAt = (value: number) => 24 + (maximum - value) / Math.max(1e-30, maximum - minimum) * 280;
const formatter = options.formatter || ((value: number) => value.toFixed(2));
for (let index = 0; index < 5; index += 1) {
const value = minimum + (maximum - minimum) * (4 - index) / 4;
const y = 24 + index / 4 * 280;
grid.append(svgNode("line", { x1: 58, x2: 880, y1: y, y2: y, class: "grid-line" }));
const text = svgNode("text", { x: 48, y: y + 4, class: "axis-label y" });
text.textContent = formatter(value);
grid.append(text);
}
labels.forEach((label, index) => {
const text = svgNode("text", { x: xAt(index), y: 326, class: "axis-label x" });
text.textContent = label;
grid.append(text);
});
if (groups && options.depth) drawGroups(groups, options.depth);
architectures.forEach((architecture) => {
const points = series[architecture].map((value, index) => `${xAt(index)},${yAt(value)}`).join(" ");
lines[architecture]?.setAttribute("points", points);
pointGroups[architecture]?.replaceChildren();
series[architecture].forEach((value, index) => {
const point = svgNode("circle", {
cx: xAt(index), cy: yAt(value), r: 3.6, fill: colors[architecture],
});
const title = svgNode("title", {});
title.textContent = `${architecture === "baseline" ? "Baseline" : "Block"} · ${labels[index]} · ${formatter(value)}`;
point.append(title);
pointGroups[architecture]?.append(point);
});
});
svg.dataset.maximum = String(maximum);
};
let spectrumDepth = 32;
let spectrumScale = "absolute";
const spectrumSeed = root.querySelector<HTMLSelectElement>("[data-spectrum-seed]")!;
const spectrumStep = root.querySelector<HTMLSelectElement>("[data-spectrum-step]")!;
const spectrumChart = root.querySelector<SVGSVGElement>("[data-spectrum-chart]")!;
const updateSpectrum = () => {
const seedKey = spectrumSeed.value;
const step = Number(spectrumStep.value);
const absolute = Object.fromEntries(architectures.map((architecture) => [
architecture,
arrayFor(spectrumDepth, architecture, seedKey, step, "activation_grad_rms_by_block"),
]));
const series = spectrumScale === "normalized"
? Object.fromEntries(architectures.map((architecture) => {
const mean = average(absolute[architecture]);
return [architecture, absolute[architecture].map((value: number) => value / mean)];
}))
: absolute;
drawChart(
spectrumChart,
spectrumChart.querySelector("[data-chart-groups]"),
spectrumChart.querySelector("[data-chart-grid]")!,
Object.fromEntries(architectures.map((architecture) => [architecture, spectrumChart.querySelector(`[data-chart-line="${architecture}"]`)])),
Object.fromEntries(architectures.map((architecture) => [architecture, spectrumChart.querySelector(`[data-chart-points="${architecture}"]`)])),
series,
Array.from({ length: spectrumDepth }, (_, index) => String(index + 1)),
{
depth: spectrumDepth,
formatter: spectrumScale === "absolute" ? formatScientific : (value) => `${value.toFixed(1)}×`,
},
);
const title = root.querySelector<HTMLElement>("[data-spectrum-title]");
if (title) title.textContent = spectrumScale === "absolute" ? "ACTIVATION GRADIENT RMS" : "NORMALIZED BY LAYER MEAN";
const state = root.querySelector<HTMLElement>("[data-spectrum-state]");
if (state) state.textContent = `DEPTH ${spectrumDepth} · ${seedKey === "mean" ? "3-SEED MEAN" : seedKey} · STEP ${formatStep(step)}`;
architectures.forEach((architecture) => {
const cv = statFor(spectrumDepth, architecture, seedKey, step, "activation_grad_statistics", "population_cv");
const ratio = statFor(spectrumDepth, architecture, seedKey, step, "activation_grad_statistics", "first_to_last_ratio");
const mean = statFor(spectrumDepth, architecture, seedKey, step, "activation_grad_statistics", "mean");
const prefix = architecture === "baseline" ? "base" : "block";
root.querySelector<HTMLElement>(`[data-spectrum-${prefix}-cv]`)!.textContent = cv.toFixed(4);
root.querySelector<HTMLElement>(`[data-spectrum-${prefix}-ratio]`)!.textContent = `${ratio.toFixed(2)}×`;
root.querySelector<HTMLElement>(`[data-spectrum-${prefix}-mean]`)!.textContent = formatScientific(mean);
});
};
root.querySelectorAll<HTMLButtonElement>("[data-spectrum-depth]").forEach((button) => button.addEventListener("click", () => {
spectrumDepth = Number(button.dataset.spectrumDepth);
root.querySelectorAll<HTMLButtonElement>("[data-spectrum-depth]").forEach((peer) => peer.setAttribute("aria-pressed", String(peer === button)));
updateSpectrum();
}));
root.querySelectorAll<HTMLButtonElement>("[data-spectrum-scale]").forEach((button) => button.addEventListener("click", () => {
spectrumScale = button.dataset.spectrumScale || "absolute";
root.querySelectorAll<HTMLButtonElement>("[data-spectrum-scale]").forEach((peer) => peer.setAttribute("aria-pressed", String(peer === button)));
updateSpectrum();
}));
spectrumSeed.addEventListener("change", updateSpectrum);
spectrumStep.addEventListener("change", updateSpectrum);
updateSpectrum();
let timeDepth = 32;
let timeMetric = "cv";
const timeSeed = root.querySelector<HTMLSelectElement>("[data-time-seed]")!;
const timeChart = root.querySelector<SVGSVGElement>("[data-time-chart]")!;
const timeMetricMap: Record<string, [string, string, (value: number) => string]> = {
cv: ["activation_grad_statistics", "population_cv", (value) => value.toFixed(2)],
imbalance: ["activation_grad_statistics", "imbalance_abs_log_ratio", (value) => value.toFixed(2)],
mean: ["activation_grad_statistics", "mean", formatScientific],
};
const updateTimeline = () => {
const seedKey = timeSeed.value;
const [group, field, formatter] = timeMetricMap[timeMetric];
const series = Object.fromEntries(architectures.map((architecture) => [
architecture,
steps.map((step: number) => statFor(timeDepth, architecture, seedKey, step, group, field)),
]));
drawChart(
timeChart,
null,
timeChart.querySelector("[data-time-grid]")!,
Object.fromEntries(architectures.map((architecture) => [architecture, timeChart.querySelector(`[data-time-line="${architecture}"]`)])),
Object.fromEntries(architectures.map((architecture) => [architecture, timeChart.querySelector(`[data-time-points="${architecture}"]`)])),
series,
steps.map((step: number) => step >= 1000 ? `${step / 1000}K` : String(step)),
{ formatter },
);
const titles: Record<string, string> = {
cv: "POPULATION CV",
imbalance: "ABS(LOG(FIRST / LAST QUARTILE))",
mean: "MEAN ACTIVATION GRADIENT RMS",
};
root.querySelector<HTMLElement>("[data-time-title]")!.textContent = titles[timeMetric];
root.querySelector<HTMLElement>("[data-time-copy]")!.textContent = `DEPTH ${timeDepth} · ${seedKey === "mean" ? "3-SEED MEAN" : seedKey}`;
};
root.querySelectorAll<HTMLButtonElement>("[data-time-depth]").forEach((button) => button.addEventListener("click", () => {
timeDepth = Number(button.dataset.timeDepth);
root.querySelectorAll<HTMLButtonElement>("[data-time-depth]").forEach((peer) => peer.setAttribute("aria-pressed", String(peer === button)));
updateTimeline();
}));
root.querySelectorAll<HTMLButtonElement>("[data-time-metric]").forEach((button) => button.addEventListener("click", () => {
timeMetric = button.dataset.timeMetric || "cv";
root.querySelectorAll<HTMLButtonElement>("[data-time-metric]").forEach((peer) => peer.setAttribute("aria-pressed", String(peer === button)));
updateTimeline();
}));
timeSeed.addEventListener("change", updateTimeline);
updateTimeline();
let outputDepth = 32;
const outputSeed = root.querySelector<HTMLSelectElement>("[data-output-seed]")!;
const outputStep = root.querySelector<HTMLSelectElement>("[data-output-step]")!;
const outputChart = root.querySelector<SVGSVGElement>("[data-output-chart]")!;
const updateOutput = () => {
const seedKey = outputSeed.value;
const step = Number(outputStep.value);
const series = Object.fromEntries(architectures.map((architecture) => [
architecture,
arrayFor(outputDepth, architecture, seedKey, step, "activation_output_rms_by_block"),
]));
drawChart(
outputChart,
outputChart.querySelector("[data-output-groups]"),
outputChart.querySelector("[data-output-grid]")!,
Object.fromEntries(architectures.map((architecture) => [architecture, outputChart.querySelector(`[data-output-line="${architecture}"]`)])),
Object.fromEntries(architectures.map((architecture) => [architecture, outputChart.querySelector(`[data-output-points="${architecture}"]`)])),
series,
Array.from({ length: outputDepth }, (_, index) => String(index + 1)),
{ depth: outputDepth, formatter: (value) => value.toFixed(2) },
);
root.querySelector<HTMLElement>("[data-output-copy]")!.textContent = `DEPTH ${outputDepth} · ${seedKey === "mean" ? "3-SEED MEAN" : seedKey} · STEP ${formatStep(step)}`;
root.querySelector<HTMLElement>("[data-rhythm-title]")!.textContent = `BLOCK ATTNRES · DEPTH ${outputDepth}`;
const bars = root.querySelector<HTMLElement>("[data-output-bars]")!;
const values: number[] = series.block;
const maximum = Math.max(...values);
const blocksPerGroup = outputDepth / 8;
bars.replaceChildren();
values.forEach((value, index) => {
const bar = document.createElement("i");
bar.style.setProperty("--bar", `${value / maximum * 100}%`);
if ((index + 1) % blocksPerGroup === 0) bar.classList.add("boundary");
const label = document.createElement("span");
label.textContent = String(index + 1);
bar.append(label);
bars.append(bar);
});
};
root.querySelectorAll<HTMLButtonElement>("[data-output-depth]").forEach((button) => button.addEventListener("click", () => {
outputDepth = Number(button.dataset.outputDepth);
root.querySelectorAll<HTMLButtonElement>("[data-output-depth]").forEach((peer) => peer.setAttribute("aria-pressed", String(peer === button)));
updateOutput();
}));
outputSeed.addEventListener("change", updateOutput);
outputStep.addEventListener("change", updateOutput);
updateOutput();
};
document.querySelectorAll<HTMLElement>("[data-gradient-lab]").forEach(initializeGradientLab);
document.addEventListener("astro:page-load", () => {
document.querySelectorAll<HTMLElement>("[data-gradient-lab]").forEach(initializeGradientLab);
});
</script>
<style>
.gradient-lab {
--g-ink: #1c201e;
--g-muted: #77746b;
--g-line: rgba(28, 32, 30, .16);
--g-paper: #f4f0e7;
--g-raised: #faf7ef;
--g-copper: #ba603b;
--g-green: #163f3b;
width: min(1120px, 100%);
margin: 42px 0;
color: var(--g-ink);
border: 1px solid var(--g-line);
background: var(--g-paper);
box-shadow: 0 30px 80px rgba(28, 32, 30, .09);
}
.gradient-head {
display: grid;
grid-template-columns: minmax(0, 1.45fr) minmax(260px, .7fr);
gap: 44px;
padding: 30px;
color: #f5efe4;
background: var(--g-green);
}
.gradient-head p { margin: 0; color: rgba(245,239,228,.7); font: .65rem/1.7 var(--mono); }
.gradient-head div > p { color: #d58a68; letter-spacing: .08em; }
.gradient-head h3 { max-width: 720px; margin: 14px 0 0; color: inherit; font-size: clamp(1.15rem, 2.2vw, 1.75rem); line-height: 1.35; }
.gradient-ledger { display: grid; grid-template-columns: repeat(6, 1fr); border-bottom: 1px solid var(--g-line); }
.gradient-ledger article { min-height: 126px; padding: 18px 15px; border-right: 1px solid var(--g-line); }
.gradient-ledger article:last-child { border-right: 0; }
.gradient-ledger span, .panel-lead span { color: var(--g-muted); font: .56rem/1.2 var(--mono); letter-spacing: .08em; }
.gradient-ledger b { display: block; margin-top: 23px; font: 700 .95rem/1 var(--mono); }
.gradient-ledger p { margin: 8px 0 0; color: var(--g-muted); font-size: .6rem; line-height: 1.45; }
.gradient-ledger .pass { color: #f7f0e6; background: var(--g-green); }
.gradient-ledger .split { color: #f7f0e6; background: var(--g-copper); }
.gradient-ledger .pass span, .gradient-ledger .pass p, .gradient-ledger .split span, .gradient-ledger .split p { color: rgba(247,240,230,.72); }
.gradient-tabs { display: grid; grid-template-columns: repeat(5, 1fr); border-bottom: 1px solid var(--g-line); background: #e9e4da; }
.gradient-tabs button { min-height: 116px; padding: 16px; text-align: left; color: inherit; border: 0; border-right: 1px solid var(--g-line); background: transparent; cursor: pointer; }
.gradient-tabs button:last-child { border-right: 0; }
.gradient-tabs button[aria-selected="true"] { color: #f7f0e6; background: var(--g-copper); }
.gradient-tabs span, .gradient-tabs small { display: block; color: var(--g-muted); font: .54rem/1.25 var(--mono); }
.gradient-tabs b { display: block; margin: 15px 0 8px; font-size: .69rem; line-height: 1.35; }
.gradient-tabs button[aria-selected="true"] span, .gradient-tabs button[aria-selected="true"] small { color: rgba(247,240,230,.72); }
.gradient-panel { padding: 30px; }
.panel-lead { display: grid; grid-template-columns: 1.05fr .95fr; gap: 48px; align-items: end; margin-bottom: 28px; }
.panel-lead h4 { max-width: 680px; margin: 10px 0 0; font-size: 1.2rem; line-height: 1.4; }
.panel-lead p { margin: 0; color: var(--g-muted); font-size: .7rem; line-height: 1.7; }
.known-grid { display: grid; grid-template-columns: repeat(2, 1fr); border: 1px solid var(--g-line); }
.known-grid article { min-height: 270px; padding: 24px; }
.known-grid article + article { border-left: 1px solid var(--g-line); background: #ece2d6; }
.known-grid span, .object-chain span, .object-compare span, .split-result span, .timeline-notes span, .output-ratios span, .metric-pairs span, .cost-compare > article > span, .replay-ledger span, .hash-ledger span, .claim-grid span {
color: var(--g-copper); font: .56rem/1 var(--mono); letter-spacing: .06em;
}
.known-grid h5 { margin: 26px 0 16px; font-size: .95rem; }
.known-grid ul { margin: 0; padding-left: 18px; }
.known-grid li { margin-top: 11px; color: var(--g-muted); font-size: .68rem; line-height: 1.55; }
.object-chain { display: grid; grid-template-columns: 1fr 30px 1fr 30px .8fr 30px 1.2fr; gap: 6px; align-items: center; margin-top: 20px; }
.object-chain div { min-height: 145px; padding: 18px; border: 1px solid var(--g-line); background: var(--g-raised); }
.object-chain .accent { color: #f7f0e6; background: var(--g-green); }
.object-chain .accent span, .object-chain .accent p { color: rgba(247,240,230,.68); }
.object-chain b { display: block; margin-top: 25px; font: 700 .77rem/1.35 var(--mono); }
.object-chain p { color: var(--g-muted); font-size: .61rem; line-height: 1.45; }
.object-chain > i, .object-compare > i, .split-result > i, .replay-ledger > i { color: var(--g-copper); font-style: normal; text-align: center; }
.object-compare { display: grid; grid-template-columns: 1fr 50px 1fr; gap: 12px; align-items: center; margin-top: 20px; padding: 20px; background: #e9e4da; }
.object-compare article { padding: 12px; }
.object-compare b { display: block; margin-top: 18px; font: 700 1.1rem/1 var(--mono); }
.object-compare p { color: var(--g-muted); font-size: .65rem; line-height: 1.55; }
.object-compare > p { grid-column: 1/-1; margin: 0; padding-top: 16px; border-top: 1px solid var(--g-line); }
.definition-boundary { display: grid; grid-template-columns: 80px 1fr; margin-top: 20px; border-top: 1px solid var(--g-line); }
.definition-boundary > * { margin: 0; padding: 15px; border-right: 1px solid var(--g-line); border-bottom: 1px solid var(--g-line); }
.definition-boundary b { color: var(--g-copper); font: .6rem/1.4 var(--mono); }
.definition-boundary p { color: var(--g-muted); font-size: .66rem; line-height: 1.55; }
.lab-controls { display: flex; flex-wrap: wrap; gap: 10px 18px; align-items: end; margin-bottom: 20px; }
.lab-controls > div { display: flex; }
.lab-controls button, .lab-controls select { min-height: 38px; padding: 10px 12px; color: var(--g-muted); font: 700 .56rem/1 var(--mono); border: 1px solid var(--g-line); background: var(--g-raised); }
.lab-controls button { cursor: pointer; }
.lab-controls button + button { border-left: 0; }
.lab-controls button[aria-pressed="true"] { color: #fff9ef; background: var(--g-green); }
.lab-controls label { display: grid; gap: 6px; color: var(--g-muted); font: .52rem/1 var(--mono); }
.spectrum-layout { display: grid; grid-template-columns: minmax(0, 1fr) 235px; border: 1px solid var(--g-line); background: var(--g-raised); }
.chart-shell { min-width: 0; padding: 18px; border: 1px solid var(--g-line); background: var(--g-raised); }
.spectrum-layout .chart-shell { border: 0; border-right: 1px solid var(--g-line); }
.chart-shell header { display: flex; justify-content: space-between; gap: 12px; color: var(--g-muted); font: .55rem/1 var(--mono); }
.chart-shell svg { display: block; width: 100%; height: auto; margin-top: 10px; overflow: visible; }
.series { fill: none; stroke-width: 3; stroke-linejoin: round; stroke-linecap: round; }
.series.baseline { stroke: #77746b; }
.series.block { stroke: #ba603b; }
.grid-line { stroke: rgba(28,32,30,.1); stroke-width: 1; }
.group-boundary { stroke: rgba(22,63,59,.22); stroke-width: 1.5; stroke-dasharray: 4 4; }
.axis-label { fill: #8b867c; font: 11px var(--mono); }
.axis-label.y { text-anchor: end; }
.axis-label.x { text-anchor: middle; }
.axis-title { fill: #8b867c; font: 11px var(--mono); text-anchor: middle; }
.chart-legend { display: flex; flex-wrap: wrap; gap: 18px; margin-top: 4px; color: var(--g-muted); font: .56rem/1 var(--mono); }
.chart-legend span { display: inline-flex; gap: 7px; align-items: center; }
.chart-legend i { width: 22px; height: 3px; }
.chart-legend i.baseline { background: #77746b; }
.chart-legend i.block { background: #ba603b; }
.chart-legend i.boundary { height: 0; border-top: 2px dashed var(--g-green); background: transparent; }
.spectrum-readout { padding: 20px 18px; }
.spectrum-readout > span { color: var(--g-copper); font: .54rem/1.4 var(--mono); }
.spectrum-readout article { padding: 18px 0; border-bottom: 1px solid var(--g-line); }
.spectrum-readout article b { color: var(--g-muted); font: .52rem/1 var(--mono); }
.spectrum-readout article div { display: flex; justify-content: space-between; margin-top: 13px; }
.spectrum-readout article span { color: var(--g-muted); font: .5rem/1 var(--mono); }
.spectrum-readout article strong { font: 700 .7rem/1 var(--mono); }
.split-result { display: grid; grid-template-columns: 1fr 45px 1fr 45px 1fr; gap: 8px; align-items: center; margin-top: 20px; }
.split-result article { min-height: 145px; padding: 20px; border: 1px solid var(--g-line); }
.split-result .accent { color: #f7f0e6; background: var(--g-green); }
.split-result .accent span, .split-result .accent p { color: rgba(247,240,230,.68); }
.split-result b { display: block; margin-top: 25px; font-size: .83rem; }
.split-result p { color: var(--g-muted); font-size: .62rem; line-height: 1.5; }
.timeline-chart { margin-top: 0; }
.timeline-notes { display: grid; grid-template-columns: repeat(4, 1fr); margin-top: 20px; border-top: 1px solid var(--g-line); border-left: 1px solid var(--g-line); }
.timeline-notes article { min-height: 145px; padding: 18px; border-right: 1px solid var(--g-line); border-bottom: 1px solid var(--g-line); background: var(--g-raised); }
.timeline-notes .counter { background: #e9e4da; }
.timeline-notes b { display: block; margin-top: 25px; font: 700 .78rem/1 var(--mono); }
.timeline-notes p { color: var(--g-muted); font-size: .61rem; line-height: 1.5; }
.output-ratios { display: grid; grid-template-columns: 1fr 1fr 1.25fr; margin-top: 20px; border-top: 1px solid var(--g-line); border-left: 1px solid var(--g-line); }
.output-ratios article { min-height: 170px; padding: 20px; border-right: 1px solid var(--g-line); border-bottom: 1px solid var(--g-line); }
.output-ratios article > div { display: flex; justify-content: space-between; margin-top: 24px; }
.output-ratios article > div b { color: var(--g-muted); font: .55rem/1 var(--mono); }
.output-ratios article > div strong { font: 700 .75rem/1 var(--mono); }
.output-ratios .accent { color: #f7f0e6; background: var(--g-green); }
.output-ratios .accent span, .output-ratios .accent p { color: rgba(247,240,230,.68); }
.output-ratios .accent b { display: block; margin-top: 25px; font-size: .8rem; }
.output-ratios p { color: var(--g-muted); font-size: .62rem; line-height: 1.5; }
.group-rhythm { margin-top: 20px; padding: 20px; color: #f7f0e6; background: var(--g-green); overflow: hidden; }
.group-rhythm header { display: flex; justify-content: space-between; color: rgba(247,240,230,.68); font: .55rem/1 var(--mono); }
.group-rhythm > div { display: grid; grid-template-columns: repeat(auto-fit, minmax(8px, 1fr)); align-items: end; height: 170px; margin-top: 18px; border-bottom: 1px solid rgba(255,255,255,.3); }
.group-rhythm i { position: relative; display: block; height: max(5px, var(--bar)); margin-right: 2px; background: #d4835d; }
.group-rhythm i.boundary { margin-right: 8px; border-right: 2px solid rgba(255,255,255,.75); }
.group-rhythm i span { position: absolute; bottom: -18px; left: 50%; color: rgba(255,255,255,.55); font: .43rem/1 var(--mono); transform: translateX(-50%); }
.group-rhythm p { margin: 35px 0 0; color: rgba(247,240,230,.72); font-size: .64rem; }
.verdict-table-wrap { overflow-x: auto; }
.verdict-table { width: 100%; min-width: 850px; border-collapse: collapse; font-size: .62rem; }
.verdict-table th, .verdict-table td { padding: 15px 12px; border-bottom: 1px solid var(--g-line); text-align: left; }
.verdict-table thead th { color: var(--g-muted); font: .53rem/1.3 var(--mono); }
.verdict-table tbody th, .verdict-table tbody td:last-child { font: 700 .66rem/1 var(--mono); }
.verdict-table .good { color: var(--g-green); }
.verdict-table .bad { color: var(--g-copper); }
.metric-pairs { display: grid; grid-template-columns: repeat(4, 1fr); margin-top: 20px; border-top: 1px solid var(--g-line); border-left: 1px solid var(--g-line); }
.metric-pairs article { min-height: 150px; padding: 18px; border-right: 1px solid var(--g-line); border-bottom: 1px solid var(--g-line); background: var(--g-raised); }
.metric-pairs article.warn { border-top: 4px solid var(--g-copper); }
.metric-pairs article.pass { border-top: 4px solid var(--g-green); }
.metric-pairs b { display: block; margin-top: 25px; font-size: .77rem; }
.metric-pairs p { color: var(--g-muted); font-size: .61rem; line-height: 1.5; }
.cost-compare { display: grid; grid-template-columns: repeat(3, 1fr); margin-top: 20px; border-top: 1px solid var(--g-line); border-left: 1px solid var(--g-line); }
.cost-compare article { min-height: 205px; padding: 20px; border-right: 1px solid var(--g-line); border-bottom: 1px solid var(--g-line); }
.cost-compare article > div { display: flex; justify-content: space-between; gap: 12px; margin-top: 25px; }
.cost-compare article > div b { color: var(--g-muted); font: .52rem/1 var(--mono); }
.cost-compare article > div strong { font: 700 .62rem/1 var(--mono); text-align: right; }
.cost-compare p { color: var(--g-copper); font: .56rem/1.5 var(--mono); }
.cost-compare .boundary { color: #f7f0e6; background: var(--g-green); }
.cost-compare .boundary span, .cost-compare .boundary p { color: rgba(247,240,230,.68); }
.cost-compare .boundary b { display: block; margin-top: 30px; font-size: .85rem; }
.replay-ledger { display: grid; grid-template-columns: 1fr 40px 1fr 40px 1fr; gap: 8px; align-items: center; margin-top: 20px; }
.replay-ledger article { min-height: 145px; padding: 18px; border: 1px solid var(--g-line); background: var(--g-raised); }
.replay-ledger article.pass { color: #f7f0e6; background: var(--g-green); }
.replay-ledger article.pass span, .replay-ledger article.pass p { color: rgba(247,240,230,.68); }
.replay-ledger b { display: block; margin-top: 25px; font-size: .7rem; }
.replay-ledger p { color: var(--g-muted); font-size: .58rem; line-height: 1.5; overflow-wrap: anywhere; }
.hash-ledger { display: grid; grid-template-columns: repeat(2, 1fr); margin-top: 20px; border-top: 1px solid var(--g-line); border-left: 1px solid var(--g-line); }
.hash-ledger article { min-width: 0; padding: 16px; border-right: 1px solid var(--g-line); border-bottom: 1px solid var(--g-line); background: #e9e4da; }
.hash-ledger code { display: block; margin-top: 12px; overflow: hidden; color: var(--g-green); font: .54rem/1.3 var(--mono); text-overflow: ellipsis; }
.claim-grid { display: grid; grid-template-columns: repeat(2, 1fr); margin-top: 20px; }
.claim-grid article { padding: 22px; }
.claim-grid .yes { color: #f7f0e6; background: var(--g-green); }
.claim-grid .no { background: #e7d8ca; }
.claim-grid ul { margin: 18px 0 0; padding-left: 17px; }
.claim-grid li { margin-top: 10px; font-size: .65rem; line-height: 1.55; }
.claim-grid .yes span, .claim-grid .yes li { color: rgba(247,240,230,.78); }
@media (max-width: 920px) {
.gradient-ledger { grid-template-columns: repeat(3, 1fr); }
.gradient-ledger article:nth-child(3) { border-right: 0; }
.gradient-tabs { grid-template-columns: repeat(3, 1fr); }
.spectrum-layout { grid-template-columns: 1fr; }
.spectrum-layout .chart-shell { border-right: 0; border-bottom: 1px solid var(--g-line); }
.timeline-notes, .metric-pairs { grid-template-columns: repeat(2, 1fr); }
.object-chain { grid-template-columns: 1fr 24px 1fr; }
.object-chain > i:nth-of-type(n+3) { display: none; }
}
@media (max-width: 680px) {
.gradient-head, .panel-lead, .known-grid { grid-template-columns: 1fr; gap: 20px; }
.gradient-head, .gradient-panel { padding: 20px; }
.known-grid article + article { border-left: 0; border-top: 1px solid var(--g-line); }
.gradient-ledger { grid-template-columns: repeat(2, 1fr); }
.gradient-ledger article:nth-child(3) { border-right: 1px solid var(--g-line); }
.gradient-ledger article:nth-child(even) { border-right: 0; }
.gradient-tabs { display: flex; overflow-x: auto; }
.gradient-tabs button { flex: 0 0 190px; }
.lab-controls { align-items: stretch; }
.lab-controls > div, .lab-controls label { flex: 1 0 100%; }
.lab-controls button { flex: 1; }
.lab-controls select { width: 100%; }
.object-chain, .object-compare, .split-result, .replay-ledger { grid-template-columns: 1fr; }
.object-chain > i, .object-compare > i, .split-result > i, .replay-ledger > i { display: block !important; transform: rotate(90deg); }
.object-compare > p { grid-column: 1; }
.definition-boundary { grid-template-columns: 70px 1fr; }
.timeline-notes, .output-ratios, .metric-pairs, .cost-compare, .hash-ledger, .claim-grid { grid-template-columns: 1fr; }
.chart-shell { padding: 12px 8px; }
.chart-shell header { padding: 0 8px; }
.axis-label { font-size: 9px; }
.group-rhythm > div { min-width: 620px; }
.group-rhythm { overflow-x: auto; }
}
</style>
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+33 -7
View File
@@ -2,6 +2,7 @@
import BaseLayout from "@/layouts/BaseLayout.astro";
import ArchitectureExplorer from "@/components/ArchitectureExplorer.astro";
import K3ArtifactLab from "@/components/K3ArtifactLab.astro";
import K3AttnResGradientLab from "@/components/K3AttnResGradientLab.astro";
import K3AttnResTraceLab from "@/components/K3AttnResTraceLab.astro";
import K3ReportLab from "@/components/K3ReportLab.astro";
import { k3FigureAtlas, k3Ledgers, k3PaperChain, k3ReportMap } from "@/data/k3";
@@ -38,7 +39,8 @@ const toc = [
["28", "lab", "八联交互实验"],
["29", "artifacts", "开放权重工件审计"],
["30", "attnres-reduced", "AttnRes 缩小机制实验"],
["31", "audit", "21 张图表审计"],
["31", "attnres-gradient", "梯度定义与深度扩展"],
["32", "audit", "21 张图表审计"],
["↳", "papers", "100 节点阅读链"],
];
@@ -107,13 +109,13 @@ const paperGroups = [
<BaseLayout
title="Kimi K3 技术报告完整深读:架构、训练、RL、系统与评测"
description="用三十二张问题账、二十一张图表审计、八个机制实验、四个开放工件视图、五个 AttnRes 独立实验视图与一百个一手阅读节点,逐节读懂 Kimi K3。"
description="用三十二张问题账、二十一张图表审计、八个机制实验、四个开放工件视图、两轮十个 AttnRes 独立实验视图与一百个一手阅读节点,逐节读懂 Kimi K3。"
section="k3"
>
<header class="page-hero k3-hero">
<div class="page-hero-inner">
<div>
<p class="eyebrow"><span>ANCHOR REPORT / ROUND 04</span> KIMI K3 · REPORT → ARTIFACTS → INDEPENDENT PROBE</p>
<p class="eyebrow"><span>ANCHOR REPORT / ROUND 05</span> KIMI K3 · REPORT → ARTIFACTS → INDEPENDENT PROBE</p>
<h1>不把报告压成摘要<br />把每个因果环节<br />重新展开</h1>
<p class="lead">
K3 同时扩展序列、深度、宽度、视觉与 Agent 轨迹。真正值得读的不是 2.8T 这个最大数字,
@@ -123,11 +125,11 @@ const paperGroups = [
<dl class="page-facts">
<div><dt>QUESTIONS</dt><dd>32 张问题账</dd></div>
<div><dt>REPORT</dt><dd>16 Figures · 5 Tables</dd></div>
<div><dt>LABS</dt><dd>8 + 4 + 5 个交互视图</dd></div>
<div><dt>LABS</dt><dd>8 + 4 + 5 + 5 个交互视图</dd></div>
<div><dt>READING</dt><dd>100 个一手 / 官方节点</dd></div>
<div><dt>MODEL</dt><dd>2.78T total / 104.2B active</dd></div>
<div><dt>ARTIFACTS</dt><dd>96 shards · 497,220 tensors</dd></div>
<div><dt>STATUS</dt><dd>K3 四轮 · AttnRes 实验</dd></div>
<div><dt>STATUS</dt><dd>K3 五轮 · 梯度定义闭环</dd></div>
</dl>
</div>
</header>
@@ -892,12 +894,36 @@ const paperGroups = [
<a class="button primary" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/research/K3_ATTNRES_REDUCED_AUDIT.md">阅读完整研究审计</a>
<a class="button" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/experiments/k3/attnres">复跑公开实验代码</a>
<a class="button" href="https://arxiv.org/abs/2603.15031">Attention Residuals 原论文</a>
<a class="button" href="https://github.com/MoonshotAI/Attention-Residuals">官方实现</a>
<a class="button" href="https://github.com/MoonshotAI/Attention-Residuals">官方论文工件</a>
</div>
</section>
<section class="article-section" id="attnres-gradient">
<p class="eyebrow"><span>31</span> GRADIENT DEFINITION × DEPTH SCALE</p>
<h2>“论文说梯度更均匀”,和上一轮参数梯度反结果,测的是同一件事吗?</h2>
<p class="lede">
第五轮先审计 Attention Residuals 官方论文与仓库:Figure 5(c) 没有公开 gradient tensor、
norm、reduction、diagnostic batch、AMP / clipping 时点或统计代码。本站因此冻结一个可复现的
post-MLP output activation-gradient 定义,把深度扩到 16 / 32 blocks、预算扩到 8,000 steps,
再用三 seed 检查“首尾平衡”和“全层离散度”是否真的同方向。
</p>
<div class="artifact-callout">
<article><span>F / FROZEN</span><b>12 × 8,000 steps</b><p>786,432,000 formal target bytes;两深度、两结构、三 seed。</p></article>
<article><span>X / OBSERVED</span><b>first/last 6 / 6 改善</b><p>depth-16 平均 61.0%;depth-32 平均 72.0%。</p></article>
<article class="warning"><span>X / COUNTEREVIDENCE</span><b>CV 6 / 6 恶化</b><p>局部尖峰让 depth-16 / 32 平均相对恶化 10.3% / 60.0%。</p></article>
<article><span>R / REPLAY</span><b>model + optimizer exact</b><p>指定 32-layer Block 格从零重训 8,000 steps,冻结字段逐项一致。</p></article>
</div>
<K3AttnResGradientLab />
<div class="hero-actions">
<a class="button primary" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/research/K3_ATTNRES_GRADIENT_SCALE_AUDIT.md">阅读完整结果审计</a>
<a class="button" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/research/K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md">核对论文定义边界</a>
<a class="button" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/experiments/k3/attnres_gradient">复跑 12 格实验</a>
<a class="button" href="https://github.com/MoonshotAI/Attention-Residuals">官方一手工件</a>
</div>
</section>
<section class="article-section" id="audit">
<p class="eyebrow"><span>31</span> FIGURE & TABLE AUDIT</p>
<p class="eyebrow"><span>32</span> FIGURE & TABLE AUDIT</p>
<h2>Figure 1–16、Table 1–5:每张图究竟支持什么,不能支持什么</h2>
<div class="figure-atlas">
{k3FigureAtlas.map(([id, report, title, contract]) => (
+9 -4
View File
@@ -9,7 +9,7 @@ const researching = chapters.filter((chapter) => ["researching", "drafting"].inc
const workstreams = [
{ label: "研究框架与规范", value: 83, next: "给 Scaling 与推理专题补逐篇图表/实验精读层级" },
{ label: "网站设计系统", value: 89, next: "打印样式与更多通用可视化组件" },
{ label: "Kimi K3 深读", value: 96, next: "对齐 AttnRes 梯度定义并扩展深度/预算;等待 A_log 官方转换合同" },
{ label: "Kimi K3 深读", value: 98, next: "对齐 layer 21–25 梯度尖峰与 mixer weights;等待 A_log 官方转换合同" },
{ label: "语言模型前史", value: 78, next: "逐图精读 Kneser–Ney、LSTM 与 Bahdanau,并加入真实小语料复现" },
{ label: "Transformer 基础", value: 79, next: "逐图精读多头电路、Pre/Post-LN 与真实 kernel / KV 配置" },
{ label: "表示、位置与残差高速公路", value: 81, next: "加入真实 hidden-state / norm traces、长上下文位置外推复现与更多深层稳定性消融" },
@@ -50,7 +50,7 @@ const workstreams = [
<div><dt>OVERALL</dt><dd>专题平均 {average}%</dd></div>
<div><dt>READABLE</dt><dd>{published} 个首版可读专题</dd></div>
<div><dt>ACTIVE</dt><dd>{researching} 个研究/写作中</dd></div>
<div><dt>UPDATED</dt><dd>2026-07-30 07:30 CST</dd></div>
<div><dt>UPDATED</dt><dd>2026-07-30 10:05 CST</dd></div>
<div><dt>MODE</dt><dd>持续迭代,不锁死版本</dd></div>
</dl>
</div>
@@ -97,7 +97,7 @@ const workstreams = [
<article><span>✓</span><h3>K3 报告已结构化拆解</h3><p>47 页报告目录、151 条参考来源和架构/后训练/系统主线已经提取。</p></article>
<article><span>✓</span><h3>17 专题知识图</h3><p>从语言模型基础到评测安全,包含先修依赖和三条贯穿案例。</p></article>
<article><span>✓</span><h3>编辑式网站系统</h3><p>响应式导航、章节模板、侧栏、进度、论文链和证据提示组件。</p></article>
<article><span>✓</span><h3>九十四个原创交互视图</h3><p>K3 三轴图、八联报告实验、四联开放工件实验与五联 AttnRes 独立实验,DeepSeek 四联公式实验、十三联 Base 工件实验、Chat 行为、completion/full-depth、multi-seed、cross-source 与 task-bootstrap CRN 五轮实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
<article><span>✓</span><h3>九十九个原创交互视图</h3><p>K3 三轴图、八联报告实验、四联开放工件实验与两轮十联 AttnRes 独立实验,DeepSeek 四联公式实验、十三联 Base 工件实验、Chat 行为、completion/full-depth、multi-seed、cross-source 与 task-bootstrap CRN 五轮实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
<article><span>✓</span><h3>十七篇首版长文</h3><p>K3、语言模型前史、Transformer、表示/位置/残差、DeepSeek、Scaling、数据工程、长上下文、MoE、后训练、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全专题。</p></article>
<article><span>✓</span><h3>语言模型前史深度专题</h3><p>八张独立问题账、33 个正式节点、20 段长文与概率—向量—记忆—对齐四联实验。</p></article>
<article><span>✓</span><h3>Transformer 深度专题</h3><p>十张独立问题账、40 个正式节点、21 段正文与 QKV—Mask—多头位置—Block 成本四联实验。</p></article>
@@ -106,6 +106,7 @@ const workstreams = [
<article><span>✓</span><h3>Kimi K3 技术报告二轮深读</h3><p>三十二张问题账、Figure 1–16 / Table 1–5 审计、100 节点阅读链,以及 Delta—Decay—AttnRes—LatentMoE—SiTU—QB—MOPD—Cache 八联实验。</p></article>
<article><span>✓</span><h3>Kimi K3 三轮开放工件里程碑</h3><p>固定官方 revisions,审计 96 个 shards、497,220 个 tensor entries 与真实 KDA / MLA / MoE / MoonViT shapes;四联实验分开显示层型、tensor anatomy、参数范围和复现边界。</p></article>
<article><span>✓</span><h3>Kimi K3 四轮 AttnRes 独立实验</h3><p>冻结三结构 × 三 seed 的 9 个 2,000-step 格;Full / Block 相对 Baseline 的平均 paired delta 为 −0.01457 / −0.04247 BPC,但核心参数梯度 CV 没有复现论文叙述。指定正式格全新进程八字段 exact,五视图同时展示结果、反证、成本与 claim boundary。</p></article>
<article><span>✓</span><h3>Kimi K3 五轮梯度定义与深度扩展</h3><p>先确认 Figure 5 没有公开唯一 gradient telemetry 合同,再冻结 16/32 blocks × Baseline/Block × 3 seeds 的 12 个 8,000-step 格。Block 的首尾失衡 6/6 改善但全层 CV 6/6 恶化,两个深度都判为 mixed;指定 32 层格完整重训的模型、优化器与全部冻结字段 exact。</p></article>
<article><span>✓</span><h3>FlashKDA RTX 5090 执行闸门</h3><p>隔离 CUDA 13.0 / glibc 2.39 编译 sm_120a wheel;6/6 官方参考逐元素相等,并完成 fixed / varlen、三种 state mode 的 1,800 个 CUDA Event samples。</p></article>
<article><span>✓</span><h3>Scaling Laws 深度专题</h3><p>九张账、29 个一手节点、DeepSeek/Kimi 双谱系与曲面—部署—复用—涌现四联实验。</p></article>
<article><span>✓</span><h3>数据工程深度专题</h3><p>十二张账、31 个一手节点、DeepSeek/Kimi 双谱系与流水线—去重—混合—改写四联实验。</p></article>
@@ -134,7 +135,7 @@ const workstreams = [
</div>
<div class="queue-table">
<div class="head"><b>优先级</b><b>专题</b><b>本轮交付</b><b>完成闸门</b></div>
<div><span>P0</span><strong>K3 四轮后续</strong><p>对齐论文梯度定义 → 增加 depth / budget → 等待 A_log 官方合同后进入真实 checkpoint forward</p><em>尺度复查 + 工件边界</em></div>
<div><span>P0</span><strong>K3 五轮后续</strong><p>对齐 layer 21–25 尖峰、pre-attention / pre-MLP 与 mixer source weights → 等待 A_log 官方合同后进入真实 checkpoint forward</p><em>局部机制 + 工件边界</em></div>
<div><span>P0</span><strong>DeepSeek 八轮后续</strong><p>干预式 mediation → SM90 FlashMLA / FP8 / pipeline traces → R1-like RL 小模型复现</p><em>运行证据 + 独立复现</em></div>
<div><span>P0</span><strong>Transformer 二轮</strong><p>多头电路逐图 → Pre/Post-LN 真实 traces → Flash/KV 配置与 kernel 对照</p><em>逐图笔记 + 实测边界</em></div>
<div><span>P0</span><strong>表示、位置与残差二轮</strong><p>真实 hidden-state / norm traces → 长上下文位置外推 → mHC / AttnRes 深层稳定性消融</p><em>可复现实验 + 逐图笔记</em></div>
@@ -220,6 +221,10 @@ const workstreams = [
<div><time>2026-07-30</time><b>AttnRes 缩小实验先冻结、后运行</b><p>三结构共享公共主干、初始化、窗口与优化器;只按三个 paired seed 和预注册 −0.010 BPC 阈值给出本协议内方向判断。</p></div>
<div><time>2026-07-30</time><b>支持结果与梯度反结果同时进入主视区</b><p>Full / Block 的最终 BPC 同向改善;核心参数 gradient RMS CV 却高于 Baseline,不换指标掩盖。</p></div>
<div><time>2026-07-30</time><b>正式重放不把 wall time 纳入 exact</b><p>Block / seed-1 的模型、优化器、曲线、历史、诊断和环境八字段 exact;计时受调度影响,单独报告。</p></div>
<div><time>2026-07-30</time><b>Figure 5 的“梯度”不再靠猜测补合同</b><p>官方未公开 gradient tensor、norm、reduction 与统计代码;本站 activation-gradient 定义只叫 operationalization,不叫论文复画。</p></div>
<div><time>2026-07-30</time><b>首尾平衡与全层 CV 永久分账</b><p>Block 在 6/6 配对中改善 first/last,却因中后段局部尖峰让 CV 在 6/6 配对中恶化;联合判定保持 mixed。</p></div>
<div><time>2026-07-30</time><b>绝对梯度尺度必须与归一化谱同屏</b><p>Block mean gradient 约为 Baseline 的 54%–57%;更接近 1 的首尾比不能偷换成各层信号更强。</p></div>
<div><time>2026-07-30</time><b>32 层完整重放扩到状态哈希</b><p>8,000-step fresh replay 的全部冻结字段以及 model / optimizer state hashes exact;额外 replay bytes 单列,不混入 formal 预算。</p></div>
<div><time>2026-07-29</time><b>32-token 对照改为同源 16→24</b><p>TNEWS 只有 105/10,000 条达到 32 tokens,强行统一会落入约 1% 极端长尾;24-token eligibility 仍保留 1,609 条中文候选。</p></div>
<div><time>2026-07-29</time><b>长度敏感性必须成对重采样</b><p>16-token 输入严格是 24-token 输入前缀,2,000 次 bootstrap 共用 prompt indices;结果只描述固定 cohort 的长度敏感性。</p></div>
<div><time>2026-07-29</time><b>三类 cohort 永久分身份</b><p>自然长度回答本批样本如何路由;matched-16 / 24 回答同一 prompt 多看 8 tokens 后如何变化,不把二者混成内容因果。</p></div>