Compare commits
5 Commits
40f6a1d631
...
f7670efcdd
| Author | SHA1 | Date | |
|---|---|---|---|
| f7670efcdd | |||
| 26fc824409 | |||
| 3f7cc1f544 | |||
| 1675ca54f3 | |||
| 091a05f0a0 |
+14
-2
@@ -41,7 +41,7 @@
|
|||||||
- [x] 完成 486 篇关键论文索引,覆盖 16 个标签专题与 Kimi/DeepSeek 聚光主线。
|
- [x] 完成 486 篇关键论文索引,覆盖 16 个标签专题与 Kimi/DeepSeek 聚光主线。
|
||||||
- [x] 完成可检索、可按专题筛选的论文库页面。
|
- [x] 完成可检索、可按专题筛选的论文库页面。
|
||||||
- [x] 完成 K3、语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全十七篇首版长文。
|
- [x] 完成 K3、语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全十七篇首版长文。
|
||||||
- [x] 完成 K3 三轴架构、八联报告实验、四联开放工件实验与五联 AttnRes 独立实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 二十二联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等九十四个原创交互视图。
|
- [x] 完成 K3 三轴架构、八联报告实验、四联开放工件实验、Round 04 / 05 各五联 AttnRes 独立实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 二十二联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等九十九个原创交互视图。
|
||||||
- [x] 完成长上下文首版:五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。
|
- [x] 完成长上下文首版:五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。
|
||||||
- [x] 核验 FlashAttention、DeepSeek-V2/V3.2/V4、Kimi Linear/K3 等六份论文原文,并建立长上下文研究账本。
|
- [x] 核验 FlashAttention、DeepSeek-V2/V3.2/V4、Kimi Linear/K3 等六份论文原文,并建立长上下文研究账本。
|
||||||
- [x] 核验 Switch、ST-MoE、DeepSeekMoE、Loss-Free、V3、LatentMoE 与 K3 原文,并建立 MoE 研究账本。
|
- [x] 核验 Switch、ST-MoE、DeepSeekMoE、Loss-Free、V3、LatentMoE 与 K3 原文,并建立 MoE 研究账本。
|
||||||
@@ -279,10 +279,18 @@
|
|||||||
- [x] K3 Round 04 五视图实验室完成:三 seed BPC 曲线、Residual RMS / Block 锯齿、Full / Block depth-weight heatmap、梯度反证、成本/哈希/claim boundary 分开展示;完整 9-run JSON、compact 数据、复现清单、协议、审计、训练与聚合代码进入公开仓库。
|
- [x] K3 Round 04 五视图实验室完成:三 seed BPC 曲线、Residual RMS / Block 锯齿、Full / Block depth-weight heatmap、梯度反证、成本/哈希/claim boundary 分开展示;完整 9-run JSON、compact 数据、复现清单、协议、审计、训练与聚合代码进入公开仓库。
|
||||||
- [x] Round 04 本地闸门通过:91 个受检文件零诊断/提示,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;冻结数据、AttnRes 专项与 K3 全量真实 Chrome 回归通过,桌面/390px 移动端零文档级溢出、零运行时异常。
|
- [x] Round 04 本地闸门通过:91 个受检文件零诊断/提示,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;冻结数据、AttnRes 专项与 K3 全量真实 Chrome 回归通过,桌面/390px 移动端零文档级溢出、零运行时异常。
|
||||||
- [x] K3 Round 04 以功能源提交 `4ce780d`、不可变镜像 `20260729T233142Z-4ce780d` 发布;OCI index digest `sha256:6e89f802…25f582`,复用 NAS `12010→8080`、NPM host 31 / cert 41 与门户 `LLM ATLAS / projects / 180`。容器 healthy、0 次重启,21/21 公网页面、HTTPS/2、gzip / immutable assets、AttnRes 专项与 K3 全量生产 Chrome 回归通过;保留 `20260729T221654Z-975ed3d` 回滚。
|
- [x] K3 Round 04 以功能源提交 `4ce780d`、不可变镜像 `20260729T233142Z-4ce780d` 发布;OCI index digest `sha256:6e89f802…25f582`,复用 NAS `12010→8080`、NPM host 31 / cert 41 与门户 `LLM ATLAS / projects / 180`。容器 healthy、0 次重启,21/21 公网页面、HTTPS/2、gzip / immutable assets、AttnRes 专项与 K3 全量生产 Chrome 回归通过;保留 `20260729T221654Z-975ed3d` 回滚。
|
||||||
|
- [x] K3 Round 05 一手定义审计确认 Figure 5(c) 未公开 gradient tensor、norm、reduction、diagnostic batch、AMP / clipping 时点或统计代码;Round 04 参数梯度与 Round 05 post-MLP output activation gradient 永久分对象记账,不把本站 operationalization 冒充作者实现。
|
||||||
|
- [x] 在任何 formal 输出前冻结 16 / 32 blocks、Baseline / Block、三个 seed、8,000 steps、六个诊断点、CV + 首尾四分位 imbalance 联合判据、FP32 residual accumulator、20-step 双 smoke 与 depth-32 Block 完整 replay;768,000-window schedule SHA-256 为 `5041e09b…f4e`。
|
||||||
|
- [x] 12 个 formal 格全部完成,共 786,432,000 target bytes;公共主干与三个 gate input tensor hashes 在 paired 架构间 exact。Block 的验证 BPC 在 6 / 6 配对中更低,depth-16 / 32 mean delta 为 `−.008938 / −.009872`,但不追加事后 BPC support 阈值。
|
||||||
|
- [x] activation-gradient 结果分裂:首/末四分位 imbalance 在 6 / 6 配对改善,depth-16 / 32 均值为 `+61.0% / +72.0%`;全层 CV 却在 6 / 6 配对恶化,均值相对 reduction 为 `−10.3% / −60.0%`。两个 depth 都按预注册规则判为 `mixed / inconclusive`,总判定 `depth-dependent or inconclusive`。
|
||||||
|
- [x] 绝对 gradient mean 仅为 Baseline 的 `57.4% / 54.4%`;参数 gradient CV 从 `0.416→0.683 / 0.397→0.772`,继续保留反结果。Output RMS 最后/第一层比则由 Baseline `4.59× / 6.08×` 降至 Block `1.17× / 1.89×`。
|
||||||
|
- [x] 指定 depth-32 / Block / seed-2026073001 从初始化完整重训 8,000 steps;排除 run-kind / timing 后冻结字段 compare SHA-256 同为 `46300a45…4817`,model / optimizer state hashes exact。正式/compact/reproduction 物理 SHA-256 为 `ad461cbe…a8d / 5377a5e7…68e3 / aedcde6a…dea6`。
|
||||||
|
- [x] K3 Round 05 五视图实验室完成:论文定义已知/未定义、绝对/归一化深度谱、六 checkpoint 时间轨迹、Output RMS 组节律、activation/parameter/BPC/成本/重放联合账全部可切换;21 个 raw JSON、完整 aggregate、compact、runner、analyzer、协议与审计进入公开树。
|
||||||
|
- [x] Round 05 本地闸门通过:94 个 Astro 文件零诊断/提示,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;冻结数据、新专项、Round 04 与 K3 全量真实 Chrome 回归通过,桌面/390px 移动端零文档级溢出、零运行时异常。
|
||||||
|
|
||||||
## 正在进行
|
## 正在进行
|
||||||
|
|
||||||
- [ ] K3 四轮下一闸门:对齐 AttnRes 论文的 activation / residual-output gradient 定义,增加模型深度与训练预算,检验本轮梯度反结果是否随尺度翻转;真实 K3 forward 继续等待 `A_log [128]→[96]` 官方转换或权重修订。
|
- [ ] K3 五轮下一闸门:对齐 layer 21–25 的 activation-gradient 尖峰、pre-attention / pre-MLP 位置与 mixer source weights,并做公开 reduction sensitivity;真实 K3 forward 继续等待 `A_log [128]→[96]` 官方转换或权重修订。
|
||||||
- [ ] DeepSeek 八轮下一闸门:推进干预式 mediation、SM90 FlashMLA、FP8 / pipeline traces 与 R1-like RL 小模型复现。
|
- [ ] DeepSeek 八轮下一闸门:推进干预式 mediation、SM90 FlashMLA、FP8 / pipeline traces 与 R1-like RL 小模型复现。
|
||||||
- [ ] 表示、位置与残差二轮:真实 hidden-state / norm traces、长上下文位置外推复现与 mHC / AttnRes 深层稳定性消融。
|
- [ ] 表示、位置与残差二轮:真实 hidden-state / norm traces、长上下文位置外推复现与 mHC / AttnRes 深层稳定性消融。
|
||||||
- [ ] 评测安全二轮:真实 cross-harness / pass@k 复跑、Judge 元评测、动态污染与过拒案例。
|
- [ ] 评测安全二轮:真实 cross-harness / pass@k 复跑、Judge 元评测、动态污染与过拒案例。
|
||||||
@@ -476,6 +484,10 @@
|
|||||||
| 2026-07-30 | 主结果与机制反结果同时发布 | Full / Block BPC 方向支持;核心参数 gradient RMS CV 却高于 Baseline,明确写成未复现论文梯度叙述 |
|
| 2026-07-30 | 主结果与机制反结果同时发布 | Full / Block BPC 方向支持;核心参数 gradient RMS CV 却高于 Baseline,明确写成未复现论文梯度叙述 |
|
||||||
| 2026-07-30 | 独立重放按数值合同而非计时合同验收 | Block / seed-1 的八组冻结字段 2,000 steps exact;wall time 受调度影响,不要求或声称 bit-exact |
|
| 2026-07-30 | 独立重放按数值合同而非计时合同验收 | Block / seed-1 的八组冻结字段 2,000 steps exact;wall time 受调度影响,不要求或声称 bit-exact |
|
||||||
| 2026-07-30 | K3 Round 04 缩小 AttnRes 里程碑发布 | 功能源 `4ce780d`、镜像 `20260729T233142Z-4ce780d`、OCI `sha256:6e89f802…25f582`;21/21 公网页面与生产专项/全量 Chrome 通过,保留 Round 08 回滚点 |
|
| 2026-07-30 | K3 Round 04 缩小 AttnRes 里程碑发布 | 功能源 `4ce780d`、镜像 `20260729T233142Z-4ce780d`、OCI `sha256:6e89f802…25f582`;21/21 公网页面与生产专项/全量 Chrome 通过,保留 Round 08 回滚点 |
|
||||||
|
| 2026-07-30 | AttnRes 的“梯度”先按公开证据拆对象 | 论文 Figure 5 没有公开唯一 telemetry 合同;参数梯度与 post-MLP activation gradient 不再互相代称 |
|
||||||
|
| 2026-07-30 | “更均匀”拆成首尾失衡与全层 CV | Block 6/6 改善 first/last,却 6/6 恶化 CV;局部尖峰与系统性早层隆起必须分开解释 |
|
||||||
|
| 2026-07-30 | 绝对尺度与归一化形状永久同报 | Block activation-gradient mean 约为 Baseline 54%–57%;不能把更接近 1 的首尾比自动解释为各层信号更强 |
|
||||||
|
| 2026-07-30 | Round 05 完整重放过闸 | depth-32 Block seed-1 从零重训 8,000 steps;全部冻结字段与 model/optimizer state hashes exact,timing 仍单独报告 |
|
||||||
|
|
||||||
## 未决问题
|
## 未决问题
|
||||||
|
|
||||||
|
|||||||
@@ -19,7 +19,7 @@
|
|||||||
|
|
||||||
当前里程碑包含 17 专题学习地图、486 篇关键论文索引、Kimi K3 完整导读,
|
当前里程碑包含 17 专题学习地图、486 篇关键论文索引、Kimi K3 完整导读,
|
||||||
语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 技术谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、工具使用与长程 Agent、原生多模态、训练系统、推理服务、数值优化,以及评测与安全深度专题,
|
语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 技术谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、工具使用与长程 Agent、原生多模态、训练系统、推理服务、数值优化,以及评测与安全深度专题,
|
||||||
以及 94 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、
|
以及 99 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、
|
||||||
8 个交互实验和 100 个一手/官方节点,完整覆盖架构、预训练、后训练、系统、评测、案例与附录。
|
8 个交互实验和 100 个一手/官方节点,完整覆盖架构、预训练、后训练、系统、评测、案例与附录。
|
||||||
第三轮已完成开放工件与首个真实 kernel 里程碑:固定官方模型与 FlashKDA revisions,审计 96 个 checkpoint shards、
|
第三轮已完成开放工件与首个真实 kernel 里程碑:固定官方模型与 FlashKDA revisions,审计 96 个 checkpoint shards、
|
||||||
497,220 个 tensor entries、真实 KDA / MLA / MoE / MoonViT shapes 与小范围参数统计,并用 4 个新视图
|
497,220 个 tensor entries、真实 KDA / MLA / MoE / MoonViT shapes 与小范围参数统计,并用 4 个新视图
|
||||||
@@ -39,6 +39,19 @@
|
|||||||
[K3_ATTNRES_REDUCED_PROTOCOL.md](./research/K3_ATTNRES_REDUCED_PROTOCOL.md)、
|
[K3_ATTNRES_REDUCED_PROTOCOL.md](./research/K3_ATTNRES_REDUCED_PROTOCOL.md)、
|
||||||
[K3_ATTNRES_REDUCED_AUDIT.md](./research/K3_ATTNRES_REDUCED_AUDIT.md) 与
|
[K3_ATTNRES_REDUCED_AUDIT.md](./research/K3_ATTNRES_REDUCED_AUDIT.md) 与
|
||||||
[AttnRes experiment](./experiments/k3/attnres/)。
|
[AttnRes experiment](./experiments/k3/attnres/)。
|
||||||
|
第五轮先审计 Attention Residuals Figure 5 的公开定义边界,再冻结
|
||||||
|
`llm-atlas-k3-attnres-gradient-scale-v1`:以 post-MLP block output activation gradient
|
||||||
|
为公开 operationalization,把深度扩为 16 / 32 blocks、预算扩为每格 8,000 steps,
|
||||||
|
完成 Baseline / Block × 三 seed 共 12 格、786,432,000 formal target bytes。结果把
|
||||||
|
“更均匀”拆成两个相反方向:Block 在 6 / 6 配对中把首/末四分位失衡改善 56%–81%,
|
||||||
|
却因中后段局部尖峰让全层 CV 在 6 / 6 配对中恶化;两个深度都按预注册联合规则判为
|
||||||
|
mixed / inconclusive。验证 BPC 仍在 6 / 6 配对中更低,但实际 step time 约 2.6×、
|
||||||
|
peak allocated memory 约 2.2×,不冒充同算力优势。指定 depth-32 / Block / seed-1
|
||||||
|
从零重训完整 8,000 steps,全部冻结字段以及 model / optimizer state hashes exact。
|
||||||
|
详见
|
||||||
|
[K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md](./research/K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md)、
|
||||||
|
[K3_ATTNRES_GRADIENT_SCALE_AUDIT.md](./research/K3_ATTNRES_GRADIENT_SCALE_AUDIT.md) 与
|
||||||
|
[gradient experiment](./experiments/k3/attnres_gradient/)。
|
||||||
DeepSeek 八轮专题以 24 张问题账、10 次技术转向、
|
DeepSeek 八轮专题以 24 张问题账、10 次技术转向、
|
||||||
22 个交互实验和 60 个一手/官方节点,串起 Dense、MoE、MLA、V3 协同、R1 与 V4;
|
22 个交互实验和 60 个一手/官方节点,串起 Dense、MoE、MLA、V3 协同、R1 与 V4;
|
||||||
并固定官方 V2-Lite revision,在 RTX 5090 上连续执行 7/27 层,记录 3,240 次真实专家选择、
|
并固定官方 V2-Lite revision,在 RTX 5090 上连续执行 7/27 层,记录 3,240 次真实专家选择、
|
||||||
|
|||||||
@@ -0,0 +1,57 @@
|
|||||||
|
# Attention Residuals activation-gradient/depth study
|
||||||
|
|
||||||
|
This directory implements preregistered protocol
|
||||||
|
`llm-atlas-k3-attnres-gradient-scale-v1`:
|
||||||
|
|
||||||
|
- `research/K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md`
|
||||||
|
- `research/K3_ATTNRES_GRADIENT_SCALE_PROTOCOL.md`
|
||||||
|
|
||||||
|
It is an independent reduced mechanism experiment. It is not a Kimi K3
|
||||||
|
checkpoint forward pass and does not claim to recover the paper's unpublished
|
||||||
|
Figure 5 telemetry definition.
|
||||||
|
|
||||||
|
## Frozen environment
|
||||||
|
|
||||||
|
```text
|
||||||
|
Python /home/wuyang/.pyenv/versions/3.10.14/envs/navi-router-cu128/bin/python
|
||||||
|
PyTorch 2.11.0+cu128
|
||||||
|
GPU NVIDIA GeForce RTX 5090
|
||||||
|
CUBLAS_WORKSPACE_CONFIG=:4096:8
|
||||||
|
```
|
||||||
|
|
||||||
|
## Build the manifest
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python experiments/k3/attnres_gradient/build_dataset.py \
|
||||||
|
--cache-dir /home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1 \
|
||||||
|
--manifest experiments/k3/attnres_gradient/manifest.json
|
||||||
|
```
|
||||||
|
|
||||||
|
## Run a smoke cell
|
||||||
|
|
||||||
|
```bash
|
||||||
|
CUBLAS_WORKSPACE_CONFIG=:4096:8 \
|
||||||
|
python experiments/k3/attnres_gradient/train.py \
|
||||||
|
--run-kind smoke \
|
||||||
|
--architecture block \
|
||||||
|
--depth 32 \
|
||||||
|
--seed 2026073001 \
|
||||||
|
--cache-dir /home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1 \
|
||||||
|
--manifest experiments/k3/attnres_gradient/manifest.json \
|
||||||
|
--output /home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1/smoke-a/depth-32-block.json
|
||||||
|
```
|
||||||
|
|
||||||
|
Smoke is fixed to 20 steps. Formal and replay runs are fixed to 8,000 steps;
|
||||||
|
the runner rejects alternative budgets. The same command uses
|
||||||
|
`--run-kind formal` or `--run-kind replay` and omits an explicit `--steps`.
|
||||||
|
|
||||||
|
Formal output keys use:
|
||||||
|
|
||||||
|
```text
|
||||||
|
formal/depth-{16|32}-{baseline|block}-seed-{seed}.json
|
||||||
|
replay/depth-32-block-seed-2026073001.json
|
||||||
|
```
|
||||||
|
|
||||||
|
Raw parquet/binary files and working runs remain in the local cache. The
|
||||||
|
manifest, runner, complete result JSON, compact website payload, reproduction
|
||||||
|
hashes, protocol, and audit enter the public repository.
|
||||||
@@ -0,0 +1,685 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Validate, aggregate, and publish K3 AttnRes Round 05 experiment data."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import hashlib
|
||||||
|
import json
|
||||||
|
import math
|
||||||
|
import os
|
||||||
|
import shutil
|
||||||
|
import statistics
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, Iterable
|
||||||
|
|
||||||
|
|
||||||
|
PROTOCOL_ID = "llm-atlas-k3-attnres-gradient-scale-v1"
|
||||||
|
ARCHITECTURES = ("baseline", "block")
|
||||||
|
DEPTHS = (16, 32)
|
||||||
|
SEEDS = (2026073001, 2026073002, 2026073003)
|
||||||
|
STEPS = (0, 100, 500, 2000, 4000, 8000)
|
||||||
|
FORMAL_STEPS = 8000
|
||||||
|
FORMAL_BATCH = 32
|
||||||
|
TARGET_BYTES_PER_RUN = 65_536_000
|
||||||
|
EXPECTED_TOTAL_TARGET_BYTES = 786_432_000
|
||||||
|
SMOKE_COMPARE_FIELDS = (
|
||||||
|
"protocol_id",
|
||||||
|
"run_kind",
|
||||||
|
"architecture",
|
||||||
|
"depth",
|
||||||
|
"seed",
|
||||||
|
"steps",
|
||||||
|
"batch_size",
|
||||||
|
"target_bytes_seen",
|
||||||
|
"manifest",
|
||||||
|
"model",
|
||||||
|
"optimizer",
|
||||||
|
"hashes",
|
||||||
|
"evaluations",
|
||||||
|
"diagnostics",
|
||||||
|
"training_history",
|
||||||
|
"gradient_gate",
|
||||||
|
"environment",
|
||||||
|
)
|
||||||
|
REPLAY_COMPARE_FIELDS = tuple(
|
||||||
|
field for field in SMOKE_COMPARE_FIELDS if field != "run_kind"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def parse_args() -> argparse.Namespace:
|
||||||
|
parser = argparse.ArgumentParser()
|
||||||
|
parser.add_argument("--formal-dir", type=Path, required=True)
|
||||||
|
parser.add_argument("--smoke-a-dir", type=Path, required=True)
|
||||||
|
parser.add_argument("--smoke-b-dir", type=Path, required=True)
|
||||||
|
parser.add_argument("--replay", type=Path, required=True)
|
||||||
|
parser.add_argument("--manifest", type=Path, required=True)
|
||||||
|
parser.add_argument("--raw-output-dir", type=Path, required=True)
|
||||||
|
parser.add_argument("--output", type=Path, required=True)
|
||||||
|
parser.add_argument("--compact-output", type=Path, required=True)
|
||||||
|
parser.add_argument("--reproduction-output", type=Path, required=True)
|
||||||
|
return parser.parse_args()
|
||||||
|
|
||||||
|
|
||||||
|
def read_json(path: Path) -> dict[str, Any]:
|
||||||
|
return json.loads(path.read_text())
|
||||||
|
|
||||||
|
|
||||||
|
def file_sha256(path: Path) -> str:
|
||||||
|
digest = hashlib.sha256()
|
||||||
|
with path.open("rb") as handle:
|
||||||
|
for block in iter(lambda: handle.read(1024 * 1024), b""):
|
||||||
|
digest.update(block)
|
||||||
|
return digest.hexdigest()
|
||||||
|
|
||||||
|
|
||||||
|
def canonical_sha256(value: Any) -> str:
|
||||||
|
payload = json.dumps(
|
||||||
|
value, ensure_ascii=False, sort_keys=True, separators=(",", ":")
|
||||||
|
).encode()
|
||||||
|
return hashlib.sha256(payload).hexdigest()
|
||||||
|
|
||||||
|
|
||||||
|
def atomic_json(path: Path, value: dict[str, Any]) -> None:
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
temporary = path.with_suffix(path.suffix + ".tmp")
|
||||||
|
temporary.write_text(
|
||||||
|
json.dumps(value, ensure_ascii=False, indent=2, sort_keys=True) + "\n"
|
||||||
|
)
|
||||||
|
os.replace(temporary, path)
|
||||||
|
|
||||||
|
|
||||||
|
def mean(values: Iterable[float]) -> float:
|
||||||
|
return statistics.fmean(values)
|
||||||
|
|
||||||
|
|
||||||
|
def require_finite(value: Any, path: str = "root") -> None:
|
||||||
|
if isinstance(value, float):
|
||||||
|
if not math.isfinite(value):
|
||||||
|
raise ValueError(f"non-finite float at {path}")
|
||||||
|
elif isinstance(value, dict):
|
||||||
|
for key, child in value.items():
|
||||||
|
require_finite(child, f"{path}.{key}")
|
||||||
|
elif isinstance(value, list):
|
||||||
|
for index, child in enumerate(value):
|
||||||
|
require_finite(child, f"{path}[{index}]")
|
||||||
|
|
||||||
|
|
||||||
|
def selected(run: dict[str, Any], fields: tuple[str, ...]) -> dict[str, Any]:
|
||||||
|
return {field: run[field] for field in fields}
|
||||||
|
|
||||||
|
|
||||||
|
def final_diagnostic(run: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
diagnostic = run["diagnostics"][-1]
|
||||||
|
if diagnostic["step"] != FORMAL_STEPS:
|
||||||
|
raise ValueError("final diagnostic is not step 8000")
|
||||||
|
return diagnostic
|
||||||
|
|
||||||
|
|
||||||
|
def validate_run(
|
||||||
|
run: dict[str, Any],
|
||||||
|
*,
|
||||||
|
path: Path,
|
||||||
|
depth: int,
|
||||||
|
architecture: str,
|
||||||
|
seed: int,
|
||||||
|
manifest: dict[str, Any],
|
||||||
|
manifest_hash: str,
|
||||||
|
) -> None:
|
||||||
|
if run["protocol_id"] != PROTOCOL_ID or run["run_kind"] != "formal":
|
||||||
|
raise ValueError(f"formal protocol/kind mismatch: {path}")
|
||||||
|
if (
|
||||||
|
run["depth"] != depth
|
||||||
|
or run["architecture"] != architecture
|
||||||
|
or run["seed"] != seed
|
||||||
|
):
|
||||||
|
raise ValueError(f"formal identity mismatch: {path}")
|
||||||
|
if (
|
||||||
|
run["steps"] != FORMAL_STEPS
|
||||||
|
or run["batch_size"] != FORMAL_BATCH
|
||||||
|
or run["target_bytes_seen"] != TARGET_BYTES_PER_RUN
|
||||||
|
):
|
||||||
|
raise ValueError(f"formal budget mismatch: {path}")
|
||||||
|
if run["manifest"]["file_sha256"] != manifest_hash:
|
||||||
|
raise ValueError(f"manifest file hash mismatch: {path}")
|
||||||
|
for key in (
|
||||||
|
"formal_schedule_sha256",
|
||||||
|
"validation_tensor_sha256",
|
||||||
|
"diagnostic_tensor_sha256",
|
||||||
|
):
|
||||||
|
if run["manifest"][key] != manifest["windows"][key]:
|
||||||
|
raise ValueError(f"manifest {key} mismatch: {path}")
|
||||||
|
if [row["step"] for row in run["evaluations"]] != list(STEPS):
|
||||||
|
raise ValueError(f"evaluation steps mismatch: {path}")
|
||||||
|
if [row["step"] for row in run["diagnostics"]] != list(STEPS):
|
||||||
|
raise ValueError(f"diagnostic steps mismatch: {path}")
|
||||||
|
if run["model"]["layers"] != depth:
|
||||||
|
raise ValueError(f"model depth mismatch: {path}")
|
||||||
|
|
||||||
|
for diagnostic in run["diagnostics"]:
|
||||||
|
capture = diagnostic["capture"]
|
||||||
|
if (
|
||||||
|
capture["count"] != depth
|
||||||
|
or capture["shape"] != [16, 256, 192]
|
||||||
|
or set(capture["dtypes"]) != {"torch.float32"}
|
||||||
|
or not capture["all_gradients_finite"]
|
||||||
|
or not capture["all_gradients_present"]
|
||||||
|
or not capture["storage_unique"]
|
||||||
|
):
|
||||||
|
raise ValueError(f"activation capture gate mismatch: {path}")
|
||||||
|
for field in (
|
||||||
|
"activation_grad_rms_by_block",
|
||||||
|
"activation_output_rms_by_block",
|
||||||
|
"core_parameter_grad_rms_by_block",
|
||||||
|
):
|
||||||
|
if len(diagnostic[field]) != depth:
|
||||||
|
raise ValueError(f"{field} length mismatch: {path}")
|
||||||
|
for field in (
|
||||||
|
"layer_input_rms_by_sublayer",
|
||||||
|
"branch_output_rms_by_sublayer",
|
||||||
|
"stream_state_rms_by_sublayer",
|
||||||
|
):
|
||||||
|
if len(diagnostic[field]) != depth * 2:
|
||||||
|
raise ValueError(f"{field} length mismatch: {path}")
|
||||||
|
if architecture == "baseline":
|
||||||
|
if diagnostic["depth_weights"] or diagnostic["output_weights"] is not None:
|
||||||
|
raise ValueError(f"unexpected Baseline mixer trace: {path}")
|
||||||
|
else:
|
||||||
|
if (
|
||||||
|
len(diagnostic["depth_weights"]) != depth * 2
|
||||||
|
or diagnostic["output_weights"]["sources"] != 9
|
||||||
|
):
|
||||||
|
raise ValueError(f"Block mixer trace mismatch: {path}")
|
||||||
|
require_finite(run, path.name)
|
||||||
|
|
||||||
|
|
||||||
|
def relative_reduction(baseline: float, block: float) -> float:
|
||||||
|
return (baseline - block) / baseline
|
||||||
|
|
||||||
|
|
||||||
|
def depth_verdict(rows: list[dict[str, Any]]) -> dict[str, Any]:
|
||||||
|
cv_reductions = [row["relative_cv_reduction"] for row in rows]
|
||||||
|
imbalance_reductions = [
|
||||||
|
row["relative_imbalance_reduction"] for row in rows
|
||||||
|
]
|
||||||
|
mean_cv = mean(cv_reductions)
|
||||||
|
mean_imbalance = mean(imbalance_reductions)
|
||||||
|
support = (
|
||||||
|
all(value > 0 for value in cv_reductions)
|
||||||
|
and mean_cv >= 0.20
|
||||||
|
and all(value > 0 for value in imbalance_reductions)
|
||||||
|
and mean_imbalance >= 0.20
|
||||||
|
)
|
||||||
|
concern = (
|
||||||
|
all(value < 0 for value in cv_reductions)
|
||||||
|
and -mean_cv >= 0.20
|
||||||
|
and all(value < 0 for value in imbalance_reductions)
|
||||||
|
and -mean_imbalance >= 0.20
|
||||||
|
)
|
||||||
|
if support:
|
||||||
|
label = "joint directional support at this depth"
|
||||||
|
elif concern:
|
||||||
|
label = "joint directional concern at this depth"
|
||||||
|
else:
|
||||||
|
label = "mixed / inconclusive at this depth"
|
||||||
|
return {
|
||||||
|
"label": label,
|
||||||
|
"threshold_relative": 0.20,
|
||||||
|
"cv_reductions": cv_reductions,
|
||||||
|
"mean_cv_reduction": mean_cv,
|
||||||
|
"imbalance_reductions": imbalance_reductions,
|
||||||
|
"mean_imbalance_reduction": mean_imbalance,
|
||||||
|
"all_cv_improve": all(value > 0 for value in cv_reductions),
|
||||||
|
"all_imbalance_improve": all(
|
||||||
|
value > 0 for value in imbalance_reductions
|
||||||
|
),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def summarize_depth(
|
||||||
|
depth: int, runs: dict[tuple[int, str, int], dict[str, Any]]
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
rows = []
|
||||||
|
for seed in SEEDS:
|
||||||
|
baseline = runs[(depth, "baseline", seed)]
|
||||||
|
block = runs[(depth, "block", seed)]
|
||||||
|
baseline_diagnostic = final_diagnostic(baseline)
|
||||||
|
block_diagnostic = final_diagnostic(block)
|
||||||
|
baseline_activation = baseline_diagnostic[
|
||||||
|
"activation_grad_statistics"
|
||||||
|
]
|
||||||
|
block_activation = block_diagnostic["activation_grad_statistics"]
|
||||||
|
baseline_bpc = baseline["evaluations"][-1]["bits_per_byte"]
|
||||||
|
block_bpc = block["evaluations"][-1]["bits_per_byte"]
|
||||||
|
rows.append(
|
||||||
|
{
|
||||||
|
"seed": seed,
|
||||||
|
"baseline_bpc": baseline_bpc,
|
||||||
|
"block_bpc": block_bpc,
|
||||||
|
"block_minus_baseline_bpc": block_bpc - baseline_bpc,
|
||||||
|
"baseline_activation_grad_mean": baseline_activation["mean"],
|
||||||
|
"block_activation_grad_mean": block_activation["mean"],
|
||||||
|
"block_to_baseline_activation_grad_mean": (
|
||||||
|
block_activation["mean"] / baseline_activation["mean"]
|
||||||
|
),
|
||||||
|
"baseline_cv": baseline_activation["population_cv"],
|
||||||
|
"block_cv": block_activation["population_cv"],
|
||||||
|
"relative_cv_reduction": relative_reduction(
|
||||||
|
baseline_activation["population_cv"],
|
||||||
|
block_activation["population_cv"],
|
||||||
|
),
|
||||||
|
"baseline_first_to_last_ratio": baseline_activation[
|
||||||
|
"first_to_last_ratio"
|
||||||
|
],
|
||||||
|
"block_first_to_last_ratio": block_activation[
|
||||||
|
"first_to_last_ratio"
|
||||||
|
],
|
||||||
|
"baseline_imbalance": baseline_activation[
|
||||||
|
"imbalance_abs_log_ratio"
|
||||||
|
],
|
||||||
|
"block_imbalance": block_activation[
|
||||||
|
"imbalance_abs_log_ratio"
|
||||||
|
],
|
||||||
|
"relative_imbalance_reduction": relative_reduction(
|
||||||
|
baseline_activation["imbalance_abs_log_ratio"],
|
||||||
|
block_activation["imbalance_abs_log_ratio"],
|
||||||
|
),
|
||||||
|
"baseline_parameter_grad_cv": baseline_diagnostic[
|
||||||
|
"core_parameter_grad_statistics"
|
||||||
|
]["population_cv"],
|
||||||
|
"block_parameter_grad_cv": block_diagnostic[
|
||||||
|
"core_parameter_grad_statistics"
|
||||||
|
]["population_cv"],
|
||||||
|
}
|
||||||
|
)
|
||||||
|
verdict = depth_verdict(rows)
|
||||||
|
return {
|
||||||
|
"depth": depth,
|
||||||
|
"by_seed": rows,
|
||||||
|
"means": {
|
||||||
|
"baseline_bpc": mean(row["baseline_bpc"] for row in rows),
|
||||||
|
"block_bpc": mean(row["block_bpc"] for row in rows),
|
||||||
|
"block_minus_baseline_bpc": mean(
|
||||||
|
row["block_minus_baseline_bpc"] for row in rows
|
||||||
|
),
|
||||||
|
"baseline_cv": mean(row["baseline_cv"] for row in rows),
|
||||||
|
"block_cv": mean(row["block_cv"] for row in rows),
|
||||||
|
"relative_cv_reduction": mean(
|
||||||
|
row["relative_cv_reduction"] for row in rows
|
||||||
|
),
|
||||||
|
"baseline_imbalance": mean(
|
||||||
|
row["baseline_imbalance"] for row in rows
|
||||||
|
),
|
||||||
|
"block_imbalance": mean(
|
||||||
|
row["block_imbalance"] for row in rows
|
||||||
|
),
|
||||||
|
"relative_imbalance_reduction": mean(
|
||||||
|
row["relative_imbalance_reduction"] for row in rows
|
||||||
|
),
|
||||||
|
"block_to_baseline_activation_grad_mean": mean(
|
||||||
|
row["block_to_baseline_activation_grad_mean"] for row in rows
|
||||||
|
),
|
||||||
|
"baseline_parameter_grad_cv": mean(
|
||||||
|
row["baseline_parameter_grad_cv"] for row in rows
|
||||||
|
),
|
||||||
|
"block_parameter_grad_cv": mean(
|
||||||
|
row["block_parameter_grad_cv"] for row in rows
|
||||||
|
),
|
||||||
|
},
|
||||||
|
"verdict": verdict,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def compact_cell(run: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
diagnostics = []
|
||||||
|
for row in run["diagnostics"]:
|
||||||
|
compact = {
|
||||||
|
"step": row["step"],
|
||||||
|
"loss_nats": row["loss_nats"],
|
||||||
|
"bits_per_byte": row["bits_per_byte"],
|
||||||
|
"activation_grad_rms_by_block": row[
|
||||||
|
"activation_grad_rms_by_block"
|
||||||
|
],
|
||||||
|
"activation_grad_statistics": row[
|
||||||
|
"activation_grad_statistics"
|
||||||
|
],
|
||||||
|
"activation_output_rms_by_block": row[
|
||||||
|
"activation_output_rms_by_block"
|
||||||
|
],
|
||||||
|
"activation_output_statistics": row[
|
||||||
|
"activation_output_statistics"
|
||||||
|
],
|
||||||
|
"core_parameter_grad_rms_by_block": row[
|
||||||
|
"core_parameter_grad_rms_by_block"
|
||||||
|
],
|
||||||
|
"core_parameter_grad_statistics": row[
|
||||||
|
"core_parameter_grad_statistics"
|
||||||
|
],
|
||||||
|
}
|
||||||
|
if row["depth_weights"]:
|
||||||
|
compact["depth_weights"] = row["depth_weights"]
|
||||||
|
compact["output_weights"] = row["output_weights"]
|
||||||
|
diagnostics.append(compact)
|
||||||
|
return {
|
||||||
|
"architecture": run["architecture"],
|
||||||
|
"depth": run["depth"],
|
||||||
|
"seed": run["seed"],
|
||||||
|
"evaluations": run["evaluations"],
|
||||||
|
"diagnostics": diagnostics,
|
||||||
|
"timing": run["timing"],
|
||||||
|
"parameters": run["model"]["parameters"],
|
||||||
|
"hashes": run["hashes"],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
args = parse_args()
|
||||||
|
manifest = read_json(args.manifest)
|
||||||
|
manifest_hash = file_sha256(args.manifest)
|
||||||
|
if manifest["protocol_id"] != PROTOCOL_ID:
|
||||||
|
raise ValueError("manifest protocol mismatch")
|
||||||
|
if (
|
||||||
|
manifest["windows"]["formal_schedule_cells"] != 768_000
|
||||||
|
or manifest["windows"]["formal_steps"] != FORMAL_STEPS
|
||||||
|
or manifest["windows"]["formal_batch"] != FORMAL_BATCH
|
||||||
|
):
|
||||||
|
raise ValueError("manifest schedule budget mismatch")
|
||||||
|
|
||||||
|
runs: dict[tuple[int, str, int], dict[str, Any]] = {}
|
||||||
|
source_paths: dict[str, Path] = {}
|
||||||
|
formal_hashes: dict[str, str] = {}
|
||||||
|
for depth in DEPTHS:
|
||||||
|
for architecture in ARCHITECTURES:
|
||||||
|
for seed in SEEDS:
|
||||||
|
name = f"depth-{depth}-{architecture}-seed-{seed}.json"
|
||||||
|
path = args.formal_dir / name
|
||||||
|
run = read_json(path)
|
||||||
|
validate_run(
|
||||||
|
run,
|
||||||
|
path=path,
|
||||||
|
depth=depth,
|
||||||
|
architecture=architecture,
|
||||||
|
seed=seed,
|
||||||
|
manifest=manifest,
|
||||||
|
manifest_hash=manifest_hash,
|
||||||
|
)
|
||||||
|
runs[(depth, architecture, seed)] = run
|
||||||
|
public_name = f"formal-{name}"
|
||||||
|
source_paths[public_name] = path
|
||||||
|
formal_hashes[public_name] = file_sha256(path)
|
||||||
|
|
||||||
|
if sum(run["target_bytes_seen"] for run in runs.values()) != (
|
||||||
|
EXPECTED_TOTAL_TARGET_BYTES
|
||||||
|
):
|
||||||
|
raise ValueError("formal total target-byte budget mismatch")
|
||||||
|
|
||||||
|
common_initial_exact: dict[str, Any] = {}
|
||||||
|
input_gate_exact: dict[str, Any] = {}
|
||||||
|
for depth in DEPTHS:
|
||||||
|
for seed in SEEDS:
|
||||||
|
baseline = runs[(depth, "baseline", seed)]
|
||||||
|
block = runs[(depth, "block", seed)]
|
||||||
|
public_fields = (
|
||||||
|
"initial_public_parameter_structure",
|
||||||
|
"initial_public_parameter_tensors",
|
||||||
|
"initial_public_parameter_elements",
|
||||||
|
"initial_public_parameters",
|
||||||
|
)
|
||||||
|
exact = all(
|
||||||
|
baseline["hashes"][field] == block["hashes"][field]
|
||||||
|
for field in public_fields
|
||||||
|
)
|
||||||
|
gate_exact = (
|
||||||
|
baseline["manifest"]["input_gate_tensor_hashes"]
|
||||||
|
== block["manifest"]["input_gate_tensor_hashes"]
|
||||||
|
)
|
||||||
|
key = f"depth-{depth}-seed-{seed}"
|
||||||
|
common_initial_exact[key] = {
|
||||||
|
"exact": exact,
|
||||||
|
"baseline": {
|
||||||
|
field: baseline["hashes"][field] for field in public_fields
|
||||||
|
},
|
||||||
|
"block": {
|
||||||
|
field: block["hashes"][field] for field in public_fields
|
||||||
|
},
|
||||||
|
}
|
||||||
|
input_gate_exact[key] = {
|
||||||
|
"exact": gate_exact,
|
||||||
|
"hashes": baseline["manifest"]["input_gate_tensor_hashes"],
|
||||||
|
}
|
||||||
|
if not exact or not gate_exact:
|
||||||
|
raise ValueError(f"paired equality gate failed: {key}")
|
||||||
|
|
||||||
|
smoke_exact: dict[str, Any] = {}
|
||||||
|
smoke_hashes: dict[str, str] = {}
|
||||||
|
for depth in DEPTHS:
|
||||||
|
for architecture in ARCHITECTURES:
|
||||||
|
name = f"depth-{depth}-{architecture}.json"
|
||||||
|
left_path = args.smoke_a_dir / name
|
||||||
|
right_path = args.smoke_b_dir / name
|
||||||
|
left = read_json(left_path)
|
||||||
|
right = read_json(right_path)
|
||||||
|
left_selected = selected(left, SMOKE_COMPARE_FIELDS)
|
||||||
|
right_selected = selected(right, SMOKE_COMPARE_FIELDS)
|
||||||
|
exact = left_selected == right_selected
|
||||||
|
if (
|
||||||
|
not exact
|
||||||
|
or left["run_kind"] != "smoke"
|
||||||
|
or left["steps"] != 20
|
||||||
|
or not left["gradient_gate"]["passed"]
|
||||||
|
):
|
||||||
|
raise ValueError(f"smoke gate failed: {name}")
|
||||||
|
key = f"depth-{depth}-{architecture}"
|
||||||
|
smoke_exact[key] = {
|
||||||
|
"exact": exact,
|
||||||
|
"compare_sha256": canonical_sha256(left_selected),
|
||||||
|
"gradient_gate": left["gradient_gate"],
|
||||||
|
}
|
||||||
|
for label, path in (("a", left_path), ("b", right_path)):
|
||||||
|
public_name = f"smoke-{label}-{name}"
|
||||||
|
source_paths[public_name] = path
|
||||||
|
smoke_hashes[public_name] = file_sha256(path)
|
||||||
|
|
||||||
|
replay = read_json(args.replay)
|
||||||
|
replay_formal = runs[(32, "block", 2026073001)]
|
||||||
|
replay_left = selected(replay_formal, REPLAY_COMPARE_FIELDS)
|
||||||
|
replay_right = selected(replay, REPLAY_COMPARE_FIELDS)
|
||||||
|
replay_exact = replay_left == replay_right
|
||||||
|
if (
|
||||||
|
replay["run_kind"] != "replay"
|
||||||
|
or replay["depth"] != 32
|
||||||
|
or replay["architecture"] != "block"
|
||||||
|
or replay["seed"] != 2026073001
|
||||||
|
or not replay_exact
|
||||||
|
):
|
||||||
|
raise ValueError("formal replay gate failed")
|
||||||
|
replay_public_name = "replay-depth-32-block-seed-2026073001.json"
|
||||||
|
source_paths[replay_public_name] = args.replay
|
||||||
|
|
||||||
|
depth_summaries = {
|
||||||
|
str(depth): summarize_depth(depth, runs) for depth in DEPTHS
|
||||||
|
}
|
||||||
|
depth_labels = [
|
||||||
|
depth_summaries[str(depth)]["verdict"]["label"] for depth in DEPTHS
|
||||||
|
]
|
||||||
|
if all(
|
||||||
|
label == "joint directional support at this depth"
|
||||||
|
for label in depth_labels
|
||||||
|
):
|
||||||
|
overall_verdict = (
|
||||||
|
"scale-consistent directional support in this operationalization"
|
||||||
|
)
|
||||||
|
elif all(
|
||||||
|
label == "joint directional concern at this depth"
|
||||||
|
for label in depth_labels
|
||||||
|
):
|
||||||
|
overall_verdict = (
|
||||||
|
"scale-consistent directional concern in this operationalization"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
overall_verdict = "depth-dependent or inconclusive"
|
||||||
|
|
||||||
|
full = {
|
||||||
|
"schema_version": 1,
|
||||||
|
"protocol_id": PROTOCOL_ID,
|
||||||
|
"manifest": manifest,
|
||||||
|
"study": {
|
||||||
|
"architectures": list(ARCHITECTURES),
|
||||||
|
"depths": list(DEPTHS),
|
||||||
|
"seeds": list(SEEDS),
|
||||||
|
"diagnostic_steps": list(STEPS),
|
||||||
|
"formal_runs": len(runs),
|
||||||
|
"formal_target_bytes": EXPECTED_TOTAL_TARGET_BYTES,
|
||||||
|
"replay_target_bytes": TARGET_BYTES_PER_RUN,
|
||||||
|
"gradient_object": (
|
||||||
|
"RMS of d(mean token CE)/d(post-MLP Transformer-block output) "
|
||||||
|
"over batch×time×channel"
|
||||||
|
),
|
||||||
|
},
|
||||||
|
"depth_summaries": depth_summaries,
|
||||||
|
"overall_verdict": overall_verdict,
|
||||||
|
"gates": {
|
||||||
|
"common_initial_parameters": common_initial_exact,
|
||||||
|
"paired_input_tensors": input_gate_exact,
|
||||||
|
"smoke_exact": smoke_exact,
|
||||||
|
"replay": {
|
||||||
|
"exact": replay_exact,
|
||||||
|
"compare_fields": list(REPLAY_COMPARE_FIELDS),
|
||||||
|
"formal_compare_sha256": canonical_sha256(replay_left),
|
||||||
|
"replay_compare_sha256": canonical_sha256(replay_right),
|
||||||
|
"formal_final_model_state": replay_formal["hashes"][
|
||||||
|
"final_model_state"
|
||||||
|
],
|
||||||
|
"replay_final_model_state": replay["hashes"][
|
||||||
|
"final_model_state"
|
||||||
|
],
|
||||||
|
"formal_final_optimizer_state": replay_formal["hashes"][
|
||||||
|
"final_optimizer_state"
|
||||||
|
],
|
||||||
|
"replay_final_optimizer_state": replay["hashes"][
|
||||||
|
"final_optimizer_state"
|
||||||
|
],
|
||||||
|
},
|
||||||
|
},
|
||||||
|
"runs": {
|
||||||
|
f"depth-{depth}-{architecture}-seed-{seed}": run
|
||||||
|
for (depth, architecture, seed), run in sorted(runs.items())
|
||||||
|
},
|
||||||
|
}
|
||||||
|
full["canonical_sha256_without_self"] = canonical_sha256(full)
|
||||||
|
|
||||||
|
compact = {
|
||||||
|
"schema_version": 1,
|
||||||
|
"protocol_id": PROTOCOL_ID,
|
||||||
|
"study": full["study"],
|
||||||
|
"manifest_summary": {
|
||||||
|
"file_sha256": manifest_hash,
|
||||||
|
"dataset_revision": manifest["dataset"]["revision"],
|
||||||
|
"train_bytes_sha256": manifest["dataset"]["splits"]["train"][
|
||||||
|
"concatenated_sha256"
|
||||||
|
],
|
||||||
|
"formal_schedule_sha256": manifest["windows"][
|
||||||
|
"formal_schedule_sha256"
|
||||||
|
],
|
||||||
|
"validation_tensor_sha256": manifest["windows"][
|
||||||
|
"validation_tensor_sha256"
|
||||||
|
],
|
||||||
|
"diagnostic_tensor_sha256": manifest["windows"][
|
||||||
|
"diagnostic_tensor_sha256"
|
||||||
|
],
|
||||||
|
},
|
||||||
|
"depth_summaries": depth_summaries,
|
||||||
|
"overall_verdict": overall_verdict,
|
||||||
|
"replay_exact": replay_exact,
|
||||||
|
"cells": [
|
||||||
|
compact_cell(runs[(depth, architecture, seed)])
|
||||||
|
for depth in DEPTHS
|
||||||
|
for architecture in ARCHITECTURES
|
||||||
|
for seed in SEEDS
|
||||||
|
],
|
||||||
|
}
|
||||||
|
compact["canonical_sha256_without_self"] = canonical_sha256(compact)
|
||||||
|
|
||||||
|
args.raw_output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
for public_name, source_path in sorted(source_paths.items()):
|
||||||
|
target = args.raw_output_dir / public_name
|
||||||
|
temporary = target.with_suffix(target.suffix + ".tmp")
|
||||||
|
shutil.copyfile(source_path, temporary)
|
||||||
|
os.replace(temporary, target)
|
||||||
|
|
||||||
|
reproduction = {
|
||||||
|
"schema_version": 1,
|
||||||
|
"protocol_id": PROTOCOL_ID,
|
||||||
|
"manifest": {
|
||||||
|
"path": str(args.manifest),
|
||||||
|
"sha256": manifest_hash,
|
||||||
|
},
|
||||||
|
"protocol_sha256": file_sha256(
|
||||||
|
Path("research/K3_ATTNRES_GRADIENT_SCALE_PROTOCOL.md")
|
||||||
|
),
|
||||||
|
"definition_audit_sha256": file_sha256(
|
||||||
|
Path("research/K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md")
|
||||||
|
),
|
||||||
|
"runner_sha256": file_sha256(
|
||||||
|
Path("experiments/k3/attnres_gradient/train.py")
|
||||||
|
),
|
||||||
|
"analyzer_sha256": file_sha256(Path(__file__)),
|
||||||
|
"formal_raw_sha256": formal_hashes,
|
||||||
|
"smoke_raw_sha256": smoke_hashes,
|
||||||
|
"replay_raw_sha256": {
|
||||||
|
replay_public_name: file_sha256(args.replay)
|
||||||
|
},
|
||||||
|
"formal_target_bytes": EXPECTED_TOTAL_TARGET_BYTES,
|
||||||
|
"replay_target_bytes": TARGET_BYTES_PER_RUN,
|
||||||
|
"smoke_exact": smoke_exact,
|
||||||
|
"replay_exact": {
|
||||||
|
"exact": replay_exact,
|
||||||
|
"compare_sha256": canonical_sha256(replay_left),
|
||||||
|
"final_model_state": replay["hashes"]["final_model_state"],
|
||||||
|
"final_optimizer_state": replay["hashes"][
|
||||||
|
"final_optimizer_state"
|
||||||
|
],
|
||||||
|
},
|
||||||
|
"aggregate_sha256": full["canonical_sha256_without_self"],
|
||||||
|
"compact_sha256": compact["canonical_sha256_without_self"],
|
||||||
|
"overall_verdict": overall_verdict,
|
||||||
|
}
|
||||||
|
reproduction["canonical_sha256_without_self"] = canonical_sha256(
|
||||||
|
reproduction
|
||||||
|
)
|
||||||
|
|
||||||
|
atomic_json(args.output, full)
|
||||||
|
atomic_json(args.compact_output, compact)
|
||||||
|
atomic_json(args.reproduction_output, reproduction)
|
||||||
|
print(
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"formal_runs": len(runs),
|
||||||
|
"formal_target_bytes": EXPECTED_TOTAL_TARGET_BYTES,
|
||||||
|
"smoke_exact": all(
|
||||||
|
row["exact"] for row in smoke_exact.values()
|
||||||
|
),
|
||||||
|
"replay_exact": replay_exact,
|
||||||
|
"depth_verdicts": {
|
||||||
|
depth: depth_summaries[str(depth)]["verdict"]["label"]
|
||||||
|
for depth in DEPTHS
|
||||||
|
},
|
||||||
|
"overall_verdict": overall_verdict,
|
||||||
|
"aggregate_sha256": full[
|
||||||
|
"canonical_sha256_without_self"
|
||||||
|
],
|
||||||
|
"compact_sha256": compact[
|
||||||
|
"canonical_sha256_without_self"
|
||||||
|
],
|
||||||
|
"reproduction_sha256": reproduction[
|
||||||
|
"canonical_sha256_without_self"
|
||||||
|
],
|
||||||
|
},
|
||||||
|
ensure_ascii=False,
|
||||||
|
indent=2,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,206 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Freeze the byte-level corpus and window schedule for K3 AttnRes Round 05."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import hashlib
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import urllib.request
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import pyarrow.parquet as pq
|
||||||
|
|
||||||
|
|
||||||
|
PROTOCOL_ID = "llm-atlas-k3-attnres-gradient-scale-v1"
|
||||||
|
DATASET_REPO = "Salesforce/wikitext"
|
||||||
|
DATASET_REVISION = "b08601e04326c79dfdd32d625aee71d232d685c3"
|
||||||
|
DATASET_VARIANT = "wikitext-2-raw-v1"
|
||||||
|
SPLITS = ("train", "validation", "test")
|
||||||
|
SEEDS = (2026073001, 2026073002, 2026073003)
|
||||||
|
CONTEXT = 256
|
||||||
|
FORMAL_STEPS = 8000
|
||||||
|
FORMAL_BATCH = 32
|
||||||
|
VALIDATION_WINDOWS = 64
|
||||||
|
DIAGNOSTIC_WINDOWS = 16
|
||||||
|
GATE_STEPS = (0, 1, 7999)
|
||||||
|
|
||||||
|
|
||||||
|
def parse_args() -> argparse.Namespace:
|
||||||
|
parser = argparse.ArgumentParser()
|
||||||
|
parser.add_argument("--cache-dir", type=Path, required=True)
|
||||||
|
parser.add_argument("--manifest", type=Path, required=True)
|
||||||
|
return parser.parse_args()
|
||||||
|
|
||||||
|
|
||||||
|
def file_sha256(path: Path) -> str:
|
||||||
|
digest = hashlib.sha256()
|
||||||
|
with path.open("rb") as handle:
|
||||||
|
for block in iter(lambda: handle.read(1024 * 1024), b""):
|
||||||
|
digest.update(block)
|
||||||
|
return digest.hexdigest()
|
||||||
|
|
||||||
|
|
||||||
|
def atomic_json(path: Path, value: dict[str, Any]) -> None:
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
temporary = path.with_suffix(path.suffix + ".tmp")
|
||||||
|
temporary.write_text(
|
||||||
|
json.dumps(value, ensure_ascii=False, indent=2, sort_keys=True) + "\n"
|
||||||
|
)
|
||||||
|
os.replace(temporary, path)
|
||||||
|
|
||||||
|
|
||||||
|
def download(url: str, path: Path) -> None:
|
||||||
|
if path.exists():
|
||||||
|
return
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
temporary = path.with_suffix(path.suffix + ".part")
|
||||||
|
request = urllib.request.Request(
|
||||||
|
url,
|
||||||
|
headers={"User-Agent": "llm-atlas-k3-attnres-gradient-scale/1.0"},
|
||||||
|
)
|
||||||
|
with urllib.request.urlopen(request, timeout=120) as response:
|
||||||
|
with temporary.open("wb") as output:
|
||||||
|
while block := response.read(1024 * 1024):
|
||||||
|
output.write(block)
|
||||||
|
os.replace(temporary, path)
|
||||||
|
|
||||||
|
|
||||||
|
def hashed_start(fields: list[str], corpus_length: int) -> int:
|
||||||
|
value = int.from_bytes(
|
||||||
|
hashlib.sha256("\0".join(fields).encode()).digest()[:8], "big"
|
||||||
|
)
|
||||||
|
return value % (corpus_length - (CONTEXT + 1))
|
||||||
|
|
||||||
|
|
||||||
|
def fixed_window_start(label: str, index: int, corpus_length: int) -> int:
|
||||||
|
return hashed_start([PROTOCOL_ID, label, str(index)], corpus_length)
|
||||||
|
|
||||||
|
|
||||||
|
def train_window_start(seed: int, step: int, row: int, corpus_length: int) -> int:
|
||||||
|
return hashed_start(
|
||||||
|
[PROTOCOL_ID, "train-window", str(seed), str(step), str(row)],
|
||||||
|
corpus_length,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def concatenate_split(parquet_path: Path) -> tuple[bytes, int]:
|
||||||
|
table = pq.read_table(parquet_path, columns=["text"])
|
||||||
|
rows = table.column("text").to_pylist()
|
||||||
|
payload = b"".join(((row or "") + "\n").encode("utf-8") for row in rows)
|
||||||
|
return payload, len(rows)
|
||||||
|
|
||||||
|
|
||||||
|
def tensor_hash(payload: bytes, starts: list[int]) -> str:
|
||||||
|
digest = hashlib.sha256()
|
||||||
|
for start in starts:
|
||||||
|
digest.update(payload[start : start + CONTEXT + 1])
|
||||||
|
return digest.hexdigest()
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
args = parse_args()
|
||||||
|
args.cache_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
split_manifest: dict[str, Any] = {}
|
||||||
|
split_bytes: dict[str, bytes] = {}
|
||||||
|
for split in SPLITS:
|
||||||
|
relative = f"{DATASET_VARIANT}/{split}-00000-of-00001.parquet"
|
||||||
|
url = (
|
||||||
|
f"https://huggingface.co/datasets/{DATASET_REPO}/resolve/"
|
||||||
|
f"{DATASET_REVISION}/{relative}"
|
||||||
|
)
|
||||||
|
parquet_path = args.cache_dir / f"{split}.parquet"
|
||||||
|
download(url, parquet_path)
|
||||||
|
payload, rows = concatenate_split(parquet_path)
|
||||||
|
binary_path = args.cache_dir / f"{split}.bin"
|
||||||
|
if not binary_path.exists() or file_sha256(binary_path) != hashlib.sha256(
|
||||||
|
payload
|
||||||
|
).hexdigest():
|
||||||
|
temporary = binary_path.with_suffix(".bin.tmp")
|
||||||
|
temporary.write_bytes(payload)
|
||||||
|
os.replace(temporary, binary_path)
|
||||||
|
split_bytes[split] = payload
|
||||||
|
split_manifest[split] = {
|
||||||
|
"source_path": relative,
|
||||||
|
"source_url": url,
|
||||||
|
"parquet_bytes": parquet_path.stat().st_size,
|
||||||
|
"parquet_sha256": file_sha256(parquet_path),
|
||||||
|
"rows": rows,
|
||||||
|
"concatenated_bytes": len(payload),
|
||||||
|
"concatenated_sha256": hashlib.sha256(payload).hexdigest(),
|
||||||
|
"binary_path": str(binary_path),
|
||||||
|
"binary_sha256": file_sha256(binary_path),
|
||||||
|
}
|
||||||
|
|
||||||
|
train = split_bytes["train"]
|
||||||
|
validation = split_bytes["validation"]
|
||||||
|
schedule_digest = hashlib.sha256()
|
||||||
|
schedule_cells = 0
|
||||||
|
for seed in SEEDS:
|
||||||
|
for step in range(1, FORMAL_STEPS + 1):
|
||||||
|
for row in range(FORMAL_BATCH):
|
||||||
|
start = train_window_start(seed, step, row, len(train))
|
||||||
|
schedule_digest.update(start.to_bytes(8, "big"))
|
||||||
|
schedule_cells += 1
|
||||||
|
|
||||||
|
validation_starts = [
|
||||||
|
fixed_window_start("validation-window", index, len(validation))
|
||||||
|
for index in range(VALIDATION_WINDOWS)
|
||||||
|
]
|
||||||
|
diagnostic_starts = [
|
||||||
|
fixed_window_start("diagnostic-window", index, len(validation))
|
||||||
|
for index in range(DIAGNOSTIC_WINDOWS)
|
||||||
|
]
|
||||||
|
gate_tensor_hashes: dict[str, dict[str, str]] = {}
|
||||||
|
for seed in SEEDS:
|
||||||
|
gate_tensor_hashes[str(seed)] = {}
|
||||||
|
for step in GATE_STEPS:
|
||||||
|
starts = [
|
||||||
|
train_window_start(seed, step, row, len(train))
|
||||||
|
for row in range(FORMAL_BATCH)
|
||||||
|
]
|
||||||
|
gate_tensor_hashes[str(seed)][str(step)] = tensor_hash(train, starts)
|
||||||
|
|
||||||
|
manifest = {
|
||||||
|
"schema_version": 1,
|
||||||
|
"protocol_id": PROTOCOL_ID,
|
||||||
|
"status": "frozen-before-model-output",
|
||||||
|
"dataset": {
|
||||||
|
"repository": DATASET_REPO,
|
||||||
|
"revision": DATASET_REVISION,
|
||||||
|
"variant": DATASET_VARIANT,
|
||||||
|
"preprocessing": (
|
||||||
|
"parquet row order; (text or empty string) + LF; UTF-8; "
|
||||||
|
"no normalization; vocabulary is raw bytes 0..255"
|
||||||
|
),
|
||||||
|
"splits": split_manifest,
|
||||||
|
},
|
||||||
|
"windows": {
|
||||||
|
"context": CONTEXT,
|
||||||
|
"target_bytes_per_window": CONTEXT,
|
||||||
|
"seeds": list(SEEDS),
|
||||||
|
"formal_steps": FORMAL_STEPS,
|
||||||
|
"formal_batch": FORMAL_BATCH,
|
||||||
|
"formal_schedule_cells": schedule_cells,
|
||||||
|
"formal_schedule_sha256": schedule_digest.hexdigest(),
|
||||||
|
"validation_starts": validation_starts,
|
||||||
|
"validation_tensor_sha256": tensor_hash(
|
||||||
|
validation, validation_starts
|
||||||
|
),
|
||||||
|
"diagnostic_starts": diagnostic_starts,
|
||||||
|
"diagnostic_tensor_sha256": tensor_hash(
|
||||||
|
validation, diagnostic_starts
|
||||||
|
),
|
||||||
|
"gate_steps": list(GATE_STEPS),
|
||||||
|
"gate_training_tensor_sha256": gate_tensor_hashes,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
atomic_json(args.manifest, manifest)
|
||||||
|
print(json.dumps(manifest, ensure_ascii=False, indent=2))
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,167 @@
|
|||||||
|
{
|
||||||
|
"dataset": {
|
||||||
|
"preprocessing": "parquet row order; (text or empty string) + LF; UTF-8; no normalization; vocabulary is raw bytes 0..255",
|
||||||
|
"repository": "Salesforce/wikitext",
|
||||||
|
"revision": "b08601e04326c79dfdd32d625aee71d232d685c3",
|
||||||
|
"splits": {
|
||||||
|
"test": {
|
||||||
|
"binary_path": "/home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1/test.bin",
|
||||||
|
"binary_sha256": "bfe9eb16ab9987fb88bde4ea9a30a00f2a45db01dfc14bad78d05325789c4f12",
|
||||||
|
"concatenated_bytes": 1292014,
|
||||||
|
"concatenated_sha256": "bfe9eb16ab9987fb88bde4ea9a30a00f2a45db01dfc14bad78d05325789c4f12",
|
||||||
|
"parquet_bytes": 732610,
|
||||||
|
"parquet_sha256": "5f1bea067869d04849c0f975a2b29c4ff47d867f484f5010ea5e861eab246d91",
|
||||||
|
"rows": 4358,
|
||||||
|
"source_path": "wikitext-2-raw-v1/test-00000-of-00001.parquet",
|
||||||
|
"source_url": "https://huggingface.co/datasets/Salesforce/wikitext/resolve/b08601e04326c79dfdd32d625aee71d232d685c3/wikitext-2-raw-v1/test-00000-of-00001.parquet"
|
||||||
|
},
|
||||||
|
"train": {
|
||||||
|
"binary_path": "/home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1/train.bin",
|
||||||
|
"binary_sha256": "0ca7d3e74dbe44564ea5942b85232f1bbcb525c9cd481cd5d28a87ee90e7e9b4",
|
||||||
|
"concatenated_bytes": 10951563,
|
||||||
|
"concatenated_sha256": "0ca7d3e74dbe44564ea5942b85232f1bbcb525c9cd481cd5d28a87ee90e7e9b4",
|
||||||
|
"parquet_bytes": 6357543,
|
||||||
|
"parquet_sha256": "e83889baabc497075506f91975be5fac0d45c5290b6b20582c8cd1e853d0c9f7",
|
||||||
|
"rows": 36718,
|
||||||
|
"source_path": "wikitext-2-raw-v1/train-00000-of-00001.parquet",
|
||||||
|
"source_url": "https://huggingface.co/datasets/Salesforce/wikitext/resolve/b08601e04326c79dfdd32d625aee71d232d685c3/wikitext-2-raw-v1/train-00000-of-00001.parquet"
|
||||||
|
},
|
||||||
|
"validation": {
|
||||||
|
"binary_path": "/home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1/validation.bin",
|
||||||
|
"binary_sha256": "a42356f6a8ff1d25daf25ec9db49e10a537c265581b61c74604bb63231dee719",
|
||||||
|
"concatenated_bytes": 1148008,
|
||||||
|
"concatenated_sha256": "a42356f6a8ff1d25daf25ec9db49e10a537c265581b61c74604bb63231dee719",
|
||||||
|
"parquet_bytes": 657209,
|
||||||
|
"parquet_sha256": "204929b7ff9d6184953f867dedb860e40aa69c078fc1e54b3baaa8fb28511c4c",
|
||||||
|
"rows": 3760,
|
||||||
|
"source_path": "wikitext-2-raw-v1/validation-00000-of-00001.parquet",
|
||||||
|
"source_url": "https://huggingface.co/datasets/Salesforce/wikitext/resolve/b08601e04326c79dfdd32d625aee71d232d685c3/wikitext-2-raw-v1/validation-00000-of-00001.parquet"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"variant": "wikitext-2-raw-v1"
|
||||||
|
},
|
||||||
|
"protocol_id": "llm-atlas-k3-attnres-gradient-scale-v1",
|
||||||
|
"schema_version": 1,
|
||||||
|
"status": "frozen-before-model-output",
|
||||||
|
"windows": {
|
||||||
|
"context": 256,
|
||||||
|
"diagnostic_starts": [
|
||||||
|
399861,
|
||||||
|
210983,
|
||||||
|
449025,
|
||||||
|
1098024,
|
||||||
|
323754,
|
||||||
|
152932,
|
||||||
|
1091551,
|
||||||
|
1078021,
|
||||||
|
415985,
|
||||||
|
612288,
|
||||||
|
910624,
|
||||||
|
530272,
|
||||||
|
827285,
|
||||||
|
765798,
|
||||||
|
1086876,
|
||||||
|
1035447
|
||||||
|
],
|
||||||
|
"diagnostic_tensor_sha256": "21117e31db302b10d67b63f035665dc8f220b879d216ccd12b7d2ba86e7b1716",
|
||||||
|
"formal_batch": 32,
|
||||||
|
"formal_schedule_cells": 768000,
|
||||||
|
"formal_schedule_sha256": "5041e09b167f229248d2462324e8c254b8f5938975f135dcd8192b00a54a4f4e",
|
||||||
|
"formal_steps": 8000,
|
||||||
|
"gate_steps": [
|
||||||
|
0,
|
||||||
|
1,
|
||||||
|
7999
|
||||||
|
],
|
||||||
|
"gate_training_tensor_sha256": {
|
||||||
|
"2026073001": {
|
||||||
|
"0": "52fdd6885cc2bef8e29cedec8c293e1ea71f63fe3cd0b83f627fe640e19a95c5",
|
||||||
|
"1": "9d0a960595a3f57cd18834d880bde56fa8dcbb0b1cbed26bcc9bbf773a67949c",
|
||||||
|
"7999": "8f4f04a889d1f8c9eb75e917dc7cd6ddef1196bdea94e87466270db7be1289ee"
|
||||||
|
},
|
||||||
|
"2026073002": {
|
||||||
|
"0": "156813ff7ab93736c8dba340711a9633b6c952d0f54df30f277e254705cc8b82",
|
||||||
|
"1": "9ee5a434bdf417d658a1485135d48749e98c86f03239eef86dbd81a9d1309da0",
|
||||||
|
"7999": "bca78ffefa000dc3693a790d65933251646c70facc64b42007d500c19bb90977"
|
||||||
|
},
|
||||||
|
"2026073003": {
|
||||||
|
"0": "8ba13ad55eb54501ec443430446793ef41dea21fd067b2a5b0c84eda202ed0ba",
|
||||||
|
"1": "326a12f07637fe70fd39fa2758b389e91dacc7f2f2ede708b80e09c9d4d1c39d",
|
||||||
|
"7999": "7d6b10c3febfb697444909f695d3a45297789022714f8b71dadfec4d6236aea6"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"seeds": [
|
||||||
|
2026073001,
|
||||||
|
2026073002,
|
||||||
|
2026073003
|
||||||
|
],
|
||||||
|
"target_bytes_per_window": 256,
|
||||||
|
"validation_starts": [
|
||||||
|
19785,
|
||||||
|
600807,
|
||||||
|
1029500,
|
||||||
|
319878,
|
||||||
|
652204,
|
||||||
|
1072662,
|
||||||
|
679625,
|
||||||
|
1027264,
|
||||||
|
500250,
|
||||||
|
944134,
|
||||||
|
187313,
|
||||||
|
834295,
|
||||||
|
968556,
|
||||||
|
550645,
|
||||||
|
239009,
|
||||||
|
452519,
|
||||||
|
356354,
|
||||||
|
134015,
|
||||||
|
64555,
|
||||||
|
397632,
|
||||||
|
203140,
|
||||||
|
346032,
|
||||||
|
314812,
|
||||||
|
10817,
|
||||||
|
1141274,
|
||||||
|
807645,
|
||||||
|
417975,
|
||||||
|
870687,
|
||||||
|
377265,
|
||||||
|
635426,
|
||||||
|
597238,
|
||||||
|
805324,
|
||||||
|
12300,
|
||||||
|
264343,
|
||||||
|
84743,
|
||||||
|
596894,
|
||||||
|
188690,
|
||||||
|
992517,
|
||||||
|
854512,
|
||||||
|
427504,
|
||||||
|
94167,
|
||||||
|
296670,
|
||||||
|
760313,
|
||||||
|
912279,
|
||||||
|
1054297,
|
||||||
|
81970,
|
||||||
|
419690,
|
||||||
|
971472,
|
||||||
|
1041491,
|
||||||
|
669963,
|
||||||
|
735537,
|
||||||
|
434513,
|
||||||
|
169153,
|
||||||
|
6229,
|
||||||
|
136413,
|
||||||
|
1098303,
|
||||||
|
400950,
|
||||||
|
457810,
|
||||||
|
659776,
|
||||||
|
911665,
|
||||||
|
909832,
|
||||||
|
532969,
|
||||||
|
555820,
|
||||||
|
1019138
|
||||||
|
],
|
||||||
|
"validation_tensor_sha256": "f459316f13078a163b47c133511bb7181e05170ab89516e196490113893ce338"
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,217 @@
|
|||||||
|
{
|
||||||
|
"aggregate_sha256": "69be133c8f5b11f8a58e831bc03f5432e737ed21efe89bf186784ead7c7f7b51",
|
||||||
|
"analyzer_sha256": "017d38d92bfd5bc26e10028015f9bf85d28516fb5c5bae06d7bf315b5eca2fc8",
|
||||||
|
"canonical_sha256_without_self": "addb2e5990af4bbec4b88b4654fc72464f4df5b57fd00e46cad1febbb042f68a",
|
||||||
|
"compact_sha256": "8cdb71808a429e5f9513c1c6c9426bdc93e30f237001a1b47a2a2fba574a044f",
|
||||||
|
"definition_audit_sha256": "79221c5648ba280d9177386d2fcfb2fa653b3dd14405185b5997754db2961cc6",
|
||||||
|
"formal_raw_sha256": {
|
||||||
|
"formal-depth-16-baseline-seed-2026073001.json": "c9144ddb458010868668ab0bb1272d4e5051300936d9b21dccee75c32a9e46de",
|
||||||
|
"formal-depth-16-baseline-seed-2026073002.json": "939cad5054960ba8a72b50b4b250b94c3d22a6b46796f669c6e4f033d59fb9de",
|
||||||
|
"formal-depth-16-baseline-seed-2026073003.json": "afd9a60ff323e97d21b24b413d9b9fa97b8d5f960bfc88d2040502aa5ad96a1e",
|
||||||
|
"formal-depth-16-block-seed-2026073001.json": "23e3c68e54ec55d981066ad5730316db9084a6637a8064763952ba109ea1f715",
|
||||||
|
"formal-depth-16-block-seed-2026073002.json": "027f8786f530d6ad5d910f0aad3d0f287887d6457c670bc8b50af7248be1487d",
|
||||||
|
"formal-depth-16-block-seed-2026073003.json": "254b637f403a4c61e8e8bb2dd083349406b6baa7a3d8a616e37a9ab6f8ebaeb9",
|
||||||
|
"formal-depth-32-baseline-seed-2026073001.json": "232ed3dfbbbc898c42622c4a9aee8400f76cb773c576bf1c522cdbb19ead998c",
|
||||||
|
"formal-depth-32-baseline-seed-2026073002.json": "f4fb5fca608a6623ce4c9707714a8a9ebd2c10f6e63f54e243ec3c06d18b1fff",
|
||||||
|
"formal-depth-32-baseline-seed-2026073003.json": "1b0bd279b165c3a2431685d8c1e6bc36ff72eb55333d5de689dca4f89506bb5e",
|
||||||
|
"formal-depth-32-block-seed-2026073001.json": "29e1d638b481619c7b32de402122523db8b17cc1fc67a8881fa1ba132a1d5c38",
|
||||||
|
"formal-depth-32-block-seed-2026073002.json": "21199deb2199395061e51220fd8c7afd04a1135a6381e406da9b5795e3ad5032",
|
||||||
|
"formal-depth-32-block-seed-2026073003.json": "c0d7f1bcfa9fa7f8f3134ca4571bdf23a951182d03b7de9611a6b4b89667d4ac"
|
||||||
|
},
|
||||||
|
"formal_target_bytes": 786432000,
|
||||||
|
"manifest": {
|
||||||
|
"path": "experiments/k3/attnres_gradient/manifest.json",
|
||||||
|
"sha256": "080afb17d1e036c0bba0a799fdb8b98ee4ad652bd42dd1b3b67110dd2ede6371"
|
||||||
|
},
|
||||||
|
"overall_verdict": "depth-dependent or inconclusive",
|
||||||
|
"protocol_id": "llm-atlas-k3-attnres-gradient-scale-v1",
|
||||||
|
"protocol_sha256": "f772629b3b82975b6756721c3a3bb57cc4171dfa26ba5b1e8043dc91c9dcce22",
|
||||||
|
"replay_exact": {
|
||||||
|
"compare_sha256": "46300a452840a9dc6a5180efe7949cf3d81cc4471ed4a942da407e9343064817",
|
||||||
|
"exact": true,
|
||||||
|
"final_model_state": "3f0b97ece3a15571ba3d656f589f512ca0bb9e20083c9f58a42ccaee14892f59",
|
||||||
|
"final_optimizer_state": "ed03e6fbd4a12d8b063dcb22e0437754285f54d585374cd52fbd534f05d24637"
|
||||||
|
},
|
||||||
|
"replay_raw_sha256": {
|
||||||
|
"replay-depth-32-block-seed-2026073001.json": "5cea76d68a642c2b4b4f189a813b2e4af9361a939ddd31a68dfdb0ab090b4aad"
|
||||||
|
},
|
||||||
|
"replay_target_bytes": 65536000,
|
||||||
|
"runner_sha256": "04ae69e10c58972c9193c2d31c7e09d924a0d0e834107afa4ba128c64ac5800f",
|
||||||
|
"schema_version": 1,
|
||||||
|
"smoke_exact": {
|
||||||
|
"depth-16-baseline": {
|
||||||
|
"compare_sha256": "525cdcd9a79ec61096ecf64fedcbf93144d3168b434e2f8fd384c4374cbebd2a",
|
||||||
|
"exact": true,
|
||||||
|
"gradient_gate": {
|
||||||
|
"first_to_last_ratio_abs_delta": 0.0,
|
||||||
|
"max_abs_scale_ratio_error": 0.0,
|
||||||
|
"normalized_spectrum_max_abs_delta": 0.0,
|
||||||
|
"passed": true,
|
||||||
|
"per_block_scale_ratios": [
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0
|
||||||
|
],
|
||||||
|
"population_cv_abs_delta": 0.0,
|
||||||
|
"thresholds": {
|
||||||
|
"scale_ratio_abs": 1e-05,
|
||||||
|
"shape_abs": 1e-06
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"depth-16-block": {
|
||||||
|
"compare_sha256": "e4330a0d878d10b474b9aeb0be58131f801ebe2e5dd9553d4713c3d16b6590cc",
|
||||||
|
"exact": true,
|
||||||
|
"gradient_gate": {
|
||||||
|
"first_to_last_ratio_abs_delta": 0.0,
|
||||||
|
"max_abs_scale_ratio_error": 0.0,
|
||||||
|
"normalized_spectrum_max_abs_delta": 0.0,
|
||||||
|
"passed": true,
|
||||||
|
"per_block_scale_ratios": [
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0
|
||||||
|
],
|
||||||
|
"population_cv_abs_delta": 0.0,
|
||||||
|
"thresholds": {
|
||||||
|
"scale_ratio_abs": 1e-05,
|
||||||
|
"shape_abs": 1e-06
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"depth-32-baseline": {
|
||||||
|
"compare_sha256": "8ee0dc37e88742b6705968f5f2845a3a1df1250f5e0fd77be6b909236ea28e56",
|
||||||
|
"exact": true,
|
||||||
|
"gradient_gate": {
|
||||||
|
"first_to_last_ratio_abs_delta": 0.0,
|
||||||
|
"max_abs_scale_ratio_error": 0.0,
|
||||||
|
"normalized_spectrum_max_abs_delta": 0.0,
|
||||||
|
"passed": true,
|
||||||
|
"per_block_scale_ratios": [
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0
|
||||||
|
],
|
||||||
|
"population_cv_abs_delta": 0.0,
|
||||||
|
"thresholds": {
|
||||||
|
"scale_ratio_abs": 1e-05,
|
||||||
|
"shape_abs": 1e-06
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"depth-32-block": {
|
||||||
|
"compare_sha256": "289073b1940d4747b7e049f8d767f9c22a49867d37403ff6ef7746a55bd1734c",
|
||||||
|
"exact": true,
|
||||||
|
"gradient_gate": {
|
||||||
|
"first_to_last_ratio_abs_delta": 0.0,
|
||||||
|
"max_abs_scale_ratio_error": 0.0,
|
||||||
|
"normalized_spectrum_max_abs_delta": 0.0,
|
||||||
|
"passed": true,
|
||||||
|
"per_block_scale_ratios": [
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0
|
||||||
|
],
|
||||||
|
"population_cv_abs_delta": 0.0,
|
||||||
|
"thresholds": {
|
||||||
|
"scale_ratio_abs": 1e-05,
|
||||||
|
"shape_abs": 1e-06
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"smoke_raw_sha256": {
|
||||||
|
"smoke-a-depth-16-baseline.json": "fe0ff3ed21bf9b86d238d7a22e99b6e07ba99135b58d29eb6b205f2602e64556",
|
||||||
|
"smoke-a-depth-16-block.json": "faa06a2bd02135858d3a2503e1ecfb34187b4ee3b4ce3428ed473a55b15b6705",
|
||||||
|
"smoke-a-depth-32-baseline.json": "e327a6f5db87515c1b29b2808205e23b81228b08ce356b98804717e2e55eef1c",
|
||||||
|
"smoke-a-depth-32-block.json": "2d9a97c2d6272c928d5b5bc01febd142d4b2f2ad3080330e849f2731a8d080c2",
|
||||||
|
"smoke-b-depth-16-baseline.json": "fe0ff3ed21bf9b86d238d7a22e99b6e07ba99135b58d29eb6b205f2602e64556",
|
||||||
|
"smoke-b-depth-16-block.json": "faa06a2bd02135858d3a2503e1ecfb34187b4ee3b4ce3428ed473a55b15b6705",
|
||||||
|
"smoke-b-depth-32-baseline.json": "e327a6f5db87515c1b29b2808205e23b81228b08ce356b98804717e2e55eef1c",
|
||||||
|
"smoke-b-depth-32-block.json": "2d9a97c2d6272c928d5b5bc01febd142d4b2f2ad3080330e849f2731a8d080c2"
|
||||||
|
}
|
||||||
|
}
|
||||||
+7366
File diff suppressed because it is too large
Load Diff
+7366
File diff suppressed because it is too large
Load Diff
+7366
File diff suppressed because it is too large
Load Diff
+9616
File diff suppressed because it is too large
Load Diff
+9616
File diff suppressed because it is too large
Load Diff
+9616
File diff suppressed because it is too large
Load Diff
+8614
File diff suppressed because it is too large
Load Diff
+8614
File diff suppressed because it is too large
Load Diff
+8614
File diff suppressed because it is too large
Load Diff
+13072
File diff suppressed because it is too large
Load Diff
+13072
File diff suppressed because it is too large
Load Diff
+13072
File diff suppressed because it is too large
Load Diff
+13072
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,700 @@
|
|||||||
|
{
|
||||||
|
"architecture": "baseline",
|
||||||
|
"batch_size": 32,
|
||||||
|
"canonical_sha256_without_self": "0ae9ad1697ca15ec4c84270ad82b9fa98056f14f20fc7372c4b9886c29845415",
|
||||||
|
"depth": 16,
|
||||||
|
"diagnostics": [
|
||||||
|
{
|
||||||
|
"activation_grad_rms_by_block": [
|
||||||
|
0.0002865509013645351,
|
||||||
|
0.00024785567075014114,
|
||||||
|
0.00021551651298068464,
|
||||||
|
0.00019683866412378848,
|
||||||
|
0.0001823508064262569,
|
||||||
|
0.0001679604029050097,
|
||||||
|
0.00015584587526973337,
|
||||||
|
0.00014737028686795384,
|
||||||
|
0.00014130656199995428,
|
||||||
|
0.00013612695329356939,
|
||||||
|
0.00013129493163432926,
|
||||||
|
0.00012807638267986476,
|
||||||
|
0.00012550120300147682,
|
||||||
|
0.00012324050476308912,
|
||||||
|
0.00012108208466088399,
|
||||||
|
0.00011888353037647903
|
||||||
|
],
|
||||||
|
"activation_grad_statistics": {
|
||||||
|
"first_quartile_mean": 0.00023669043730478734,
|
||||||
|
"first_to_last_ratio": 1.9372775995887173,
|
||||||
|
"imbalance_abs_log_ratio": 0.6612836883477485,
|
||||||
|
"last_quartile_mean": 0.00012217683070048224,
|
||||||
|
"mean": 0.00016411257956860936,
|
||||||
|
"normalized": [
|
||||||
|
1.7460629899168627,
|
||||||
|
1.5102783187106137,
|
||||||
|
1.313223602646409,
|
||||||
|
1.1994124072706904,
|
||||||
|
1.1111324123086055,
|
||||||
|
1.0234462424910682,
|
||||||
|
0.9496278449793059,
|
||||||
|
0.8979828801383491,
|
||||||
|
0.8610343117596252,
|
||||||
|
0.8294729974472175,
|
||||||
|
0.8000296624393728,
|
||||||
|
0.7804178266926869,
|
||||||
|
0.7647262832098098,
|
||||||
|
0.7509509940495869,
|
||||||
|
0.7377989242455608,
|
||||||
|
0.7244023016942358
|
||||||
|
],
|
||||||
|
"population_cv": 0.29361872380036635
|
||||||
|
},
|
||||||
|
"activation_output_rms_by_block": [
|
||||||
|
0.02866268903017044,
|
||||||
|
0.029099803417921066,
|
||||||
|
0.02953805774450302,
|
||||||
|
0.030088091269135475,
|
||||||
|
0.030644793063402176,
|
||||||
|
0.03131929785013199,
|
||||||
|
0.03213750571012497,
|
||||||
|
0.03286394104361534,
|
||||||
|
0.033869802951812744,
|
||||||
|
0.03499744459986687,
|
||||||
|
0.03588006645441055,
|
||||||
|
0.037365153431892395,
|
||||||
|
0.03789033368229866,
|
||||||
|
0.039445169270038605,
|
||||||
|
0.04050002992153168,
|
||||||
|
0.04121527820825577
|
||||||
|
],
|
||||||
|
"activation_output_statistics": {
|
||||||
|
"first_quartile_mean": 0.0293471603654325,
|
||||||
|
"first_to_last_ratio": 0.7380574840395957,
|
||||||
|
"imbalance_abs_log_ratio": 0.3037335657624931,
|
||||||
|
"last_quartile_mean": 0.03976270277053118,
|
||||||
|
"mean": 0.034094841103069484,
|
||||||
|
"normalized": [
|
||||||
|
0.8406752488894867,
|
||||||
|
0.8534957922212247,
|
||||||
|
0.8663497699023965,
|
||||||
|
0.8824822259232266,
|
||||||
|
0.8988102619619861,
|
||||||
|
0.9185934539320196,
|
||||||
|
0.942591449919727,
|
||||||
|
0.9638977622528551,
|
||||||
|
0.9933996421752943,
|
||||||
|
1.0264733158329964,
|
||||||
|
1.0523605710888722,
|
||||||
|
1.0959180985456622,
|
||||||
|
1.1113216092650298,
|
||||||
|
1.1569248600043318,
|
||||||
|
1.1878638706395246,
|
||||||
|
1.2088420674453664
|
||||||
|
],
|
||||||
|
"population_cv": 0.11952987076282337
|
||||||
|
},
|
||||||
|
"bits_per_byte": 8.096463027059821,
|
||||||
|
"branch_output_rms_by_sublayer": [
|
||||||
|
0.0033055038657039404,
|
||||||
|
0.0037518907338380814,
|
||||||
|
0.0033751516602933407,
|
||||||
|
0.003911525942385197,
|
||||||
|
0.004229962360113859,
|
||||||
|
0.0038461871445178986,
|
||||||
|
0.004243654198944569,
|
||||||
|
0.0038366704247891903,
|
||||||
|
0.004372371360659599,
|
||||||
|
0.003851204412057996,
|
||||||
|
0.005078981164842844,
|
||||||
|
0.003818925702944398,
|
||||||
|
0.0054668826051056385,
|
||||||
|
0.0038749484810978174,
|
||||||
|
0.005823117680847645,
|
||||||
|
0.0037566579412668943,
|
||||||
|
0.0065947300754487514,
|
||||||
|
0.003876061411574483,
|
||||||
|
0.006763988174498081,
|
||||||
|
0.003795720636844635,
|
||||||
|
0.006863070651888847,
|
||||||
|
0.0038706199266016483,
|
||||||
|
0.007856340147554874,
|
||||||
|
0.003839300014078617,
|
||||||
|
0.00830968376249075,
|
||||||
|
0.003927029203623533,
|
||||||
|
0.009655521251261234,
|
||||||
|
0.0038513424806296825,
|
||||||
|
0.009004893712699413,
|
||||||
|
0.0039484030567109585,
|
||||||
|
0.008608299307525158,
|
||||||
|
0.003791053779423237
|
||||||
|
],
|
||||||
|
"capture": {
|
||||||
|
"all_gradients_finite": true,
|
||||||
|
"all_gradients_present": true,
|
||||||
|
"count": 16,
|
||||||
|
"dtypes": [
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32"
|
||||||
|
],
|
||||||
|
"position": "post-MLP Transformer-block output; Block AttnRes is captured before aggregation-partial reset",
|
||||||
|
"shape": [
|
||||||
|
16,
|
||||||
|
256,
|
||||||
|
192
|
||||||
|
],
|
||||||
|
"storage_unique": true
|
||||||
|
},
|
||||||
|
"core_parameter_grad_rms_by_block": [
|
||||||
|
0.00809059897248305,
|
||||||
|
0.007313812062277777,
|
||||||
|
0.007090710224412361,
|
||||||
|
0.0067033738367094624,
|
||||||
|
0.006535407867454384,
|
||||||
|
0.006785021795736036,
|
||||||
|
0.006295993569919972,
|
||||||
|
0.005509300326230647,
|
||||||
|
0.006231736048910802,
|
||||||
|
0.006047390980174007,
|
||||||
|
0.005629405359532731,
|
||||||
|
0.005784042737956458,
|
||||||
|
0.0059591736113607996,
|
||||||
|
0.0061983173180850636,
|
||||||
|
0.006571690039252182,
|
||||||
|
0.005060205437758869
|
||||||
|
],
|
||||||
|
"core_parameter_grad_statistics": {
|
||||||
|
"first_quartile_mean": 0.007299623773970663,
|
||||||
|
"first_to_last_ratio": 1.227374872012571,
|
||||||
|
"imbalance_abs_log_ratio": 0.2048776382299483,
|
||||||
|
"last_quartile_mean": 0.005947346601614228,
|
||||||
|
"mean": 0.006362886261765913,
|
||||||
|
"normalized": [
|
||||||
|
1.271529717747562,
|
||||||
|
1.1494488132258316,
|
||||||
|
1.1143858200043408,
|
||||||
|
1.05351149791715,
|
||||||
|
1.0271137340180259,
|
||||||
|
1.066343404015675,
|
||||||
|
0.9894870520870547,
|
||||||
|
0.8658492545019393,
|
||||||
|
0.9793882512652816,
|
||||||
|
0.9504163254515969,
|
||||||
|
0.8847251275509024,
|
||||||
|
0.909028151691524,
|
||||||
|
0.9365519618304368,
|
||||||
|
0.9741361173356609,
|
||||||
|
1.0328158902888074,
|
||||||
|
0.7952688810682109
|
||||||
|
],
|
||||||
|
"population_cv": 0.11384977604868386
|
||||||
|
},
|
||||||
|
"depth_weights": [],
|
||||||
|
"layer_input_rms_by_sublayer": [
|
||||||
|
0.02823694422841072,
|
||||||
|
0.028404271230101585,
|
||||||
|
0.02866268903017044,
|
||||||
|
0.02886480651795864,
|
||||||
|
0.029099803417921066,
|
||||||
|
0.029365191236138344,
|
||||||
|
0.02953805774450302,
|
||||||
|
0.029881296679377556,
|
||||||
|
0.030088091269135475,
|
||||||
|
0.030441828072071075,
|
||||||
|
0.030644793063402176,
|
||||||
|
0.03106885403394699,
|
||||||
|
0.03131929785013199,
|
||||||
|
0.031892918050289154,
|
||||||
|
0.03213750571012497,
|
||||||
|
0.03266792371869087,
|
||||||
|
0.03286394104361534,
|
||||||
|
0.03364725783467293,
|
||||||
|
0.033869802951812744,
|
||||||
|
0.03485054895281792,
|
||||||
|
0.03499744459986687,
|
||||||
|
0.03562505170702934,
|
||||||
|
0.03588006645441055,
|
||||||
|
0.03719272464513779,
|
||||||
|
0.037365153431892395,
|
||||||
|
0.03768136352300644,
|
||||||
|
0.03789033368229866,
|
||||||
|
0.03927159309387207,
|
||||||
|
0.039445169270038605,
|
||||||
|
0.04041972756385803,
|
||||||
|
0.04050002992153168,
|
||||||
|
0.04110744968056679
|
||||||
|
],
|
||||||
|
"loss_nats": 5.6120405197143555,
|
||||||
|
"loss_scale": 1.0,
|
||||||
|
"output_weights": null,
|
||||||
|
"step": 0,
|
||||||
|
"stream_state_rms_by_sublayer": [
|
||||||
|
0.028404271230101585,
|
||||||
|
0.02866268903017044,
|
||||||
|
0.02886480651795864,
|
||||||
|
0.029099803417921066,
|
||||||
|
0.029365191236138344,
|
||||||
|
0.02953805774450302,
|
||||||
|
0.029881296679377556,
|
||||||
|
0.030088091269135475,
|
||||||
|
0.030441828072071075,
|
||||||
|
0.030644793063402176,
|
||||||
|
0.03106885403394699,
|
||||||
|
0.03131929785013199,
|
||||||
|
0.031892918050289154,
|
||||||
|
0.03213750571012497,
|
||||||
|
0.03266792371869087,
|
||||||
|
0.03286394104361534,
|
||||||
|
0.03364725783467293,
|
||||||
|
0.033869802951812744,
|
||||||
|
0.03485054895281792,
|
||||||
|
0.03499744459986687,
|
||||||
|
0.03562505170702934,
|
||||||
|
0.03588006645441055,
|
||||||
|
0.03719272464513779,
|
||||||
|
0.037365153431892395,
|
||||||
|
0.03768136352300644,
|
||||||
|
0.03789033368229866,
|
||||||
|
0.03927159309387207,
|
||||||
|
0.039445169270038605,
|
||||||
|
0.04041972756385803,
|
||||||
|
0.04050002992153168,
|
||||||
|
0.04110744968056679,
|
||||||
|
0.04121527820825577
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"activation_grad_rms_by_block": [
|
||||||
|
5.69471885683015e-05,
|
||||||
|
5.527245593839325e-05,
|
||||||
|
5.393734318204224e-05,
|
||||||
|
5.292302375892177e-05,
|
||||||
|
5.23102717124857e-05,
|
||||||
|
5.178990977583453e-05,
|
||||||
|
5.1564093155320734e-05,
|
||||||
|
5.129844430484809e-05,
|
||||||
|
5.121947833686136e-05,
|
||||||
|
5.116340616950765e-05,
|
||||||
|
5.109804988023825e-05,
|
||||||
|
5.104695082991384e-05,
|
||||||
|
5.09646451973822e-05,
|
||||||
|
5.100828275317326e-05,
|
||||||
|
5.106439857627265e-05,
|
||||||
|
5.11144389747642e-05
|
||||||
|
],
|
||||||
|
"activation_grad_statistics": {
|
||||||
|
"first_quartile_mean": 5.477000286191469e-05,
|
||||||
|
"first_to_last_ratio": 1.0731232762518041,
|
||||||
|
"imbalance_abs_log_ratio": 0.07057334637995329,
|
||||||
|
"last_quartile_mean": 5.103794137539808e-05,
|
||||||
|
"mean": 5.2170148819641327e-05,
|
||||||
|
"normalized": [
|
||||||
|
1.0915665348238701,
|
||||||
|
1.0594651767139285,
|
||||||
|
1.033873669184083,
|
||||||
|
1.0144311441756324,
|
||||||
|
1.002685882559561,
|
||||||
|
0.9927115591500164,
|
||||||
|
0.988383094968431,
|
||||||
|
0.9832911246274795,
|
||||||
|
0.9817775010367221,
|
||||||
|
0.9807027069519371,
|
||||||
|
0.9794499543578176,
|
||||||
|
0.9784704852268963,
|
||||||
|
0.9768928467805085,
|
||||||
|
0.9777292936141551,
|
||||||
|
0.9788049244944386,
|
||||||
|
0.9797641013345227
|
||||||
|
],
|
||||||
|
"population_cv": 0.032815404485150704
|
||||||
|
},
|
||||||
|
"activation_output_rms_by_block": [
|
||||||
|
0.028743742033839226,
|
||||||
|
0.02942933700978756,
|
||||||
|
0.030596865341067314,
|
||||||
|
0.03249936178326607,
|
||||||
|
0.03542664647102356,
|
||||||
|
0.03894684836268425,
|
||||||
|
0.04318666458129883,
|
||||||
|
0.04936420917510986,
|
||||||
|
0.05370106175541878,
|
||||||
|
0.06048284471035004,
|
||||||
|
0.0664696991443634,
|
||||||
|
0.07231792062520981,
|
||||||
|
0.07627613097429276,
|
||||||
|
0.07997097074985504,
|
||||||
|
0.08494995534420013,
|
||||||
|
0.08972473442554474
|
||||||
|
],
|
||||||
|
"activation_output_statistics": {
|
||||||
|
"first_quartile_mean": 0.03031732654199004,
|
||||||
|
"first_to_last_ratio": 0.3664591129538783,
|
||||||
|
"imbalance_abs_log_ratio": 1.0038683247140607,
|
||||||
|
"last_quartile_mean": 0.08273044787347317,
|
||||||
|
"mean": 0.05450543703045696,
|
||||||
|
"normalized": [
|
||||||
|
0.5273555006590916,
|
||||||
|
0.5399339701348108,
|
||||||
|
0.5613543713807885,
|
||||||
|
0.5962590808162097,
|
||||||
|
0.6499653686150171,
|
||||||
|
0.7145497859400932,
|
||||||
|
0.7923368187501488,
|
||||||
|
0.9056749539963465,
|
||||||
|
0.9852422929002714,
|
||||||
|
1.1096662646068314,
|
||||||
|
1.21950584686113,
|
||||||
|
1.326801958945847,
|
||||||
|
1.399422427007981,
|
||||||
|
1.4672108895332452,
|
||||||
|
1.5585592919240543,
|
||||||
|
1.6461611779281335
|
||||||
|
],
|
||||||
|
"population_cv": 0.3808757530834043
|
||||||
|
},
|
||||||
|
"bits_per_byte": 6.7509657011324835,
|
||||||
|
"branch_output_rms_by_sublayer": [
|
||||||
|
0.003436450148001313,
|
||||||
|
0.003791899885982275,
|
||||||
|
0.003964710980653763,
|
||||||
|
0.003938332665711641,
|
||||||
|
0.0056790695525705814,
|
||||||
|
0.003919641952961683,
|
||||||
|
0.006960175931453705,
|
||||||
|
0.0038887569680809975,
|
||||||
|
0.008646421134471893,
|
||||||
|
0.003928068559616804,
|
||||||
|
0.009785857982933521,
|
||||||
|
0.0038193853106349707,
|
||||||
|
0.010625048540532589,
|
||||||
|
0.003996006678789854,
|
||||||
|
0.012892307713627815,
|
||||||
|
0.0039014811627566814,
|
||||||
|
0.012757784686982632,
|
||||||
|
0.003861474571749568,
|
||||||
|
0.014726920053362846,
|
||||||
|
0.003962590359151363,
|
||||||
|
0.012820222415030003,
|
||||||
|
0.004551138263195753,
|
||||||
|
0.01401793584227562,
|
||||||
|
0.004161422606557608,
|
||||||
|
0.01284075528383255,
|
||||||
|
0.004107494372874498,
|
||||||
|
0.012963366694748402,
|
||||||
|
0.0034911984112113714,
|
||||||
|
0.014392351731657982,
|
||||||
|
0.004168955609202385,
|
||||||
|
0.013408888131380081,
|
||||||
|
0.003909120801836252
|
||||||
|
],
|
||||||
|
"capture": {
|
||||||
|
"all_gradients_finite": true,
|
||||||
|
"all_gradients_present": true,
|
||||||
|
"count": 16,
|
||||||
|
"dtypes": [
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32"
|
||||||
|
],
|
||||||
|
"position": "post-MLP Transformer-block output; Block AttnRes is captured before aggregation-partial reset",
|
||||||
|
"shape": [
|
||||||
|
16,
|
||||||
|
256,
|
||||||
|
192
|
||||||
|
],
|
||||||
|
"storage_unique": true
|
||||||
|
},
|
||||||
|
"core_parameter_grad_rms_by_block": [
|
||||||
|
0.0004507690637370199,
|
||||||
|
0.0004530669153600251,
|
||||||
|
0.0005174622333717066,
|
||||||
|
0.0005579942111201043,
|
||||||
|
0.0006494724560479704,
|
||||||
|
0.0007477190266084469,
|
||||||
|
0.0008433132451421535,
|
||||||
|
0.0008983201732199723,
|
||||||
|
0.001008715304343127,
|
||||||
|
0.0011143118429534733,
|
||||||
|
0.0010428854825857738,
|
||||||
|
0.0010916116075341876,
|
||||||
|
0.0011085001103672855,
|
||||||
|
0.0012404769719650786,
|
||||||
|
0.0013545740271648265,
|
||||||
|
0.0014243271893498689
|
||||||
|
],
|
||||||
|
"core_parameter_grad_statistics": {
|
||||||
|
"first_quartile_mean": 0.000494823105897214,
|
||||||
|
"first_to_last_ratio": 0.385986622193018,
|
||||||
|
"imbalance_abs_log_ratio": 0.9519525676489338,
|
||||||
|
"last_quartile_mean": 0.001281969574711765,
|
||||||
|
"mean": 0.0009064699913044387,
|
||||||
|
"normalized": [
|
||||||
|
0.49727963204644987,
|
||||||
|
0.4998145771025995,
|
||||||
|
0.5708542349284638,
|
||||||
|
0.6155683215912455,
|
||||||
|
0.7164853357289404,
|
||||||
|
0.824869034585972,
|
||||||
|
0.930326710461313,
|
||||||
|
0.99100927977468,
|
||||||
|
1.112795033503044,
|
||||||
|
1.2292870736404011,
|
||||||
|
1.1504909071341995,
|
||||||
|
1.2042446170372658,
|
||||||
|
1.2228756836970622,
|
||||||
|
1.3684699812069823,
|
||||||
|
1.494339625314625,
|
||||||
|
1.571289952246756
|
||||||
|
],
|
||||||
|
"population_cv": 0.338812252534038
|
||||||
|
},
|
||||||
|
"depth_weights": [],
|
||||||
|
"layer_input_rms_by_sublayer": [
|
||||||
|
0.028241552412509918,
|
||||||
|
0.028533460572361946,
|
||||||
|
0.028743742033839226,
|
||||||
|
0.029236938804388046,
|
||||||
|
0.02942933700978756,
|
||||||
|
0.03035305254161358,
|
||||||
|
0.030596865341067314,
|
||||||
|
0.032195623964071274,
|
||||||
|
0.03249936178326607,
|
||||||
|
0.03506496921181679,
|
||||||
|
0.03542664647102356,
|
||||||
|
0.03852483630180359,
|
||||||
|
0.03894684836268425,
|
||||||
|
0.04251723736524582,
|
||||||
|
0.04318666458129883,
|
||||||
|
0.04861506074666977,
|
||||||
|
0.04936420917510986,
|
||||||
|
0.052915990352630615,
|
||||||
|
0.05370106175541878,
|
||||||
|
0.0598941408097744,
|
||||||
|
0.06048284471035004,
|
||||||
|
0.06538087129592896,
|
||||||
|
0.0664696991443634,
|
||||||
|
0.07142146676778793,
|
||||||
|
0.07231792062520981,
|
||||||
|
0.07521561533212662,
|
||||||
|
0.07627613097429276,
|
||||||
|
0.07935076206922531,
|
||||||
|
0.07997097074985504,
|
||||||
|
0.08368266373872757,
|
||||||
|
0.08494995534420013,
|
||||||
|
0.08860929310321808
|
||||||
|
],
|
||||||
|
"loss_nats": 4.679412841796875,
|
||||||
|
"loss_scale": 1.0,
|
||||||
|
"output_weights": null,
|
||||||
|
"step": 20,
|
||||||
|
"stream_state_rms_by_sublayer": [
|
||||||
|
0.028533460572361946,
|
||||||
|
0.028743742033839226,
|
||||||
|
0.029236938804388046,
|
||||||
|
0.02942933700978756,
|
||||||
|
0.03035305254161358,
|
||||||
|
0.030596865341067314,
|
||||||
|
0.032195623964071274,
|
||||||
|
0.03249936178326607,
|
||||||
|
0.03506496921181679,
|
||||||
|
0.03542664647102356,
|
||||||
|
0.03852483630180359,
|
||||||
|
0.03894684836268425,
|
||||||
|
0.04251723736524582,
|
||||||
|
0.04318666458129883,
|
||||||
|
0.04861506074666977,
|
||||||
|
0.04936420917510986,
|
||||||
|
0.052915990352630615,
|
||||||
|
0.05370106175541878,
|
||||||
|
0.0598941408097744,
|
||||||
|
0.06048284471035004,
|
||||||
|
0.06538087129592896,
|
||||||
|
0.0664696991443634,
|
||||||
|
0.07142146676778793,
|
||||||
|
0.07231792062520981,
|
||||||
|
0.07521561533212662,
|
||||||
|
0.07627613097429276,
|
||||||
|
0.07935076206922531,
|
||||||
|
0.07997097074985504,
|
||||||
|
0.08368266373872757,
|
||||||
|
0.08494995534420013,
|
||||||
|
0.08860929310321808,
|
||||||
|
0.08972473442554474
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"environment": {
|
||||||
|
"autocast": "cuda-bfloat16-forward-fp32-cross-entropy",
|
||||||
|
"compile": false,
|
||||||
|
"compute_capability": [
|
||||||
|
12,
|
||||||
|
0
|
||||||
|
],
|
||||||
|
"cublas_workspace_config": ":4096:8",
|
||||||
|
"cuda": "12.8",
|
||||||
|
"deterministic_algorithms": true,
|
||||||
|
"gpu": "NVIDIA GeForce RTX 5090",
|
||||||
|
"python": "3.10.14",
|
||||||
|
"torch": "2.11.0+cu128"
|
||||||
|
},
|
||||||
|
"evaluations": [
|
||||||
|
{
|
||||||
|
"bits_per_byte": 8.076786148009306,
|
||||||
|
"cross_entropy_nats": 5.5984015464782715,
|
||||||
|
"step": 0
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"bits_per_byte": 6.752401670275862,
|
||||||
|
"cross_entropy_nats": 4.680408179759979,
|
||||||
|
"step": 20
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"gradient_gate": {
|
||||||
|
"first_to_last_ratio_abs_delta": 0.0,
|
||||||
|
"max_abs_scale_ratio_error": 0.0,
|
||||||
|
"normalized_spectrum_max_abs_delta": 0.0,
|
||||||
|
"passed": true,
|
||||||
|
"per_block_scale_ratios": [
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0
|
||||||
|
],
|
||||||
|
"population_cv_abs_delta": 0.0,
|
||||||
|
"thresholds": {
|
||||||
|
"scale_ratio_abs": 1e-05,
|
||||||
|
"shape_abs": 1e-06
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"hashes": {
|
||||||
|
"final_mixer_parameters": null,
|
||||||
|
"final_model_state": "d53dab0fa7c76215ab91de676b2aef3f9ef14cb0cc1819b7c4a887915bed97c0",
|
||||||
|
"final_optimizer_state": "a7bce44db1478ce53933758aa5033bbb1e0aa21296e9615c33146920f30a2057",
|
||||||
|
"final_public_parameters": "d53dab0fa7c76215ab91de676b2aef3f9ef14cb0cc1819b7c4a887915bed97c0",
|
||||||
|
"initial_mixer_parameters": null,
|
||||||
|
"initial_public_parameter_elements": 9541824,
|
||||||
|
"initial_public_parameter_structure": "e732db766f25e01f6ce1772cc182ced9de2c56c4a2384130117242b9444f0abe",
|
||||||
|
"initial_public_parameter_tensors": 115,
|
||||||
|
"initial_public_parameters": "af2724a1c34bcfd61d8a8bef402246430898c815e56b6e5c5257949a4eb0e7b1"
|
||||||
|
},
|
||||||
|
"manifest": {
|
||||||
|
"diagnostic_tensor_sha256": "21117e31db302b10d67b63f035665dc8f220b879d216ccd12b7d2ba86e7b1716",
|
||||||
|
"file_sha256": "080afb17d1e036c0bba0a799fdb8b98ee4ad652bd42dd1b3b67110dd2ede6371",
|
||||||
|
"formal_schedule_sha256": "5041e09b167f229248d2462324e8c254b8f5938975f135dcd8192b00a54a4f4e",
|
||||||
|
"input_gate_tensor_hashes": {
|
||||||
|
"0": "65136111a29a042e61a7909132560d95cd4bcf0f9b52f64d0fb2e57773856434",
|
||||||
|
"1": "d995676b4e7dec8f661cd8c2345fe7fc7a513c17f528c02fc946a441a6995a94",
|
||||||
|
"7999": "2345e7ac3decca2bdaebf13094fdc92fcabef42e3e461a2fefd6e8e7e76baccc"
|
||||||
|
},
|
||||||
|
"path": "experiments/k3/attnres_gradient/manifest.json",
|
||||||
|
"validation_tensor_sha256": "f459316f13078a163b47c133511bb7181e05170ab89516e196490113893ce338"
|
||||||
|
},
|
||||||
|
"model": {
|
||||||
|
"attnres_aggregation_groups": 8,
|
||||||
|
"context": 256,
|
||||||
|
"d_ff": 768,
|
||||||
|
"d_head": 32,
|
||||||
|
"d_model": 192,
|
||||||
|
"heads": 6,
|
||||||
|
"layers": 16,
|
||||||
|
"parameters": {
|
||||||
|
"core": 9541824,
|
||||||
|
"embedding": 98304,
|
||||||
|
"mixer": 0,
|
||||||
|
"total": 9541824
|
||||||
|
},
|
||||||
|
"sublayers": 32,
|
||||||
|
"sublayers_per_attnres_group": 4,
|
||||||
|
"transformer_blocks_per_attnres_group": 2,
|
||||||
|
"vocabulary": 256
|
||||||
|
},
|
||||||
|
"optimizer": {
|
||||||
|
"betas": [
|
||||||
|
0.9,
|
||||||
|
0.95
|
||||||
|
],
|
||||||
|
"epsilon": 1e-08,
|
||||||
|
"grad_clip": 1.0,
|
||||||
|
"min_lr": 3e-05,
|
||||||
|
"name": "AdamW",
|
||||||
|
"peak_lr": 0.0003,
|
||||||
|
"warmup_steps": 400,
|
||||||
|
"weight_decay_ndim_ge_2": 0.1
|
||||||
|
},
|
||||||
|
"protocol_id": "llm-atlas-k3-attnres-gradient-scale-v1",
|
||||||
|
"run_kind": "smoke",
|
||||||
|
"schema_version": 1,
|
||||||
|
"seed": 2026073001,
|
||||||
|
"steps": 20,
|
||||||
|
"target_bytes_seen": 163840,
|
||||||
|
"timing": {
|
||||||
|
"mean_ms": null,
|
||||||
|
"measured_steps": 0,
|
||||||
|
"median_ms": null,
|
||||||
|
"p95_ms": null,
|
||||||
|
"peak_allocated_bytes": 1648265728,
|
||||||
|
"peak_reserved_bytes": 3282042880,
|
||||||
|
"warmup_steps_excluded": 20
|
||||||
|
},
|
||||||
|
"training_history": [
|
||||||
|
{
|
||||||
|
"bits_per_byte": 8.088097790921855,
|
||||||
|
"learning_rate": 7.499999999999999e-07,
|
||||||
|
"loss_nats": 5.6062421798706055,
|
||||||
|
"step": 1,
|
||||||
|
"unclipped_grad_norm": 19.475919723510742
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"bits_per_byte": 7.474302716882146,
|
||||||
|
"learning_rate": 7.499999999999999e-06,
|
||||||
|
"loss_nats": 5.180791854858398,
|
||||||
|
"step": 10,
|
||||||
|
"unclipped_grad_norm": 12.938376426696777
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"bits_per_byte": 6.789615018295581,
|
||||||
|
"learning_rate": 1.4999999999999999e-05,
|
||||||
|
"loss_nats": 4.706202507019043,
|
||||||
|
"step": 20,
|
||||||
|
"unclipped_grad_norm": 4.482712745666504
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,700 @@
|
|||||||
|
{
|
||||||
|
"architecture": "baseline",
|
||||||
|
"batch_size": 32,
|
||||||
|
"canonical_sha256_without_self": "0ae9ad1697ca15ec4c84270ad82b9fa98056f14f20fc7372c4b9886c29845415",
|
||||||
|
"depth": 16,
|
||||||
|
"diagnostics": [
|
||||||
|
{
|
||||||
|
"activation_grad_rms_by_block": [
|
||||||
|
0.0002865509013645351,
|
||||||
|
0.00024785567075014114,
|
||||||
|
0.00021551651298068464,
|
||||||
|
0.00019683866412378848,
|
||||||
|
0.0001823508064262569,
|
||||||
|
0.0001679604029050097,
|
||||||
|
0.00015584587526973337,
|
||||||
|
0.00014737028686795384,
|
||||||
|
0.00014130656199995428,
|
||||||
|
0.00013612695329356939,
|
||||||
|
0.00013129493163432926,
|
||||||
|
0.00012807638267986476,
|
||||||
|
0.00012550120300147682,
|
||||||
|
0.00012324050476308912,
|
||||||
|
0.00012108208466088399,
|
||||||
|
0.00011888353037647903
|
||||||
|
],
|
||||||
|
"activation_grad_statistics": {
|
||||||
|
"first_quartile_mean": 0.00023669043730478734,
|
||||||
|
"first_to_last_ratio": 1.9372775995887173,
|
||||||
|
"imbalance_abs_log_ratio": 0.6612836883477485,
|
||||||
|
"last_quartile_mean": 0.00012217683070048224,
|
||||||
|
"mean": 0.00016411257956860936,
|
||||||
|
"normalized": [
|
||||||
|
1.7460629899168627,
|
||||||
|
1.5102783187106137,
|
||||||
|
1.313223602646409,
|
||||||
|
1.1994124072706904,
|
||||||
|
1.1111324123086055,
|
||||||
|
1.0234462424910682,
|
||||||
|
0.9496278449793059,
|
||||||
|
0.8979828801383491,
|
||||||
|
0.8610343117596252,
|
||||||
|
0.8294729974472175,
|
||||||
|
0.8000296624393728,
|
||||||
|
0.7804178266926869,
|
||||||
|
0.7647262832098098,
|
||||||
|
0.7509509940495869,
|
||||||
|
0.7377989242455608,
|
||||||
|
0.7244023016942358
|
||||||
|
],
|
||||||
|
"population_cv": 0.29361872380036635
|
||||||
|
},
|
||||||
|
"activation_output_rms_by_block": [
|
||||||
|
0.02866268903017044,
|
||||||
|
0.029099803417921066,
|
||||||
|
0.02953805774450302,
|
||||||
|
0.030088091269135475,
|
||||||
|
0.030644793063402176,
|
||||||
|
0.03131929785013199,
|
||||||
|
0.03213750571012497,
|
||||||
|
0.03286394104361534,
|
||||||
|
0.033869802951812744,
|
||||||
|
0.03499744459986687,
|
||||||
|
0.03588006645441055,
|
||||||
|
0.037365153431892395,
|
||||||
|
0.03789033368229866,
|
||||||
|
0.039445169270038605,
|
||||||
|
0.04050002992153168,
|
||||||
|
0.04121527820825577
|
||||||
|
],
|
||||||
|
"activation_output_statistics": {
|
||||||
|
"first_quartile_mean": 0.0293471603654325,
|
||||||
|
"first_to_last_ratio": 0.7380574840395957,
|
||||||
|
"imbalance_abs_log_ratio": 0.3037335657624931,
|
||||||
|
"last_quartile_mean": 0.03976270277053118,
|
||||||
|
"mean": 0.034094841103069484,
|
||||||
|
"normalized": [
|
||||||
|
0.8406752488894867,
|
||||||
|
0.8534957922212247,
|
||||||
|
0.8663497699023965,
|
||||||
|
0.8824822259232266,
|
||||||
|
0.8988102619619861,
|
||||||
|
0.9185934539320196,
|
||||||
|
0.942591449919727,
|
||||||
|
0.9638977622528551,
|
||||||
|
0.9933996421752943,
|
||||||
|
1.0264733158329964,
|
||||||
|
1.0523605710888722,
|
||||||
|
1.0959180985456622,
|
||||||
|
1.1113216092650298,
|
||||||
|
1.1569248600043318,
|
||||||
|
1.1878638706395246,
|
||||||
|
1.2088420674453664
|
||||||
|
],
|
||||||
|
"population_cv": 0.11952987076282337
|
||||||
|
},
|
||||||
|
"bits_per_byte": 8.096463027059821,
|
||||||
|
"branch_output_rms_by_sublayer": [
|
||||||
|
0.0033055038657039404,
|
||||||
|
0.0037518907338380814,
|
||||||
|
0.0033751516602933407,
|
||||||
|
0.003911525942385197,
|
||||||
|
0.004229962360113859,
|
||||||
|
0.0038461871445178986,
|
||||||
|
0.004243654198944569,
|
||||||
|
0.0038366704247891903,
|
||||||
|
0.004372371360659599,
|
||||||
|
0.003851204412057996,
|
||||||
|
0.005078981164842844,
|
||||||
|
0.003818925702944398,
|
||||||
|
0.0054668826051056385,
|
||||||
|
0.0038749484810978174,
|
||||||
|
0.005823117680847645,
|
||||||
|
0.0037566579412668943,
|
||||||
|
0.0065947300754487514,
|
||||||
|
0.003876061411574483,
|
||||||
|
0.006763988174498081,
|
||||||
|
0.003795720636844635,
|
||||||
|
0.006863070651888847,
|
||||||
|
0.0038706199266016483,
|
||||||
|
0.007856340147554874,
|
||||||
|
0.003839300014078617,
|
||||||
|
0.00830968376249075,
|
||||||
|
0.003927029203623533,
|
||||||
|
0.009655521251261234,
|
||||||
|
0.0038513424806296825,
|
||||||
|
0.009004893712699413,
|
||||||
|
0.0039484030567109585,
|
||||||
|
0.008608299307525158,
|
||||||
|
0.003791053779423237
|
||||||
|
],
|
||||||
|
"capture": {
|
||||||
|
"all_gradients_finite": true,
|
||||||
|
"all_gradients_present": true,
|
||||||
|
"count": 16,
|
||||||
|
"dtypes": [
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32"
|
||||||
|
],
|
||||||
|
"position": "post-MLP Transformer-block output; Block AttnRes is captured before aggregation-partial reset",
|
||||||
|
"shape": [
|
||||||
|
16,
|
||||||
|
256,
|
||||||
|
192
|
||||||
|
],
|
||||||
|
"storage_unique": true
|
||||||
|
},
|
||||||
|
"core_parameter_grad_rms_by_block": [
|
||||||
|
0.00809059897248305,
|
||||||
|
0.007313812062277777,
|
||||||
|
0.007090710224412361,
|
||||||
|
0.0067033738367094624,
|
||||||
|
0.006535407867454384,
|
||||||
|
0.006785021795736036,
|
||||||
|
0.006295993569919972,
|
||||||
|
0.005509300326230647,
|
||||||
|
0.006231736048910802,
|
||||||
|
0.006047390980174007,
|
||||||
|
0.005629405359532731,
|
||||||
|
0.005784042737956458,
|
||||||
|
0.0059591736113607996,
|
||||||
|
0.0061983173180850636,
|
||||||
|
0.006571690039252182,
|
||||||
|
0.005060205437758869
|
||||||
|
],
|
||||||
|
"core_parameter_grad_statistics": {
|
||||||
|
"first_quartile_mean": 0.007299623773970663,
|
||||||
|
"first_to_last_ratio": 1.227374872012571,
|
||||||
|
"imbalance_abs_log_ratio": 0.2048776382299483,
|
||||||
|
"last_quartile_mean": 0.005947346601614228,
|
||||||
|
"mean": 0.006362886261765913,
|
||||||
|
"normalized": [
|
||||||
|
1.271529717747562,
|
||||||
|
1.1494488132258316,
|
||||||
|
1.1143858200043408,
|
||||||
|
1.05351149791715,
|
||||||
|
1.0271137340180259,
|
||||||
|
1.066343404015675,
|
||||||
|
0.9894870520870547,
|
||||||
|
0.8658492545019393,
|
||||||
|
0.9793882512652816,
|
||||||
|
0.9504163254515969,
|
||||||
|
0.8847251275509024,
|
||||||
|
0.909028151691524,
|
||||||
|
0.9365519618304368,
|
||||||
|
0.9741361173356609,
|
||||||
|
1.0328158902888074,
|
||||||
|
0.7952688810682109
|
||||||
|
],
|
||||||
|
"population_cv": 0.11384977604868386
|
||||||
|
},
|
||||||
|
"depth_weights": [],
|
||||||
|
"layer_input_rms_by_sublayer": [
|
||||||
|
0.02823694422841072,
|
||||||
|
0.028404271230101585,
|
||||||
|
0.02866268903017044,
|
||||||
|
0.02886480651795864,
|
||||||
|
0.029099803417921066,
|
||||||
|
0.029365191236138344,
|
||||||
|
0.02953805774450302,
|
||||||
|
0.029881296679377556,
|
||||||
|
0.030088091269135475,
|
||||||
|
0.030441828072071075,
|
||||||
|
0.030644793063402176,
|
||||||
|
0.03106885403394699,
|
||||||
|
0.03131929785013199,
|
||||||
|
0.031892918050289154,
|
||||||
|
0.03213750571012497,
|
||||||
|
0.03266792371869087,
|
||||||
|
0.03286394104361534,
|
||||||
|
0.03364725783467293,
|
||||||
|
0.033869802951812744,
|
||||||
|
0.03485054895281792,
|
||||||
|
0.03499744459986687,
|
||||||
|
0.03562505170702934,
|
||||||
|
0.03588006645441055,
|
||||||
|
0.03719272464513779,
|
||||||
|
0.037365153431892395,
|
||||||
|
0.03768136352300644,
|
||||||
|
0.03789033368229866,
|
||||||
|
0.03927159309387207,
|
||||||
|
0.039445169270038605,
|
||||||
|
0.04041972756385803,
|
||||||
|
0.04050002992153168,
|
||||||
|
0.04110744968056679
|
||||||
|
],
|
||||||
|
"loss_nats": 5.6120405197143555,
|
||||||
|
"loss_scale": 1.0,
|
||||||
|
"output_weights": null,
|
||||||
|
"step": 0,
|
||||||
|
"stream_state_rms_by_sublayer": [
|
||||||
|
0.028404271230101585,
|
||||||
|
0.02866268903017044,
|
||||||
|
0.02886480651795864,
|
||||||
|
0.029099803417921066,
|
||||||
|
0.029365191236138344,
|
||||||
|
0.02953805774450302,
|
||||||
|
0.029881296679377556,
|
||||||
|
0.030088091269135475,
|
||||||
|
0.030441828072071075,
|
||||||
|
0.030644793063402176,
|
||||||
|
0.03106885403394699,
|
||||||
|
0.03131929785013199,
|
||||||
|
0.031892918050289154,
|
||||||
|
0.03213750571012497,
|
||||||
|
0.03266792371869087,
|
||||||
|
0.03286394104361534,
|
||||||
|
0.03364725783467293,
|
||||||
|
0.033869802951812744,
|
||||||
|
0.03485054895281792,
|
||||||
|
0.03499744459986687,
|
||||||
|
0.03562505170702934,
|
||||||
|
0.03588006645441055,
|
||||||
|
0.03719272464513779,
|
||||||
|
0.037365153431892395,
|
||||||
|
0.03768136352300644,
|
||||||
|
0.03789033368229866,
|
||||||
|
0.03927159309387207,
|
||||||
|
0.039445169270038605,
|
||||||
|
0.04041972756385803,
|
||||||
|
0.04050002992153168,
|
||||||
|
0.04110744968056679,
|
||||||
|
0.04121527820825577
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"activation_grad_rms_by_block": [
|
||||||
|
5.69471885683015e-05,
|
||||||
|
5.527245593839325e-05,
|
||||||
|
5.393734318204224e-05,
|
||||||
|
5.292302375892177e-05,
|
||||||
|
5.23102717124857e-05,
|
||||||
|
5.178990977583453e-05,
|
||||||
|
5.1564093155320734e-05,
|
||||||
|
5.129844430484809e-05,
|
||||||
|
5.121947833686136e-05,
|
||||||
|
5.116340616950765e-05,
|
||||||
|
5.109804988023825e-05,
|
||||||
|
5.104695082991384e-05,
|
||||||
|
5.09646451973822e-05,
|
||||||
|
5.100828275317326e-05,
|
||||||
|
5.106439857627265e-05,
|
||||||
|
5.11144389747642e-05
|
||||||
|
],
|
||||||
|
"activation_grad_statistics": {
|
||||||
|
"first_quartile_mean": 5.477000286191469e-05,
|
||||||
|
"first_to_last_ratio": 1.0731232762518041,
|
||||||
|
"imbalance_abs_log_ratio": 0.07057334637995329,
|
||||||
|
"last_quartile_mean": 5.103794137539808e-05,
|
||||||
|
"mean": 5.2170148819641327e-05,
|
||||||
|
"normalized": [
|
||||||
|
1.0915665348238701,
|
||||||
|
1.0594651767139285,
|
||||||
|
1.033873669184083,
|
||||||
|
1.0144311441756324,
|
||||||
|
1.002685882559561,
|
||||||
|
0.9927115591500164,
|
||||||
|
0.988383094968431,
|
||||||
|
0.9832911246274795,
|
||||||
|
0.9817775010367221,
|
||||||
|
0.9807027069519371,
|
||||||
|
0.9794499543578176,
|
||||||
|
0.9784704852268963,
|
||||||
|
0.9768928467805085,
|
||||||
|
0.9777292936141551,
|
||||||
|
0.9788049244944386,
|
||||||
|
0.9797641013345227
|
||||||
|
],
|
||||||
|
"population_cv": 0.032815404485150704
|
||||||
|
},
|
||||||
|
"activation_output_rms_by_block": [
|
||||||
|
0.028743742033839226,
|
||||||
|
0.02942933700978756,
|
||||||
|
0.030596865341067314,
|
||||||
|
0.03249936178326607,
|
||||||
|
0.03542664647102356,
|
||||||
|
0.03894684836268425,
|
||||||
|
0.04318666458129883,
|
||||||
|
0.04936420917510986,
|
||||||
|
0.05370106175541878,
|
||||||
|
0.06048284471035004,
|
||||||
|
0.0664696991443634,
|
||||||
|
0.07231792062520981,
|
||||||
|
0.07627613097429276,
|
||||||
|
0.07997097074985504,
|
||||||
|
0.08494995534420013,
|
||||||
|
0.08972473442554474
|
||||||
|
],
|
||||||
|
"activation_output_statistics": {
|
||||||
|
"first_quartile_mean": 0.03031732654199004,
|
||||||
|
"first_to_last_ratio": 0.3664591129538783,
|
||||||
|
"imbalance_abs_log_ratio": 1.0038683247140607,
|
||||||
|
"last_quartile_mean": 0.08273044787347317,
|
||||||
|
"mean": 0.05450543703045696,
|
||||||
|
"normalized": [
|
||||||
|
0.5273555006590916,
|
||||||
|
0.5399339701348108,
|
||||||
|
0.5613543713807885,
|
||||||
|
0.5962590808162097,
|
||||||
|
0.6499653686150171,
|
||||||
|
0.7145497859400932,
|
||||||
|
0.7923368187501488,
|
||||||
|
0.9056749539963465,
|
||||||
|
0.9852422929002714,
|
||||||
|
1.1096662646068314,
|
||||||
|
1.21950584686113,
|
||||||
|
1.326801958945847,
|
||||||
|
1.399422427007981,
|
||||||
|
1.4672108895332452,
|
||||||
|
1.5585592919240543,
|
||||||
|
1.6461611779281335
|
||||||
|
],
|
||||||
|
"population_cv": 0.3808757530834043
|
||||||
|
},
|
||||||
|
"bits_per_byte": 6.7509657011324835,
|
||||||
|
"branch_output_rms_by_sublayer": [
|
||||||
|
0.003436450148001313,
|
||||||
|
0.003791899885982275,
|
||||||
|
0.003964710980653763,
|
||||||
|
0.003938332665711641,
|
||||||
|
0.0056790695525705814,
|
||||||
|
0.003919641952961683,
|
||||||
|
0.006960175931453705,
|
||||||
|
0.0038887569680809975,
|
||||||
|
0.008646421134471893,
|
||||||
|
0.003928068559616804,
|
||||||
|
0.009785857982933521,
|
||||||
|
0.0038193853106349707,
|
||||||
|
0.010625048540532589,
|
||||||
|
0.003996006678789854,
|
||||||
|
0.012892307713627815,
|
||||||
|
0.0039014811627566814,
|
||||||
|
0.012757784686982632,
|
||||||
|
0.003861474571749568,
|
||||||
|
0.014726920053362846,
|
||||||
|
0.003962590359151363,
|
||||||
|
0.012820222415030003,
|
||||||
|
0.004551138263195753,
|
||||||
|
0.01401793584227562,
|
||||||
|
0.004161422606557608,
|
||||||
|
0.01284075528383255,
|
||||||
|
0.004107494372874498,
|
||||||
|
0.012963366694748402,
|
||||||
|
0.0034911984112113714,
|
||||||
|
0.014392351731657982,
|
||||||
|
0.004168955609202385,
|
||||||
|
0.013408888131380081,
|
||||||
|
0.003909120801836252
|
||||||
|
],
|
||||||
|
"capture": {
|
||||||
|
"all_gradients_finite": true,
|
||||||
|
"all_gradients_present": true,
|
||||||
|
"count": 16,
|
||||||
|
"dtypes": [
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32",
|
||||||
|
"torch.float32"
|
||||||
|
],
|
||||||
|
"position": "post-MLP Transformer-block output; Block AttnRes is captured before aggregation-partial reset",
|
||||||
|
"shape": [
|
||||||
|
16,
|
||||||
|
256,
|
||||||
|
192
|
||||||
|
],
|
||||||
|
"storage_unique": true
|
||||||
|
},
|
||||||
|
"core_parameter_grad_rms_by_block": [
|
||||||
|
0.0004507690637370199,
|
||||||
|
0.0004530669153600251,
|
||||||
|
0.0005174622333717066,
|
||||||
|
0.0005579942111201043,
|
||||||
|
0.0006494724560479704,
|
||||||
|
0.0007477190266084469,
|
||||||
|
0.0008433132451421535,
|
||||||
|
0.0008983201732199723,
|
||||||
|
0.001008715304343127,
|
||||||
|
0.0011143118429534733,
|
||||||
|
0.0010428854825857738,
|
||||||
|
0.0010916116075341876,
|
||||||
|
0.0011085001103672855,
|
||||||
|
0.0012404769719650786,
|
||||||
|
0.0013545740271648265,
|
||||||
|
0.0014243271893498689
|
||||||
|
],
|
||||||
|
"core_parameter_grad_statistics": {
|
||||||
|
"first_quartile_mean": 0.000494823105897214,
|
||||||
|
"first_to_last_ratio": 0.385986622193018,
|
||||||
|
"imbalance_abs_log_ratio": 0.9519525676489338,
|
||||||
|
"last_quartile_mean": 0.001281969574711765,
|
||||||
|
"mean": 0.0009064699913044387,
|
||||||
|
"normalized": [
|
||||||
|
0.49727963204644987,
|
||||||
|
0.4998145771025995,
|
||||||
|
0.5708542349284638,
|
||||||
|
0.6155683215912455,
|
||||||
|
0.7164853357289404,
|
||||||
|
0.824869034585972,
|
||||||
|
0.930326710461313,
|
||||||
|
0.99100927977468,
|
||||||
|
1.112795033503044,
|
||||||
|
1.2292870736404011,
|
||||||
|
1.1504909071341995,
|
||||||
|
1.2042446170372658,
|
||||||
|
1.2228756836970622,
|
||||||
|
1.3684699812069823,
|
||||||
|
1.494339625314625,
|
||||||
|
1.571289952246756
|
||||||
|
],
|
||||||
|
"population_cv": 0.338812252534038
|
||||||
|
},
|
||||||
|
"depth_weights": [],
|
||||||
|
"layer_input_rms_by_sublayer": [
|
||||||
|
0.028241552412509918,
|
||||||
|
0.028533460572361946,
|
||||||
|
0.028743742033839226,
|
||||||
|
0.029236938804388046,
|
||||||
|
0.02942933700978756,
|
||||||
|
0.03035305254161358,
|
||||||
|
0.030596865341067314,
|
||||||
|
0.032195623964071274,
|
||||||
|
0.03249936178326607,
|
||||||
|
0.03506496921181679,
|
||||||
|
0.03542664647102356,
|
||||||
|
0.03852483630180359,
|
||||||
|
0.03894684836268425,
|
||||||
|
0.04251723736524582,
|
||||||
|
0.04318666458129883,
|
||||||
|
0.04861506074666977,
|
||||||
|
0.04936420917510986,
|
||||||
|
0.052915990352630615,
|
||||||
|
0.05370106175541878,
|
||||||
|
0.0598941408097744,
|
||||||
|
0.06048284471035004,
|
||||||
|
0.06538087129592896,
|
||||||
|
0.0664696991443634,
|
||||||
|
0.07142146676778793,
|
||||||
|
0.07231792062520981,
|
||||||
|
0.07521561533212662,
|
||||||
|
0.07627613097429276,
|
||||||
|
0.07935076206922531,
|
||||||
|
0.07997097074985504,
|
||||||
|
0.08368266373872757,
|
||||||
|
0.08494995534420013,
|
||||||
|
0.08860929310321808
|
||||||
|
],
|
||||||
|
"loss_nats": 4.679412841796875,
|
||||||
|
"loss_scale": 1.0,
|
||||||
|
"output_weights": null,
|
||||||
|
"step": 20,
|
||||||
|
"stream_state_rms_by_sublayer": [
|
||||||
|
0.028533460572361946,
|
||||||
|
0.028743742033839226,
|
||||||
|
0.029236938804388046,
|
||||||
|
0.02942933700978756,
|
||||||
|
0.03035305254161358,
|
||||||
|
0.030596865341067314,
|
||||||
|
0.032195623964071274,
|
||||||
|
0.03249936178326607,
|
||||||
|
0.03506496921181679,
|
||||||
|
0.03542664647102356,
|
||||||
|
0.03852483630180359,
|
||||||
|
0.03894684836268425,
|
||||||
|
0.04251723736524582,
|
||||||
|
0.04318666458129883,
|
||||||
|
0.04861506074666977,
|
||||||
|
0.04936420917510986,
|
||||||
|
0.052915990352630615,
|
||||||
|
0.05370106175541878,
|
||||||
|
0.0598941408097744,
|
||||||
|
0.06048284471035004,
|
||||||
|
0.06538087129592896,
|
||||||
|
0.0664696991443634,
|
||||||
|
0.07142146676778793,
|
||||||
|
0.07231792062520981,
|
||||||
|
0.07521561533212662,
|
||||||
|
0.07627613097429276,
|
||||||
|
0.07935076206922531,
|
||||||
|
0.07997097074985504,
|
||||||
|
0.08368266373872757,
|
||||||
|
0.08494995534420013,
|
||||||
|
0.08860929310321808,
|
||||||
|
0.08972473442554474
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"environment": {
|
||||||
|
"autocast": "cuda-bfloat16-forward-fp32-cross-entropy",
|
||||||
|
"compile": false,
|
||||||
|
"compute_capability": [
|
||||||
|
12,
|
||||||
|
0
|
||||||
|
],
|
||||||
|
"cublas_workspace_config": ":4096:8",
|
||||||
|
"cuda": "12.8",
|
||||||
|
"deterministic_algorithms": true,
|
||||||
|
"gpu": "NVIDIA GeForce RTX 5090",
|
||||||
|
"python": "3.10.14",
|
||||||
|
"torch": "2.11.0+cu128"
|
||||||
|
},
|
||||||
|
"evaluations": [
|
||||||
|
{
|
||||||
|
"bits_per_byte": 8.076786148009306,
|
||||||
|
"cross_entropy_nats": 5.5984015464782715,
|
||||||
|
"step": 0
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"bits_per_byte": 6.752401670275862,
|
||||||
|
"cross_entropy_nats": 4.680408179759979,
|
||||||
|
"step": 20
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"gradient_gate": {
|
||||||
|
"first_to_last_ratio_abs_delta": 0.0,
|
||||||
|
"max_abs_scale_ratio_error": 0.0,
|
||||||
|
"normalized_spectrum_max_abs_delta": 0.0,
|
||||||
|
"passed": true,
|
||||||
|
"per_block_scale_ratios": [
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0,
|
||||||
|
2.0
|
||||||
|
],
|
||||||
|
"population_cv_abs_delta": 0.0,
|
||||||
|
"thresholds": {
|
||||||
|
"scale_ratio_abs": 1e-05,
|
||||||
|
"shape_abs": 1e-06
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"hashes": {
|
||||||
|
"final_mixer_parameters": null,
|
||||||
|
"final_model_state": "d53dab0fa7c76215ab91de676b2aef3f9ef14cb0cc1819b7c4a887915bed97c0",
|
||||||
|
"final_optimizer_state": "a7bce44db1478ce53933758aa5033bbb1e0aa21296e9615c33146920f30a2057",
|
||||||
|
"final_public_parameters": "d53dab0fa7c76215ab91de676b2aef3f9ef14cb0cc1819b7c4a887915bed97c0",
|
||||||
|
"initial_mixer_parameters": null,
|
||||||
|
"initial_public_parameter_elements": 9541824,
|
||||||
|
"initial_public_parameter_structure": "e732db766f25e01f6ce1772cc182ced9de2c56c4a2384130117242b9444f0abe",
|
||||||
|
"initial_public_parameter_tensors": 115,
|
||||||
|
"initial_public_parameters": "af2724a1c34bcfd61d8a8bef402246430898c815e56b6e5c5257949a4eb0e7b1"
|
||||||
|
},
|
||||||
|
"manifest": {
|
||||||
|
"diagnostic_tensor_sha256": "21117e31db302b10d67b63f035665dc8f220b879d216ccd12b7d2ba86e7b1716",
|
||||||
|
"file_sha256": "080afb17d1e036c0bba0a799fdb8b98ee4ad652bd42dd1b3b67110dd2ede6371",
|
||||||
|
"formal_schedule_sha256": "5041e09b167f229248d2462324e8c254b8f5938975f135dcd8192b00a54a4f4e",
|
||||||
|
"input_gate_tensor_hashes": {
|
||||||
|
"0": "65136111a29a042e61a7909132560d95cd4bcf0f9b52f64d0fb2e57773856434",
|
||||||
|
"1": "d995676b4e7dec8f661cd8c2345fe7fc7a513c17f528c02fc946a441a6995a94",
|
||||||
|
"7999": "2345e7ac3decca2bdaebf13094fdc92fcabef42e3e461a2fefd6e8e7e76baccc"
|
||||||
|
},
|
||||||
|
"path": "experiments/k3/attnres_gradient/manifest.json",
|
||||||
|
"validation_tensor_sha256": "f459316f13078a163b47c133511bb7181e05170ab89516e196490113893ce338"
|
||||||
|
},
|
||||||
|
"model": {
|
||||||
|
"attnres_aggregation_groups": 8,
|
||||||
|
"context": 256,
|
||||||
|
"d_ff": 768,
|
||||||
|
"d_head": 32,
|
||||||
|
"d_model": 192,
|
||||||
|
"heads": 6,
|
||||||
|
"layers": 16,
|
||||||
|
"parameters": {
|
||||||
|
"core": 9541824,
|
||||||
|
"embedding": 98304,
|
||||||
|
"mixer": 0,
|
||||||
|
"total": 9541824
|
||||||
|
},
|
||||||
|
"sublayers": 32,
|
||||||
|
"sublayers_per_attnres_group": 4,
|
||||||
|
"transformer_blocks_per_attnres_group": 2,
|
||||||
|
"vocabulary": 256
|
||||||
|
},
|
||||||
|
"optimizer": {
|
||||||
|
"betas": [
|
||||||
|
0.9,
|
||||||
|
0.95
|
||||||
|
],
|
||||||
|
"epsilon": 1e-08,
|
||||||
|
"grad_clip": 1.0,
|
||||||
|
"min_lr": 3e-05,
|
||||||
|
"name": "AdamW",
|
||||||
|
"peak_lr": 0.0003,
|
||||||
|
"warmup_steps": 400,
|
||||||
|
"weight_decay_ndim_ge_2": 0.1
|
||||||
|
},
|
||||||
|
"protocol_id": "llm-atlas-k3-attnres-gradient-scale-v1",
|
||||||
|
"run_kind": "smoke",
|
||||||
|
"schema_version": 1,
|
||||||
|
"seed": 2026073001,
|
||||||
|
"steps": 20,
|
||||||
|
"target_bytes_seen": 163840,
|
||||||
|
"timing": {
|
||||||
|
"mean_ms": null,
|
||||||
|
"measured_steps": 0,
|
||||||
|
"median_ms": null,
|
||||||
|
"p95_ms": null,
|
||||||
|
"peak_allocated_bytes": 1648265728,
|
||||||
|
"peak_reserved_bytes": 3282042880,
|
||||||
|
"warmup_steps_excluded": 20
|
||||||
|
},
|
||||||
|
"training_history": [
|
||||||
|
{
|
||||||
|
"bits_per_byte": 8.088097790921855,
|
||||||
|
"learning_rate": 7.499999999999999e-07,
|
||||||
|
"loss_nats": 5.6062421798706055,
|
||||||
|
"step": 1,
|
||||||
|
"unclipped_grad_norm": 19.475919723510742
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"bits_per_byte": 7.474302716882146,
|
||||||
|
"learning_rate": 7.499999999999999e-06,
|
||||||
|
"loss_nats": 5.180791854858398,
|
||||||
|
"step": 10,
|
||||||
|
"unclipped_grad_norm": 12.938376426696777
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"bits_per_byte": 6.789615018295581,
|
||||||
|
"learning_rate": 1.4999999999999999e-05,
|
||||||
|
"loss_nats": 4.706202507019043,
|
||||||
|
"step": 20,
|
||||||
|
"unclipped_grad_norm": 4.482712745666504
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,827 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Run one preregistered AttnRes activation-gradient/depth experiment cell."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import hashlib
|
||||||
|
import importlib.util
|
||||||
|
import json
|
||||||
|
import math
|
||||||
|
import os
|
||||||
|
import platform
|
||||||
|
import statistics
|
||||||
|
import sys
|
||||||
|
import time
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, Iterable
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
import torch
|
||||||
|
import torch.nn as nn
|
||||||
|
import torch.nn.functional as F
|
||||||
|
|
||||||
|
|
||||||
|
PROTOCOL_ID = "llm-atlas-k3-attnres-gradient-scale-v1"
|
||||||
|
ARCHITECTURES = ("baseline", "block")
|
||||||
|
DEPTHS = (16, 32)
|
||||||
|
EXPECTED_SEEDS = (2026073001, 2026073002, 2026073003)
|
||||||
|
DIAGNOSTIC_STEPS = (0, 100, 500, 2000, 4000, 8000)
|
||||||
|
FORMAL_STEPS = 8000
|
||||||
|
SMOKE_STEPS = 20
|
||||||
|
CONTEXT = 256
|
||||||
|
VOCABULARY = 256
|
||||||
|
BLOCK_GROUPS = 8
|
||||||
|
PEAK_LR = 3e-4
|
||||||
|
MIN_LR = 3e-5
|
||||||
|
WARMUP_STEPS = 400
|
||||||
|
WEIGHT_DECAY = 0.1
|
||||||
|
BETAS = (0.9, 0.95)
|
||||||
|
ADAM_EPS = 1e-8
|
||||||
|
GRAD_CLIP = 1.0
|
||||||
|
|
||||||
|
|
||||||
|
def load_round04_module() -> Any:
|
||||||
|
path = Path(__file__).resolve().parents[1] / "attnres" / "train.py"
|
||||||
|
spec = importlib.util.spec_from_file_location("k3_attnres_round04_train", path)
|
||||||
|
if spec is None or spec.loader is None:
|
||||||
|
raise RuntimeError(f"cannot import Round 04 runner from {path}")
|
||||||
|
module = importlib.util.module_from_spec(spec)
|
||||||
|
sys.modules[spec.name] = module
|
||||||
|
spec.loader.exec_module(module)
|
||||||
|
return module
|
||||||
|
|
||||||
|
|
||||||
|
round04 = load_round04_module()
|
||||||
|
|
||||||
|
|
||||||
|
def configure_round04_globals(depth: int) -> None:
|
||||||
|
round04.PROTOCOL_ID = PROTOCOL_ID
|
||||||
|
round04.LAYERS = depth
|
||||||
|
round04.SUBLAYERS = depth * 2
|
||||||
|
round04.BLOCKS = BLOCK_GROUPS
|
||||||
|
round04.SUBLAYERS_PER_BLOCK = (depth * 2) // BLOCK_GROUPS
|
||||||
|
round04.WARMUP_STEPS = WARMUP_STEPS
|
||||||
|
round04.EVAL_STEPS = DIAGNOSTIC_STEPS
|
||||||
|
|
||||||
|
|
||||||
|
def parse_args() -> argparse.Namespace:
|
||||||
|
parser = argparse.ArgumentParser()
|
||||||
|
parser.add_argument("--architecture", choices=ARCHITECTURES, required=True)
|
||||||
|
parser.add_argument("--depth", type=int, choices=DEPTHS, required=True)
|
||||||
|
parser.add_argument("--seed", type=int, required=True)
|
||||||
|
parser.add_argument("--steps", type=int)
|
||||||
|
parser.add_argument("--batch-size", type=int, default=32)
|
||||||
|
parser.add_argument("--cache-dir", type=Path, required=True)
|
||||||
|
parser.add_argument("--manifest", type=Path, required=True)
|
||||||
|
parser.add_argument("--output", type=Path, required=True)
|
||||||
|
parser.add_argument("--validation-windows", type=int, default=64)
|
||||||
|
parser.add_argument("--diagnostic-windows", type=int, default=16)
|
||||||
|
parser.add_argument("--eval-batch-size", type=int, default=8)
|
||||||
|
parser.add_argument("--timing-warmup", type=int, default=20)
|
||||||
|
parser.add_argument(
|
||||||
|
"--run-kind", choices=("smoke", "formal", "replay"), default="formal"
|
||||||
|
)
|
||||||
|
args = parser.parse_args()
|
||||||
|
expected_steps = SMOKE_STEPS if args.run_kind == "smoke" else FORMAL_STEPS
|
||||||
|
if args.steps is None:
|
||||||
|
args.steps = expected_steps
|
||||||
|
if args.steps != expected_steps:
|
||||||
|
raise ValueError(
|
||||||
|
f"{args.run_kind} must run exactly {expected_steps} steps, got {args.steps}"
|
||||||
|
)
|
||||||
|
if args.batch_size != 32:
|
||||||
|
raise ValueError("the frozen protocol requires batch size 32")
|
||||||
|
if args.validation_windows != 64 or args.diagnostic_windows != 16:
|
||||||
|
raise ValueError("the frozen protocol requires 64 validation / 16 diagnostic windows")
|
||||||
|
return args
|
||||||
|
|
||||||
|
|
||||||
|
def configure_determinism(seed: int) -> None:
|
||||||
|
if os.environ.get("CUBLAS_WORKSPACE_CONFIG") != ":4096:8":
|
||||||
|
raise RuntimeError("CUBLAS_WORKSPACE_CONFIG must be :4096:8 before Python starts")
|
||||||
|
torch.manual_seed(seed)
|
||||||
|
torch.cuda.manual_seed_all(seed)
|
||||||
|
torch.use_deterministic_algorithms(True)
|
||||||
|
torch.backends.cudnn.benchmark = False
|
||||||
|
torch.backends.cudnn.deterministic = True
|
||||||
|
torch.backends.cuda.matmul.allow_tf32 = False
|
||||||
|
torch.backends.cudnn.allow_tf32 = False
|
||||||
|
torch.set_float32_matmul_precision("highest")
|
||||||
|
|
||||||
|
|
||||||
|
def canonical_sha256(value: Any) -> str:
|
||||||
|
payload = json.dumps(
|
||||||
|
value, ensure_ascii=False, sort_keys=True, separators=(",", ":")
|
||||||
|
).encode()
|
||||||
|
return hashlib.sha256(payload).hexdigest()
|
||||||
|
|
||||||
|
|
||||||
|
def file_sha256(path: Path) -> str:
|
||||||
|
digest = hashlib.sha256()
|
||||||
|
with path.open("rb") as handle:
|
||||||
|
for block in iter(lambda: handle.read(1024 * 1024), b""):
|
||||||
|
digest.update(block)
|
||||||
|
return digest.hexdigest()
|
||||||
|
|
||||||
|
|
||||||
|
def tensor_bytes(tensor: torch.Tensor) -> bytes:
|
||||||
|
value = tensor.detach().cpu().contiguous()
|
||||||
|
return (
|
||||||
|
f"{value.dtype}|{tuple(value.shape)}|".encode()
|
||||||
|
+ value.reshape(-1).view(torch.uint8).numpy().tobytes()
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def named_state_hash(
|
||||||
|
model: nn.Module, *, include_mixers: bool | None
|
||||||
|
) -> str:
|
||||||
|
digest = hashlib.sha256()
|
||||||
|
for name, tensor in sorted(model.state_dict().items()):
|
||||||
|
is_mixer = name.startswith("mixers.") or name.startswith("output_mixer.")
|
||||||
|
if include_mixers is not None and is_mixer != include_mixers:
|
||||||
|
continue
|
||||||
|
digest.update(name.encode())
|
||||||
|
digest.update(b"\0")
|
||||||
|
digest.update(tensor_bytes(tensor))
|
||||||
|
return digest.hexdigest()
|
||||||
|
|
||||||
|
|
||||||
|
def state_structure_hash(
|
||||||
|
model: nn.Module, *, include_mixers: bool | None
|
||||||
|
) -> tuple[str, int, int]:
|
||||||
|
digest = hashlib.sha256()
|
||||||
|
tensor_count = 0
|
||||||
|
element_count = 0
|
||||||
|
for name, tensor in sorted(model.state_dict().items()):
|
||||||
|
is_mixer = name.startswith("mixers.") or name.startswith("output_mixer.")
|
||||||
|
if include_mixers is not None and is_mixer != include_mixers:
|
||||||
|
continue
|
||||||
|
digest.update(
|
||||||
|
f"{name}|{tuple(tensor.shape)}|{tensor.dtype}|{tensor.numel()}\n".encode()
|
||||||
|
)
|
||||||
|
tensor_count += 1
|
||||||
|
element_count += tensor.numel()
|
||||||
|
return digest.hexdigest(), tensor_count, element_count
|
||||||
|
|
||||||
|
|
||||||
|
def recursive_state_hash(value: Any) -> str:
|
||||||
|
digest = hashlib.sha256()
|
||||||
|
|
||||||
|
def visit(path: str, item: Any) -> None:
|
||||||
|
if torch.is_tensor(item):
|
||||||
|
digest.update(f"{path}|tensor|".encode())
|
||||||
|
digest.update(tensor_bytes(item))
|
||||||
|
elif isinstance(item, dict):
|
||||||
|
digest.update(f"{path}|dict|{len(item)}\n".encode())
|
||||||
|
for key in sorted(item, key=lambda candidate: str(candidate)):
|
||||||
|
visit(f"{path}/{key}", item[key])
|
||||||
|
elif isinstance(item, (list, tuple)):
|
||||||
|
digest.update(f"{path}|sequence|{len(item)}\n".encode())
|
||||||
|
for index, child in enumerate(item):
|
||||||
|
visit(f"{path}/{index}", child)
|
||||||
|
else:
|
||||||
|
digest.update(f"{path}|scalar|{repr(item)}\n".encode())
|
||||||
|
|
||||||
|
visit("root", value)
|
||||||
|
return digest.hexdigest()
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class ActivationTrace:
|
||||||
|
block_outputs: list[torch.Tensor]
|
||||||
|
layer_input_rms: list[float]
|
||||||
|
branch_output_rms: list[float]
|
||||||
|
stream_state_rms: list[float]
|
||||||
|
depth_weights: list[dict[str, Any]]
|
||||||
|
output_weights: dict[str, Any] | None = None
|
||||||
|
|
||||||
|
|
||||||
|
def rms(value: torch.Tensor) -> float:
|
||||||
|
return value.float().square().mean().sqrt().detach().cpu().item()
|
||||||
|
|
||||||
|
|
||||||
|
class GradientLanguageModel(round04.ReducedLanguageModel):
|
||||||
|
"""Round 04 trunk with aligned post-MLP activation capture."""
|
||||||
|
|
||||||
|
def forward(
|
||||||
|
self, input_ids: torch.Tensor, capture: bool = False
|
||||||
|
) -> tuple[torch.Tensor, ActivationTrace | None]:
|
||||||
|
embedded = self.embed(input_ids)
|
||||||
|
trace = ActivationTrace([], [], [], [], []) if capture else None
|
||||||
|
|
||||||
|
if self.architecture == "baseline":
|
||||||
|
hidden = embedded
|
||||||
|
for block in self.blocks:
|
||||||
|
attention_input = hidden
|
||||||
|
attention_output = block.attention(block.attention_norm(attention_input))
|
||||||
|
hidden = hidden + attention_output
|
||||||
|
if trace is not None:
|
||||||
|
trace.layer_input_rms.append(rms(attention_input))
|
||||||
|
trace.branch_output_rms.append(rms(attention_output))
|
||||||
|
trace.stream_state_rms.append(rms(hidden))
|
||||||
|
mlp_input = hidden
|
||||||
|
mlp_output = block.mlp(block.mlp_norm(mlp_input))
|
||||||
|
hidden = hidden + mlp_output
|
||||||
|
if trace is not None:
|
||||||
|
hidden.retain_grad()
|
||||||
|
trace.block_outputs.append(hidden)
|
||||||
|
trace.layer_input_rms.append(rms(mlp_input))
|
||||||
|
trace.branch_output_rms.append(rms(mlp_output))
|
||||||
|
trace.stream_state_rms.append(rms(hidden))
|
||||||
|
else:
|
||||||
|
completed = [embedded]
|
||||||
|
partial: torch.Tensor | None = None
|
||||||
|
mixer_index = 0
|
||||||
|
for block in self.blocks:
|
||||||
|
for branch_index in range(2):
|
||||||
|
sources = completed + ([] if partial is None else [partial])
|
||||||
|
branch_input, weights = self.mixers[mixer_index](
|
||||||
|
sources, capture
|
||||||
|
)
|
||||||
|
mixer_index += 1
|
||||||
|
if branch_index == 0:
|
||||||
|
branch_output = block.attention(
|
||||||
|
block.attention_norm(branch_input)
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
branch_output = block.mlp(block.mlp_norm(branch_input))
|
||||||
|
branch_for_residual = branch_output.float()
|
||||||
|
partial = (
|
||||||
|
branch_for_residual
|
||||||
|
if partial is None
|
||||||
|
else partial + branch_for_residual
|
||||||
|
)
|
||||||
|
if trace is not None:
|
||||||
|
trace.layer_input_rms.append(rms(branch_input))
|
||||||
|
trace.branch_output_rms.append(rms(branch_output))
|
||||||
|
trace.stream_state_rms.append(rms(partial))
|
||||||
|
trace.depth_weights.append(weights or {})
|
||||||
|
if branch_index == 1:
|
||||||
|
partial.retain_grad()
|
||||||
|
trace.block_outputs.append(partial)
|
||||||
|
if mixer_index % round04.SUBLAYERS_PER_BLOCK == 0:
|
||||||
|
completed.append(partial)
|
||||||
|
partial = None
|
||||||
|
if partial is not None or len(completed) != BLOCK_GROUPS + 1:
|
||||||
|
raise RuntimeError("Block AttnRes aggregation contract failed")
|
||||||
|
if self.output_mixer is None:
|
||||||
|
raise RuntimeError("Block AttnRes output mixer missing")
|
||||||
|
hidden, output_weights = self.output_mixer(completed, capture)
|
||||||
|
if trace is not None:
|
||||||
|
trace.output_weights = output_weights
|
||||||
|
|
||||||
|
normalized = self.final_norm(hidden)
|
||||||
|
logits = F.linear(normalized, self.token_embedding.weight)
|
||||||
|
return logits, trace
|
||||||
|
|
||||||
|
|
||||||
|
def cross_entropy(logits: torch.Tensor, targets: torch.Tensor) -> torch.Tensor:
|
||||||
|
return F.cross_entropy(
|
||||||
|
logits.float().reshape(-1, VOCABULARY), targets.reshape(-1)
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def learning_rate(step: int, total_steps: int) -> float:
|
||||||
|
if step <= WARMUP_STEPS:
|
||||||
|
return PEAK_LR * step / WARMUP_STEPS
|
||||||
|
progress = (step - WARMUP_STEPS) / max(1, total_steps - WARMUP_STEPS)
|
||||||
|
cosine = 0.5 * (1 + math.cos(math.pi * progress))
|
||||||
|
return MIN_LR + (PEAK_LR - MIN_LR) * cosine
|
||||||
|
|
||||||
|
|
||||||
|
@torch.no_grad()
|
||||||
|
def evaluate(
|
||||||
|
model: GradientLanguageModel,
|
||||||
|
corpus: Any,
|
||||||
|
window_count: int,
|
||||||
|
eval_batch_size: int,
|
||||||
|
) -> dict[str, float]:
|
||||||
|
model.eval()
|
||||||
|
loss_sum = 0.0
|
||||||
|
target_count = 0
|
||||||
|
for begin in range(0, window_count, eval_batch_size):
|
||||||
|
end = min(begin + eval_batch_size, window_count)
|
||||||
|
inputs, targets = corpus.fixed_batch(
|
||||||
|
corpus.validation_starts, begin, end
|
||||||
|
)
|
||||||
|
with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
|
||||||
|
logits, _ = model(inputs)
|
||||||
|
loss = F.cross_entropy(
|
||||||
|
logits.float().reshape(-1, VOCABULARY),
|
||||||
|
targets.reshape(-1),
|
||||||
|
reduction="sum",
|
||||||
|
)
|
||||||
|
loss_sum += loss.detach().cpu().item()
|
||||||
|
target_count += targets.numel()
|
||||||
|
nats = loss_sum / target_count
|
||||||
|
return {"cross_entropy_nats": nats, "bits_per_byte": nats / math.log(2)}
|
||||||
|
|
||||||
|
|
||||||
|
def mean(values: Iterable[float]) -> float:
|
||||||
|
return statistics.fmean(values)
|
||||||
|
|
||||||
|
|
||||||
|
def depth_statistics(values: list[float]) -> dict[str, Any]:
|
||||||
|
average = mean(values)
|
||||||
|
variance = mean((value - average) ** 2 for value in values)
|
||||||
|
quartile = len(values) // 4
|
||||||
|
first = mean(values[:quartile])
|
||||||
|
last = mean(values[-quartile:])
|
||||||
|
ratio = first / last
|
||||||
|
return {
|
||||||
|
"mean": average,
|
||||||
|
"population_cv": math.sqrt(variance) / average,
|
||||||
|
"normalized": [value / average for value in values],
|
||||||
|
"first_quartile_mean": first,
|
||||||
|
"last_quartile_mean": last,
|
||||||
|
"first_to_last_ratio": ratio,
|
||||||
|
"imbalance_abs_log_ratio": abs(math.log(ratio)),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def core_parameter_gradient_rms(model: GradientLanguageModel) -> list[float]:
|
||||||
|
values = []
|
||||||
|
for block in model.blocks:
|
||||||
|
sum_square = 0.0
|
||||||
|
count = 0
|
||||||
|
for parameter in block.parameters():
|
||||||
|
if parameter.grad is None:
|
||||||
|
raise RuntimeError("missing core parameter gradient")
|
||||||
|
gradient = parameter.grad.detach().float()
|
||||||
|
if not torch.isfinite(gradient).all():
|
||||||
|
raise RuntimeError("non-finite core parameter gradient")
|
||||||
|
sum_square += gradient.square().sum().detach().cpu().item()
|
||||||
|
count += gradient.numel()
|
||||||
|
values.append(math.sqrt(sum_square / count))
|
||||||
|
return values
|
||||||
|
|
||||||
|
|
||||||
|
def activation_storage_unique(outputs: list[torch.Tensor]) -> bool:
|
||||||
|
pointers = [output.untyped_storage().data_ptr() for output in outputs]
|
||||||
|
return len(pointers) == len(set(pointers))
|
||||||
|
|
||||||
|
|
||||||
|
def diagnostic(
|
||||||
|
model: GradientLanguageModel,
|
||||||
|
corpus: Any,
|
||||||
|
window_count: int,
|
||||||
|
*,
|
||||||
|
loss_scale: float = 1.0,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
model.eval()
|
||||||
|
model.zero_grad(set_to_none=True)
|
||||||
|
inputs, targets = corpus.fixed_batch(
|
||||||
|
corpus.diagnostic_starts, 0, window_count
|
||||||
|
)
|
||||||
|
with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
|
||||||
|
logits, trace = model(inputs, capture=True)
|
||||||
|
unscaled_loss = cross_entropy(logits, targets)
|
||||||
|
loss = unscaled_loss * loss_scale
|
||||||
|
if trace is None or len(trace.block_outputs) != len(model.blocks):
|
||||||
|
raise RuntimeError("aligned activation capture count mismatch")
|
||||||
|
expected_shape = (window_count, CONTEXT, round04.D_MODEL)
|
||||||
|
if any(tuple(output.shape) != expected_shape for output in trace.block_outputs):
|
||||||
|
raise RuntimeError("aligned activation capture shape mismatch")
|
||||||
|
if any(output.dtype != torch.float32 for output in trace.block_outputs):
|
||||||
|
raise RuntimeError("aligned activation capture must use FP32 residual state")
|
||||||
|
if not activation_storage_unique(trace.block_outputs):
|
||||||
|
raise RuntimeError("captured block outputs alias storage")
|
||||||
|
loss.backward()
|
||||||
|
|
||||||
|
activation_grad_rms = []
|
||||||
|
activation_output_rms = []
|
||||||
|
activation_dtypes = []
|
||||||
|
for output in trace.block_outputs:
|
||||||
|
if output.grad is None:
|
||||||
|
raise RuntimeError("captured activation gradient is None")
|
||||||
|
gradient = output.grad.detach().float()
|
||||||
|
if not torch.isfinite(gradient).all():
|
||||||
|
raise RuntimeError("captured activation gradient is non-finite")
|
||||||
|
activation_grad_rms.append(
|
||||||
|
gradient.square().mean().sqrt().detach().cpu().item()
|
||||||
|
)
|
||||||
|
activation_output_rms.append(rms(output))
|
||||||
|
activation_dtypes.append(str(output.dtype))
|
||||||
|
parameter_grad_rms = core_parameter_gradient_rms(model)
|
||||||
|
|
||||||
|
return {
|
||||||
|
"loss_nats": unscaled_loss.detach().cpu().item(),
|
||||||
|
"bits_per_byte": unscaled_loss.detach().cpu().item() / math.log(2),
|
||||||
|
"loss_scale": loss_scale,
|
||||||
|
"capture": {
|
||||||
|
"count": len(trace.block_outputs),
|
||||||
|
"shape": list(expected_shape),
|
||||||
|
"dtypes": activation_dtypes,
|
||||||
|
"all_gradients_finite": True,
|
||||||
|
"all_gradients_present": True,
|
||||||
|
"storage_unique": True,
|
||||||
|
"position": (
|
||||||
|
"post-MLP Transformer-block output; Block AttnRes is captured "
|
||||||
|
"before aggregation-partial reset"
|
||||||
|
),
|
||||||
|
},
|
||||||
|
"activation_grad_rms_by_block": activation_grad_rms,
|
||||||
|
"activation_grad_statistics": depth_statistics(activation_grad_rms),
|
||||||
|
"activation_output_rms_by_block": activation_output_rms,
|
||||||
|
"activation_output_statistics": depth_statistics(activation_output_rms),
|
||||||
|
"core_parameter_grad_rms_by_block": parameter_grad_rms,
|
||||||
|
"core_parameter_grad_statistics": depth_statistics(parameter_grad_rms),
|
||||||
|
"layer_input_rms_by_sublayer": trace.layer_input_rms,
|
||||||
|
"branch_output_rms_by_sublayer": trace.branch_output_rms,
|
||||||
|
"stream_state_rms_by_sublayer": trace.stream_state_rms,
|
||||||
|
"depth_weights": trace.depth_weights,
|
||||||
|
"output_weights": trace.output_weights,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def loss_scale_gate(
|
||||||
|
model: GradientLanguageModel, corpus: Any, window_count: int
|
||||||
|
) -> tuple[dict[str, Any], dict[str, Any]]:
|
||||||
|
base = diagnostic(model, corpus, window_count, loss_scale=1.0)
|
||||||
|
doubled = diagnostic(model, corpus, window_count, loss_scale=2.0)
|
||||||
|
base_values = base["activation_grad_rms_by_block"]
|
||||||
|
doubled_values = doubled["activation_grad_rms_by_block"]
|
||||||
|
ratios = [
|
||||||
|
doubled_value / base_value
|
||||||
|
for base_value, doubled_value in zip(base_values, doubled_values)
|
||||||
|
]
|
||||||
|
base_stats = base["activation_grad_statistics"]
|
||||||
|
doubled_stats = doubled["activation_grad_statistics"]
|
||||||
|
cv_delta = abs(
|
||||||
|
doubled_stats["population_cv"] - base_stats["population_cv"]
|
||||||
|
)
|
||||||
|
ratio_delta = abs(
|
||||||
|
doubled_stats["first_to_last_ratio"]
|
||||||
|
- base_stats["first_to_last_ratio"]
|
||||||
|
)
|
||||||
|
normalized_max_delta = max(
|
||||||
|
abs(left - right)
|
||||||
|
for left, right in zip(
|
||||||
|
base_stats["normalized"], doubled_stats["normalized"]
|
||||||
|
)
|
||||||
|
)
|
||||||
|
passed = (
|
||||||
|
all(abs(ratio - 2.0) <= 1e-5 for ratio in ratios)
|
||||||
|
and cv_delta <= 1e-6
|
||||||
|
and ratio_delta <= 1e-6
|
||||||
|
and normalized_max_delta <= 1e-6
|
||||||
|
)
|
||||||
|
gate = {
|
||||||
|
"passed": passed,
|
||||||
|
"per_block_scale_ratios": ratios,
|
||||||
|
"max_abs_scale_ratio_error": max(abs(ratio - 2.0) for ratio in ratios),
|
||||||
|
"population_cv_abs_delta": cv_delta,
|
||||||
|
"first_to_last_ratio_abs_delta": ratio_delta,
|
||||||
|
"normalized_spectrum_max_abs_delta": normalized_max_delta,
|
||||||
|
"thresholds": {
|
||||||
|
"scale_ratio_abs": 1e-5,
|
||||||
|
"shape_abs": 1e-6,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
if not passed:
|
||||||
|
raise RuntimeError(f"loss-scale diagnostic gate failed: {gate}")
|
||||||
|
return base, gate
|
||||||
|
|
||||||
|
|
||||||
|
def percentile(values: list[float], quantile: float) -> float:
|
||||||
|
return float(np.quantile(np.asarray(values, dtype=np.float64), quantile))
|
||||||
|
|
||||||
|
|
||||||
|
def parameter_inventory(model: GradientLanguageModel) -> dict[str, int]:
|
||||||
|
total = sum(parameter.numel() for parameter in model.parameters())
|
||||||
|
mixer = sum(
|
||||||
|
parameter.numel()
|
||||||
|
for name, parameter in model.named_parameters()
|
||||||
|
if name.startswith("mixers.") or name.startswith("output_mixer.")
|
||||||
|
)
|
||||||
|
return {
|
||||||
|
"total": total,
|
||||||
|
"core": total - mixer,
|
||||||
|
"mixer": mixer,
|
||||||
|
"embedding": (
|
||||||
|
model.token_embedding.weight.numel()
|
||||||
|
+ model.position_embedding.weight.numel()
|
||||||
|
),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def model_input_gate_hashes(
|
||||||
|
corpus: Any, manifest: dict[str, Any], seed: int, batch_size: int
|
||||||
|
) -> dict[str, str]:
|
||||||
|
values: dict[str, str] = {}
|
||||||
|
for step in manifest["windows"]["gate_steps"]:
|
||||||
|
raw_digest = hashlib.sha256()
|
||||||
|
for row in range(batch_size):
|
||||||
|
start = round04.window_start(seed, step, row, len(corpus.train))
|
||||||
|
raw_digest.update(
|
||||||
|
np.asarray(
|
||||||
|
corpus.train[start : start + CONTEXT + 1], dtype=np.uint8
|
||||||
|
).tobytes()
|
||||||
|
)
|
||||||
|
expected_raw_hash = manifest["windows"][
|
||||||
|
"gate_training_tensor_sha256"
|
||||||
|
][str(seed)][str(step)]
|
||||||
|
if raw_digest.hexdigest() != expected_raw_hash:
|
||||||
|
raise RuntimeError(f"manifest gate tensor mismatch at step {step}")
|
||||||
|
inputs, targets = corpus.training_batch(seed, step, batch_size)
|
||||||
|
digest = hashlib.sha256()
|
||||||
|
digest.update(tensor_bytes(inputs))
|
||||||
|
digest.update(tensor_bytes(targets))
|
||||||
|
values[str(step)] = digest.hexdigest()
|
||||||
|
return values
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
args = parse_args()
|
||||||
|
if not torch.cuda.is_available():
|
||||||
|
raise RuntimeError("CUDA is required by the frozen protocol")
|
||||||
|
if args.seed not in EXPECTED_SEEDS:
|
||||||
|
raise ValueError(f"seed is not preregistered: {args.seed}")
|
||||||
|
configure_round04_globals(args.depth)
|
||||||
|
configure_determinism(args.seed)
|
||||||
|
device = torch.device("cuda")
|
||||||
|
|
||||||
|
manifest = json.loads(args.manifest.read_text())
|
||||||
|
if manifest["protocol_id"] != PROTOCOL_ID:
|
||||||
|
raise ValueError("manifest protocol mismatch")
|
||||||
|
if manifest["windows"]["formal_steps"] != FORMAL_STEPS:
|
||||||
|
raise ValueError("manifest formal-step mismatch")
|
||||||
|
corpus = round04.ByteCorpus(args.cache_dir, manifest, device)
|
||||||
|
|
||||||
|
model = GradientLanguageModel(args.architecture).to(device)
|
||||||
|
public_structure_hash, public_tensors, public_elements = state_structure_hash(
|
||||||
|
model, include_mixers=False
|
||||||
|
)
|
||||||
|
initial_public_hash = named_state_hash(model, include_mixers=False)
|
||||||
|
initial_mixer_hash = (
|
||||||
|
named_state_hash(model, include_mixers=True)
|
||||||
|
if args.architecture == "block"
|
||||||
|
else None
|
||||||
|
)
|
||||||
|
input_gate_hashes = model_input_gate_hashes(
|
||||||
|
corpus, manifest, args.seed, args.batch_size
|
||||||
|
)
|
||||||
|
|
||||||
|
decay_parameters: list[nn.Parameter] = []
|
||||||
|
no_decay_parameters: list[nn.Parameter] = []
|
||||||
|
for parameter in model.parameters():
|
||||||
|
(decay_parameters if parameter.ndim >= 2 else no_decay_parameters).append(
|
||||||
|
parameter
|
||||||
|
)
|
||||||
|
optimizer = torch.optim.AdamW(
|
||||||
|
[
|
||||||
|
{"params": decay_parameters, "weight_decay": WEIGHT_DECAY},
|
||||||
|
{"params": no_decay_parameters, "weight_decay": 0.0},
|
||||||
|
],
|
||||||
|
lr=PEAK_LR,
|
||||||
|
betas=BETAS,
|
||||||
|
eps=ADAM_EPS,
|
||||||
|
)
|
||||||
|
|
||||||
|
evaluation_steps = sorted(
|
||||||
|
set(step for step in DIAGNOSTIC_STEPS if step <= args.steps)
|
||||||
|
| {0, args.steps}
|
||||||
|
)
|
||||||
|
evaluations = [
|
||||||
|
{
|
||||||
|
"step": 0,
|
||||||
|
**evaluate(
|
||||||
|
model, corpus, args.validation_windows, args.eval_batch_size
|
||||||
|
),
|
||||||
|
}
|
||||||
|
]
|
||||||
|
if args.run_kind == "smoke":
|
||||||
|
initial_diagnostic, gradient_gate = loss_scale_gate(
|
||||||
|
model, corpus, args.diagnostic_windows
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
initial_diagnostic = diagnostic(
|
||||||
|
model, corpus, args.diagnostic_windows
|
||||||
|
)
|
||||||
|
gradient_gate = None
|
||||||
|
diagnostics = [{"step": 0, **initial_diagnostic}]
|
||||||
|
print(
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"event": "diagnostic",
|
||||||
|
"step": 0,
|
||||||
|
"architecture": args.architecture,
|
||||||
|
"depth": args.depth,
|
||||||
|
"validation_bpc": evaluations[0]["bits_per_byte"],
|
||||||
|
"activation_gradient_cv": initial_diagnostic[
|
||||||
|
"activation_grad_statistics"
|
||||||
|
]["population_cv"],
|
||||||
|
},
|
||||||
|
sort_keys=True,
|
||||||
|
),
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
model.zero_grad(set_to_none=True)
|
||||||
|
|
||||||
|
training_history: list[dict[str, float | int]] = []
|
||||||
|
step_times: list[float] = []
|
||||||
|
model.train()
|
||||||
|
for step in range(1, args.steps + 1):
|
||||||
|
lr = learning_rate(step, args.steps)
|
||||||
|
for group in optimizer.param_groups:
|
||||||
|
group["lr"] = lr
|
||||||
|
inputs, targets = corpus.training_batch(args.seed, step, args.batch_size)
|
||||||
|
optimizer.zero_grad(set_to_none=True)
|
||||||
|
|
||||||
|
torch.cuda.synchronize()
|
||||||
|
started = time.perf_counter()
|
||||||
|
with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
|
||||||
|
logits, _ = model(inputs)
|
||||||
|
loss = cross_entropy(logits, targets)
|
||||||
|
if not torch.isfinite(loss):
|
||||||
|
raise RuntimeError(f"non-finite loss at step {step}: {loss}")
|
||||||
|
loss.backward()
|
||||||
|
unclipped_norm = torch.nn.utils.clip_grad_norm_(
|
||||||
|
model.parameters(), GRAD_CLIP
|
||||||
|
)
|
||||||
|
optimizer.step()
|
||||||
|
torch.cuda.synchronize()
|
||||||
|
elapsed_ms = (time.perf_counter() - started) * 1000
|
||||||
|
|
||||||
|
if step == args.timing_warmup:
|
||||||
|
torch.cuda.reset_peak_memory_stats()
|
||||||
|
elif step > args.timing_warmup:
|
||||||
|
step_times.append(elapsed_ms)
|
||||||
|
if step == 1 or step % 10 == 0 or step == args.steps:
|
||||||
|
training_history.append(
|
||||||
|
{
|
||||||
|
"step": step,
|
||||||
|
"loss_nats": loss.detach().cpu().item(),
|
||||||
|
"bits_per_byte": loss.detach().cpu().item() / math.log(2),
|
||||||
|
"learning_rate": lr,
|
||||||
|
"unclipped_grad_norm": float(unclipped_norm.detach().cpu()),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
if step in evaluation_steps and step != 0:
|
||||||
|
evaluations.append(
|
||||||
|
{
|
||||||
|
"step": step,
|
||||||
|
**evaluate(
|
||||||
|
model,
|
||||||
|
corpus,
|
||||||
|
args.validation_windows,
|
||||||
|
args.eval_batch_size,
|
||||||
|
),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
diagnostics.append(
|
||||||
|
{
|
||||||
|
"step": step,
|
||||||
|
**diagnostic(
|
||||||
|
model, corpus, args.diagnostic_windows
|
||||||
|
),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
print(
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"event": "diagnostic",
|
||||||
|
"step": step,
|
||||||
|
"architecture": args.architecture,
|
||||||
|
"depth": args.depth,
|
||||||
|
"validation_bpc": evaluations[-1]["bits_per_byte"],
|
||||||
|
"activation_gradient_cv": diagnostics[-1][
|
||||||
|
"activation_grad_statistics"
|
||||||
|
]["population_cv"],
|
||||||
|
},
|
||||||
|
sort_keys=True,
|
||||||
|
),
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
model.zero_grad(set_to_none=True)
|
||||||
|
model.train()
|
||||||
|
|
||||||
|
training_peak_allocated = torch.cuda.max_memory_allocated()
|
||||||
|
training_peak_reserved = torch.cuda.max_memory_reserved()
|
||||||
|
final_public_hash = named_state_hash(model, include_mixers=False)
|
||||||
|
final_mixer_hash = (
|
||||||
|
named_state_hash(model, include_mixers=True)
|
||||||
|
if args.architecture == "block"
|
||||||
|
else None
|
||||||
|
)
|
||||||
|
final_full_hash = named_state_hash(model, include_mixers=None)
|
||||||
|
optimizer_hash = recursive_state_hash(optimizer.state_dict())
|
||||||
|
timing = {
|
||||||
|
"warmup_steps_excluded": args.timing_warmup,
|
||||||
|
"measured_steps": len(step_times),
|
||||||
|
"mean_ms": mean(step_times) if step_times else None,
|
||||||
|
"median_ms": statistics.median(step_times) if step_times else None,
|
||||||
|
"p95_ms": percentile(step_times, 0.95) if step_times else None,
|
||||||
|
"peak_allocated_bytes": training_peak_allocated,
|
||||||
|
"peak_reserved_bytes": training_peak_reserved,
|
||||||
|
}
|
||||||
|
result = {
|
||||||
|
"schema_version": 1,
|
||||||
|
"protocol_id": PROTOCOL_ID,
|
||||||
|
"run_kind": args.run_kind,
|
||||||
|
"architecture": args.architecture,
|
||||||
|
"depth": args.depth,
|
||||||
|
"seed": args.seed,
|
||||||
|
"steps": args.steps,
|
||||||
|
"batch_size": args.batch_size,
|
||||||
|
"target_bytes_seen": args.steps * args.batch_size * CONTEXT,
|
||||||
|
"manifest": {
|
||||||
|
"path": str(args.manifest),
|
||||||
|
"file_sha256": file_sha256(args.manifest),
|
||||||
|
"formal_schedule_sha256": manifest["windows"][
|
||||||
|
"formal_schedule_sha256"
|
||||||
|
],
|
||||||
|
"validation_tensor_sha256": manifest["windows"][
|
||||||
|
"validation_tensor_sha256"
|
||||||
|
],
|
||||||
|
"diagnostic_tensor_sha256": manifest["windows"][
|
||||||
|
"diagnostic_tensor_sha256"
|
||||||
|
],
|
||||||
|
"input_gate_tensor_hashes": input_gate_hashes,
|
||||||
|
},
|
||||||
|
"model": {
|
||||||
|
"layers": args.depth,
|
||||||
|
"sublayers": args.depth * 2,
|
||||||
|
"attnres_aggregation_groups": BLOCK_GROUPS,
|
||||||
|
"sublayers_per_attnres_group": args.depth * 2 // BLOCK_GROUPS,
|
||||||
|
"transformer_blocks_per_attnres_group": args.depth // BLOCK_GROUPS,
|
||||||
|
"d_model": round04.D_MODEL,
|
||||||
|
"heads": round04.HEADS,
|
||||||
|
"d_head": round04.D_HEAD,
|
||||||
|
"d_ff": round04.D_FF,
|
||||||
|
"context": CONTEXT,
|
||||||
|
"vocabulary": VOCABULARY,
|
||||||
|
"parameters": parameter_inventory(model),
|
||||||
|
},
|
||||||
|
"optimizer": {
|
||||||
|
"name": "AdamW",
|
||||||
|
"betas": list(BETAS),
|
||||||
|
"epsilon": ADAM_EPS,
|
||||||
|
"weight_decay_ndim_ge_2": WEIGHT_DECAY,
|
||||||
|
"peak_lr": PEAK_LR,
|
||||||
|
"min_lr": MIN_LR,
|
||||||
|
"warmup_steps": WARMUP_STEPS,
|
||||||
|
"grad_clip": GRAD_CLIP,
|
||||||
|
},
|
||||||
|
"hashes": {
|
||||||
|
"initial_public_parameter_structure": public_structure_hash,
|
||||||
|
"initial_public_parameter_tensors": public_tensors,
|
||||||
|
"initial_public_parameter_elements": public_elements,
|
||||||
|
"initial_public_parameters": initial_public_hash,
|
||||||
|
"initial_mixer_parameters": initial_mixer_hash,
|
||||||
|
"final_public_parameters": final_public_hash,
|
||||||
|
"final_mixer_parameters": final_mixer_hash,
|
||||||
|
"final_model_state": final_full_hash,
|
||||||
|
"final_optimizer_state": optimizer_hash,
|
||||||
|
},
|
||||||
|
"evaluations": evaluations,
|
||||||
|
"diagnostics": diagnostics,
|
||||||
|
"training_history": training_history,
|
||||||
|
"gradient_gate": gradient_gate,
|
||||||
|
"timing": timing,
|
||||||
|
"environment": {
|
||||||
|
"python": platform.python_version(),
|
||||||
|
"torch": torch.__version__,
|
||||||
|
"cuda": torch.version.cuda,
|
||||||
|
"gpu": torch.cuda.get_device_name(0),
|
||||||
|
"compute_capability": list(torch.cuda.get_device_capability(0)),
|
||||||
|
"cublas_workspace_config": os.environ["CUBLAS_WORKSPACE_CONFIG"],
|
||||||
|
"deterministic_algorithms": torch.are_deterministic_algorithms_enabled(),
|
||||||
|
"autocast": "cuda-bfloat16-forward-fp32-cross-entropy",
|
||||||
|
"compile": False,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
result["canonical_sha256_without_self"] = canonical_sha256(result)
|
||||||
|
args.output.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
temporary = args.output.with_suffix(args.output.suffix + ".tmp")
|
||||||
|
temporary.write_text(
|
||||||
|
json.dumps(result, ensure_ascii=False, indent=2, sort_keys=True) + "\n"
|
||||||
|
)
|
||||||
|
os.replace(temporary, args.output)
|
||||||
|
print(
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"output": str(args.output),
|
||||||
|
"run_kind": args.run_kind,
|
||||||
|
"architecture": args.architecture,
|
||||||
|
"depth": args.depth,
|
||||||
|
"seed": args.seed,
|
||||||
|
"steps": args.steps,
|
||||||
|
"final_bpc": evaluations[-1]["bits_per_byte"],
|
||||||
|
"final_activation_gradient_cv": diagnostics[-1][
|
||||||
|
"activation_grad_statistics"
|
||||||
|
]["population_cv"],
|
||||||
|
"canonical_sha256": result["canonical_sha256_without_self"],
|
||||||
|
"timing": timing,
|
||||||
|
},
|
||||||
|
ensure_ascii=False,
|
||||||
|
indent=2,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -20,6 +20,7 @@
|
|||||||
"build:data:deepseek-chat-task-bootstrap": "node scripts/build-deepseek-chat-task-bootstrap-crn-compact.mjs",
|
"build:data:deepseek-chat-task-bootstrap": "node scripts/build-deepseek-chat-task-bootstrap-crn-compact.mjs",
|
||||||
"check:data:deepseek-chat-task-bootstrap": "node scripts/check-deepseek-chat-task-bootstrap-crn-data.mjs",
|
"check:data:deepseek-chat-task-bootstrap": "node scripts/check-deepseek-chat-task-bootstrap-crn-data.mjs",
|
||||||
"check:data:k3-attnres": "node scripts/check-k3-attnres-data.mjs",
|
"check:data:k3-attnres": "node scripts/check-k3-attnres-data.mjs",
|
||||||
|
"check:data:k3-attnres-gradient": "node scripts/check-k3-attnres-gradient-data.mjs",
|
||||||
"check:site": "node scripts/check-site.mjs",
|
"check:site": "node scripts/check-site.mjs",
|
||||||
"check:moe-browser": "node scripts/check-moe-browser.mjs",
|
"check:moe-browser": "node scripts/check-moe-browser.mjs",
|
||||||
"check:reasoning-browser": "node scripts/check-reasoning-browser.mjs",
|
"check:reasoning-browser": "node scripts/check-reasoning-browser.mjs",
|
||||||
@@ -40,6 +41,7 @@
|
|||||||
"check:deepseek-cross-source-sampling-browser": "node scripts/check-deepseek-cross-source-sampling-browser.mjs",
|
"check:deepseek-cross-source-sampling-browser": "node scripts/check-deepseek-cross-source-sampling-browser.mjs",
|
||||||
"check:deepseek-task-bootstrap-browser": "node scripts/check-deepseek-task-bootstrap-browser.mjs",
|
"check:deepseek-task-bootstrap-browser": "node scripts/check-deepseek-task-bootstrap-browser.mjs",
|
||||||
"check:k3-attnres-browser": "node scripts/check-k3-attnres-browser.mjs",
|
"check:k3-attnres-browser": "node scripts/check-k3-attnres-browser.mjs",
|
||||||
|
"check:k3-attnres-gradient-browser": "node scripts/check-k3-attnres-gradient-browser.mjs",
|
||||||
"check:k3-browser": "node scripts/check-k3-browser.mjs"
|
"check:k3-browser": "node scripts/check-k3-browser.mjs"
|
||||||
},
|
},
|
||||||
"dependencies": {
|
"dependencies": {
|
||||||
|
|||||||
@@ -0,0 +1,198 @@
|
|||||||
|
# Kimi K3 第五轮前置审计:Attention Residuals 的“梯度更均匀”到底指什么
|
||||||
|
|
||||||
|
> 审计日期:2026-07-30(Asia/Shanghai)
|
||||||
|
> 官方仓库:`MoonshotAI/Attention-Residuals@85e22310fe5ee860b4a023de312d791de8a5a5e6`
|
||||||
|
> 官方 PDF SHA-256:`e5831b0db1347606453b5176b0142115a18887b6a9c2e1d05a266d4805a26b2f`
|
||||||
|
> 结论性质:一手工件审计,不是新的实验结果
|
||||||
|
|
||||||
|
## 0. 先说结论
|
||||||
|
|
||||||
|
K3 Round 04 的反结果——Block AttnRes 的**核心参数梯度 RMS 跨层 CV 更高**——不能直接
|
||||||
|
反驳 Attention Residuals 论文 Figure 5(c) 所说的“梯度分布更均匀”,因为两边很可能测的
|
||||||
|
不是同一个对象:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Round 04:
|
||||||
|
每个 Transformer block 内所有核心参数梯度拼接后的 RMS
|
||||||
|
∇θL,θ = attention + MLP + 两个输入 norm 的参数
|
||||||
|
|
||||||
|
论文 Figure 5(c):
|
||||||
|
图题只写 “Each transformer block's gradient magnitude”
|
||||||
|
结合 Figure 5(b) 的 block output magnitude 和正文,最自然的操作化是
|
||||||
|
每个 block 输出 activation 的梯度 ∂L/∂h_l
|
||||||
|
```
|
||||||
|
|
||||||
|
但“最自然”不等于“官方已经明确定义”。论文和当前官方仓库都没有给出足以唯一重建
|
||||||
|
Figure 5(c) 的测量合同,也没有发布训练代码。因此,下一轮不会把自己的 activation-gradient
|
||||||
|
定义冒充成论文原始实现,而会把它命名为:
|
||||||
|
|
||||||
|
> **与 Figure 5 叙述对齐的一种公开、冻结、可复现的 operationalization**
|
||||||
|
|
||||||
|
这一区分很重要:参数梯度回答“这一层的权重此刻收到多大更新信号”,activation 梯度回答
|
||||||
|
“损失对这一深度的表征有多敏感”。二者相关,但不会因为链式法则而自动同方向。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. 官方工件实际提供了什么
|
||||||
|
|
||||||
|
官方仓库在固定 revision 下只包含:
|
||||||
|
|
||||||
|
- `README.md`;
|
||||||
|
- `Attention_Residuals.pdf`;
|
||||||
|
- 论文图片资产;
|
||||||
|
- citation 与外部入口。
|
||||||
|
|
||||||
|
仓库**不包含**:
|
||||||
|
|
||||||
|
- 模型或 residual mixer 的可执行实现;
|
||||||
|
- Figure 5 的统计脚本;
|
||||||
|
- 训练配置、日志或 checkpoint;
|
||||||
|
- gradient hook、norm 和 reduction 定义;
|
||||||
|
- 用于复画 Figure 5 的原始数组。
|
||||||
|
|
||||||
|
因此,本审计能固定论文的文字、公式、图与模型尺度,不能从官方代码恢复一个不存在的
|
||||||
|
隐藏测量合同。
|
||||||
|
|
||||||
|
## 2. Figure 5 能确认的事实
|
||||||
|
|
||||||
|
官方 `training_dynamics.png` 和 PDF Figure 5 有三个并列面板:
|
||||||
|
|
||||||
|
| 面板 | 图题 | 横轴 |
|
||||||
|
|---|---|---|
|
||||||
|
| (a) | Validation Loss | training step |
|
||||||
|
| (b) | Output Magnitude | Layer / Transformer Block Index |
|
||||||
|
| (c) | Gradient Magnitude ×10⁻⁵ | Layer / Transformer Block Index |
|
||||||
|
|
||||||
|
Figure 5 caption 对 (b) 与 (c) 的完整对象描述分别是:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Each transformer block's output magnitude at the end of training.
|
||||||
|
Each transformer block's gradient magnitude.
|
||||||
|
```
|
||||||
|
|
||||||
|
正文明确表达了两个方向:
|
||||||
|
|
||||||
|
1. Baseline 的 hidden-state magnitude 随深度单调增长;Block AttnRes 把增长约束在局部
|
||||||
|
residual block 内,形成周期性的深度图案。
|
||||||
|
2. Baseline 的早期层梯度“不成比例地大”;Block AttnRes 的可学习 softmax 权重产生了
|
||||||
|
“明显更均匀”的梯度分布。
|
||||||
|
|
||||||
|
图中最终大模型约有 27 个 Transformer blocks。论文的最终模型描述也是 27 个 Transformer
|
||||||
|
blocks / 54 个 residual layers;Block AttnRes 每 6 个 residual layers 聚合一次,共 9 个
|
||||||
|
聚合块,外加 embedding 形成 10 个跨块来源。
|
||||||
|
|
||||||
|
## 3. Figure 5 不能确认的事项
|
||||||
|
|
||||||
|
下面每一项都会改变曲线,却没有在论文或官方仓库中被唯一指定:
|
||||||
|
|
||||||
|
| 未定义项 | 至少两种合理解释 |
|
||||||
|
|---|---|
|
||||||
|
| 梯度对象 | block 输出 activation 梯度;block 参数梯度;分支输出梯度 |
|
||||||
|
| block 输出位置 | attention+MLP 后;只在 MLP 后;进入下一层 norm 前;聚合块边界后 |
|
||||||
|
| norm | L2 norm;RMS;mean absolute value;每 token norm 后再平均 |
|
||||||
|
| reduction | batch/token/channel 联合;先按 token 再按 batch;只取末 token |
|
||||||
|
| loss | token mean;sample mean;未归一化 sum;带或不带 mask |
|
||||||
|
| 采样 | 一个 batch;多 batch 平均;训练流中的 moving average |
|
||||||
|
| 时间点 | “训练结束”单点;末段平均;某个 checkpoint |
|
||||||
|
| 数值阶段 | AMP 缩放前/后;gradient clipping 前/后;BF16 或 FP32 |
|
||||||
|
| 运行模式 | train 或 eval;dropout 是否开启 |
|
||||||
|
| 归一化 | 绝对值;再除全层均值;再除 Baseline |
|
||||||
|
|
||||||
|
Figure 5(c) 的纵轴是绝对 magnitude 标度,并不等价于 CV。只报告 CV 还会丢失两个信息:
|
||||||
|
|
||||||
|
- 全部层梯度是否一起缩小或放大;
|
||||||
|
- 不均匀来自“早层系统性偏大”,还是某一个中间/末端尖峰。
|
||||||
|
|
||||||
|
所以 Round 05 必须同时公开绝对曲线、按层均值归一化曲线、CV 与前后深度分位比。
|
||||||
|
|
||||||
|
## 4. 为什么 activation gradient 是合理推断,但仍只是推断
|
||||||
|
|
||||||
|
把 Figure 5(c) 操作化为 `∂L/∂h_l` 有三条证据:
|
||||||
|
|
||||||
|
1. 它与 Figure 5(b) 的 “transformer block output magnitude” 在横轴和叙述上成对;
|
||||||
|
2. “早期层的梯度”在表示传播语境中通常可由对 block output 保留梯度直接比较;
|
||||||
|
3. 参数张量的大小和类型在 attention 与 MLP 间差异很大,若把参数拼接,论文通常需要说明
|
||||||
|
聚合口径,否则 “each transformer block” 不是天然的单一标量。
|
||||||
|
|
||||||
|
但也有无法排除的替代解释:
|
||||||
|
|
||||||
|
- 论文作者可能测 block 参数梯度;
|
||||||
|
- 可能测 residual branch output 而非完整 block output;
|
||||||
|
- 可能先对每个 token 做 L2 norm,再跨 token 平均;
|
||||||
|
- 可能在内部训练系统中有未公开的统一 telemetry 定义。
|
||||||
|
|
||||||
|
因此,网站和审计只说“与论文叙述对齐的公开定义”,不说“论文就是这样算的”,也不把
|
||||||
|
数值和 Figure 5 纵轴直接对齐。
|
||||||
|
|
||||||
|
## 5. Round 04 与 Round 05 的对象对照
|
||||||
|
|
||||||
|
| 维度 | Round 04 已测对象 | Round 05 主对象 |
|
||||||
|
|---|---|---|
|
||||||
|
| 数学对象 | `∇θ_l L` | `∂L/∂h_l` |
|
||||||
|
| `l` 的单位 | Transformer block | Transformer block |
|
||||||
|
| 张量内容 | block 的核心参数 | block 的 post-MLP output activation |
|
||||||
|
| 聚合 | 参数元素联合 RMS | batch×time×channel 联合 RMS |
|
||||||
|
| 是否含 AttnRes 参数 | 否 | 不适用;梯度穿过 mixer |
|
||||||
|
| 时间 | final diagnostic batch | 全部预注册 diagnostic steps |
|
||||||
|
| 目的 | 权重更新信号是否均匀 | 表征深度的反向信号是否均匀 |
|
||||||
|
|
||||||
|
Round 04 的参数梯度结果不会被改名、删去或用新指标覆盖。它仍是一个有效反结果,只是不能
|
||||||
|
代表论文未定义清楚的 Figure 5(c)。
|
||||||
|
|
||||||
|
## 6. 两种结构怎样取得真正对齐的 16 / 32 个位置
|
||||||
|
|
||||||
|
Grok Headless 被用作一次对抗式方法审阅,不作为事实来源。它正确指出了 non-leaf tensor、
|
||||||
|
alias、AMP、clip 时点和只看 CV 的风险;但它也提出了一个不适用于本实现的担忧:
|
||||||
|
“Block AttnRes 只有约 8 个 block 输出,无法与 Baseline 的 16 / 32 层对齐”。
|
||||||
|
|
||||||
|
这里要区分两种 block:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Transformer block:
|
||||||
|
attention + MLP;depth=16 时始终有 16 个,depth=32 时始终有 32 个
|
||||||
|
|
||||||
|
AttnRes aggregation group:
|
||||||
|
把若干 residual sublayers 的 partial sum 保存为一个跨组 source;
|
||||||
|
两种深度都约为 8 组
|
||||||
|
```
|
||||||
|
|
||||||
|
Round 05 在**每个 Transformer block 的 MLP 分支完成后**取 `h_l`。所以 Baseline 与 Block
|
||||||
|
都有完全相同的 `l=1..depth`:
|
||||||
|
|
||||||
|
| 架构 | Round 05 的 `h_l` | shape |
|
||||||
|
|---|---|---|
|
||||||
|
| Baseline | 第 `l` 个 attention residual 与 MLP residual 都完成后的 hidden state | `[B,T,C]` |
|
||||||
|
| Block | 第 `l` 个 MLP branch 加入后、可能保存并清空 aggregation partial **之前**的 partial output | `[B,T,C]` |
|
||||||
|
|
||||||
|
Block 的 `h_l` 是局部 residual partial,而不是跨组 source 列表;这正对应论文 Figure 5(b)
|
||||||
|
所描述的“增长被限制在每个 block 内”的周期性图案。组边界前取值也避免把 reset 后的零张量
|
||||||
|
错误当成 Transformer block output。
|
||||||
|
|
||||||
|
## 7. 实验实现必须通过的梯度测量闸门
|
||||||
|
|
||||||
|
正式训练前,四个结构格(2 个深度 × 2 个 residual graph)都必须证明:
|
||||||
|
|
||||||
|
1. 所有 `h_l` 都是不同的、非别名的捕获对象,数量严格等于 Transformer depth;
|
||||||
|
2. `retain_grad()` 或 hook 后所有梯度非 `None`、finite,shape 与 activation 完全一致;
|
||||||
|
3. diagnostic loss 使用固定输入、固定 token-mean CE、`eval()`、无 optimizer step;
|
||||||
|
4. backward 发生在 parameter gradient clip 之前,且不经过 `GradScaler`;
|
||||||
|
5. 将同一 diagnostic loss 精确乘 2 后,每层 activation-gradient RMS 也乘 2;
|
||||||
|
6. 乘 2 前后的 CV、归一化曲线与深度分位比在数值容差内不变;
|
||||||
|
7. 同配置全新进程重复运行,冻结诊断字段 exact。
|
||||||
|
|
||||||
|
若任一项失败,正式 8,000-step grid 不得开始。
|
||||||
|
|
||||||
|
## 8. 本审计带来的研究决策
|
||||||
|
|
||||||
|
下一轮不再把一个宽度 192、深度 16、训练 2,000 step 的参数梯度 CV 与论文最终模型图强行
|
||||||
|
放在同一条结论线上,而是:
|
||||||
|
|
||||||
|
- 增加 depth 32;
|
||||||
|
- 把训练预算扩为 8,000 step;
|
||||||
|
- Baseline / Block 使用相同的 Transformer block index;
|
||||||
|
- 在 6 个固定时点测 post-MLP activation gradient;
|
||||||
|
- 绝对标度、归一化形状、CV、前后四分位失衡一起公开;
|
||||||
|
- 参数梯度作为次要指标保留;
|
||||||
|
- 预注册“支持 / 混合 / 不支持”规则,并公开全部反结果。
|
||||||
|
|
||||||
|
精确协议见 `research/K3_ATTNRES_GRADIENT_SCALE_PROTOCOL.md`。
|
||||||
@@ -0,0 +1,545 @@
|
|||||||
|
# Kimi K3 第五轮:Attention Residuals 梯度定义与深度扩展实验审计
|
||||||
|
|
||||||
|
> 协议:`llm-atlas-k3-attnres-gradient-scale-v1`
|
||||||
|
> 前置定义审计:`research/K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md`
|
||||||
|
> 预注册协议:`research/K3_ATTNRES_GRADIENT_SCALE_PROTOCOL.md`
|
||||||
|
> 数据清单:`experiments/k3/attnres_gradient/manifest.json`
|
||||||
|
> 执行日期:2026-07-30
|
||||||
|
> 设备:NVIDIA GeForce RTX 5090;PyTorch `2.11.0+cu128`
|
||||||
|
|
||||||
|
## 0. 先说结论
|
||||||
|
|
||||||
|
这一轮本来想澄清一个看似矛盾的问题:
|
||||||
|
|
||||||
|
```text
|
||||||
|
论文 Figure 5:
|
||||||
|
Block AttnRes 的梯度沿深度“明显更均匀”
|
||||||
|
|
||||||
|
本站 Round 04:
|
||||||
|
Block AttnRes 的核心参数梯度 RMS 跨层 CV 反而更高
|
||||||
|
```
|
||||||
|
|
||||||
|
一手工件审计先确认:论文没有公开 Figure 5(c) 的确切 gradient tensor、norm、reduction、
|
||||||
|
diagnostic batch、AMP / clip 时点或统计代码。因此,Round 05 没有假装恢复作者的隐藏实现,
|
||||||
|
而是冻结一个与 Figure 5 的 output/gradient 并列叙述对齐、可复现的定义:
|
||||||
|
|
||||||
|
```text
|
||||||
|
h_l:
|
||||||
|
第 l 个 Transformer block 完成 attention + MLP 后的 FP32 residual output
|
||||||
|
|
||||||
|
m_l:
|
||||||
|
sqrt(mean((∂L / ∂h_l)² over batch × time × channel))
|
||||||
|
|
||||||
|
L:
|
||||||
|
固定 16 × 256 targets 的 token-mean cross entropy
|
||||||
|
```
|
||||||
|
|
||||||
|
结果不是一句“是”或“不是”,而是一个更有信息量的分解:
|
||||||
|
|
||||||
|
1. **Block 确实大幅缓解了早层整体偏大。**
|
||||||
|
Baseline 首四分位梯度平均是末四分位的 `3.01–3.90×`;Block 把它改到
|
||||||
|
`0.55–1.40×`。预注册的首尾失衡指标在 6 / 6 个 depth×seed 配对中都改善,
|
||||||
|
depth-16 平均改善 `61.0%`,depth-32 平均改善 `72.0%`。
|
||||||
|
|
||||||
|
2. **但 Block 没有让整条深度谱更平。**
|
||||||
|
它在中后段形成了局部尖峰,所以 population CV 在 6 / 6 个配对中都恶化:
|
||||||
|
depth-16 平均相对恶化 `10.3%`,depth-32 平均相对恶化 `60.0%`。
|
||||||
|
|
||||||
|
3. **Block 的绝对 activation-gradient 平均尺度更小。**
|
||||||
|
它只有 Baseline 的 `57.4%`(depth-16)和 `54.4%`(depth-32)。这意味着“首尾更接近”
|
||||||
|
不能自动解释为所有层都获得更强更新信号。
|
||||||
|
|
||||||
|
4. **参数梯度仍与 Round 04 同方向。**
|
||||||
|
核心参数梯度 CV 从 `0.416→0.683`(depth-16),从 `0.397→0.772`
|
||||||
|
(depth-32);Block 更不均匀。
|
||||||
|
|
||||||
|
5. **验证 BPC 在 6 / 6 配对中都更低,但不是同算力优势。**
|
||||||
|
平均改善 `0.00894 BPC`(depth-16)与 `0.00987 BPC`(depth-32);Block 实际 step
|
||||||
|
time 是 Baseline 的约 `2.55–2.60×`,peak allocated memory 约 `2.15–2.16×`。
|
||||||
|
|
||||||
|
按看结果前冻结的联合判据,CV 与首尾失衡必须同时改善才算 support;必须同时恶化才算
|
||||||
|
concern。这里二者方向相反,所以:
|
||||||
|
|
||||||
|
| depth | 预注册判定 |
|
||||||
|
|---:|---|
|
||||||
|
| 16 | **mixed / inconclusive at this depth** |
|
||||||
|
| 32 | **mixed / inconclusive at this depth** |
|
||||||
|
| 总判定 | **depth-dependent or inconclusive** |
|
||||||
|
|
||||||
|
最准确的中文总结是:
|
||||||
|
|
||||||
|
> 在这套公开 operationalization 中,Block AttnRes 把“早层系统性偏大”变成了“首尾更接近、
|
||||||
|
> 但中后段有局部尖峰”的另一种梯度分布。它修正了一类失衡,却没有降低全层离散度。
|
||||||
|
|
||||||
|
这不复现论文 Figure 5 的数值,也不反驳一个没有公开测量合同的隐藏实现。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. 为什么要另开一轮,而不是改写 Round 04
|
||||||
|
|
||||||
|
Round 04 的梯度对象是:
|
||||||
|
|
||||||
|
```text
|
||||||
|
每个 Transformer block 的 attention、MLP 与两个 input norm
|
||||||
|
所有核心参数梯度拼接后的 RMS
|
||||||
|
```
|
||||||
|
|
||||||
|
这回答“该层权重收到多大更新信号”。Round 05 的对象是 `∂L/∂h_l`,回答“损失对该深度
|
||||||
|
表征有多敏感”。链式法则把两者联系起来,但不会保证跨层形状同方向。
|
||||||
|
|
||||||
|
因此:
|
||||||
|
|
||||||
|
- Round 04 参数梯度反结果继续有效;
|
||||||
|
- Round 05 不把它改名为 activation gradient;
|
||||||
|
- 两种对象在网站并排展示;
|
||||||
|
- 任何方向冲突都保留,而不是选择更像论文的一种。
|
||||||
|
|
||||||
|
官方工件边界见前置定义审计。固定 revision 为:
|
||||||
|
|
||||||
|
```text
|
||||||
|
MoonshotAI/Attention-Residuals
|
||||||
|
85e22310fe5ee860b4a023de312d791de8a5a5e6
|
||||||
|
|
||||||
|
Attention_Residuals.pdf SHA-256
|
||||||
|
e5831b0db1347606453b5176b0142115a18887b6a9c2e1d05a266d4805a26b2f
|
||||||
|
```
|
||||||
|
|
||||||
|
官方仓库没有可执行训练代码或 Figure 5 原始数组。
|
||||||
|
|
||||||
|
## 2. 实验规模与配对合同
|
||||||
|
|
||||||
|
正式网格:
|
||||||
|
|
||||||
|
```text
|
||||||
|
2 depths
|
||||||
|
× 2 residual graphs
|
||||||
|
× 3 seeds
|
||||||
|
× 8,000 steps
|
||||||
|
× 32 windows
|
||||||
|
× 256 target bytes
|
||||||
|
= 12 independent runs
|
||||||
|
= 786,432,000 formal target bytes
|
||||||
|
```
|
||||||
|
|
||||||
|
| 项 | depth-16 | depth-32 |
|
||||||
|
|---|---:|---:|
|
||||||
|
| Transformer blocks | 16 | 32 |
|
||||||
|
| residual sublayers | 32 | 64 |
|
||||||
|
| AttnRes aggregation groups | 8 | 8 |
|
||||||
|
| sublayers / group | 4 | 8 |
|
||||||
|
| Transformer blocks / group | 2 | 4 |
|
||||||
|
| width / heads / FFN | 192 / 6 / 768 | 192 / 6 / 768 |
|
||||||
|
|
||||||
|
每个 depth / seed 的 Baseline 与 Block:
|
||||||
|
|
||||||
|
- 使用逐 tensor exact 的公共主干初始化;
|
||||||
|
- 使用逐 step / row exact 的 byte windows;
|
||||||
|
- 使用相同 optimizer、LR schedule、batch、context 与 target-byte budget;
|
||||||
|
- 分支线性计算走 BF16 autocast;
|
||||||
|
- residual accumulator 与被测 `h_l` 都为 FP32;
|
||||||
|
- 每个结构在全新进程中从零训练。
|
||||||
|
|
||||||
|
二者**不匹配**:
|
||||||
|
|
||||||
|
- mixer 参数;
|
||||||
|
- mixer FLOPs;
|
||||||
|
- step wall time;
|
||||||
|
- activation memory。
|
||||||
|
|
||||||
|
所以 BPC 只能叫“同 token / step 预算对比”,不能叫“同算力优势”。
|
||||||
|
|
||||||
|
## 3. 数据与日程
|
||||||
|
|
||||||
|
| 对象 | 固定值 |
|
||||||
|
|---|---|
|
||||||
|
| dataset | `Salesforce/wikitext` |
|
||||||
|
| revision | `b08601e04326c79dfdd32d625aee71d232d685c3` |
|
||||||
|
| variant | `wikitext-2-raw-v1` |
|
||||||
|
| train bytes | 10,951,563 |
|
||||||
|
| train SHA-256 | `0ca7d3e7…e9b4` |
|
||||||
|
| validation bytes | 1,148,008 |
|
||||||
|
| validation SHA-256 | `a42356f6…e719` |
|
||||||
|
| formal schedule cells | 768,000 |
|
||||||
|
| schedule SHA-256 | `5041e09b…f4e` |
|
||||||
|
| validation tensor SHA-256 | `f459316f…338` |
|
||||||
|
| diagnostic tensor SHA-256 | `21117e31…716` |
|
||||||
|
|
||||||
|
固定诊断时点:
|
||||||
|
|
||||||
|
```text
|
||||||
|
0, 100, 500, 2,000, 4,000, 8,000
|
||||||
|
```
|
||||||
|
|
||||||
|
六个时点的全部 activation gradient、output RMS、parameter gradient、mixer 权重与验证 BPC
|
||||||
|
都进入 raw JSON;没有只挑“最好看”的 checkpoint。
|
||||||
|
|
||||||
|
## 4. 正式训练前的故障与修订
|
||||||
|
|
||||||
|
### 4.1 第一次 smoke 的标量序列化错误
|
||||||
|
|
||||||
|
首个 depth-16 Baseline 20-step smoke 已完成数值计算,但在写 JSON 前失败:
|
||||||
|
|
||||||
|
```text
|
||||||
|
RuntimeError:
|
||||||
|
self.dim() cannot be 0 to view Float as Byte
|
||||||
|
```
|
||||||
|
|
||||||
|
原因是 optimizer 的 step 是 0 维 tensor,hash helper 直接把它 `view(torch.uint8)`。修复为:
|
||||||
|
|
||||||
|
```text
|
||||||
|
tensor.reshape(-1).view(torch.uint8)
|
||||||
|
```
|
||||||
|
|
||||||
|
当时:
|
||||||
|
|
||||||
|
- 没有 formal 运行;
|
||||||
|
- 没有输出 JSON;
|
||||||
|
- 没有可供选择的正式结果。
|
||||||
|
|
||||||
|
### 4.2 smoke 发现被测 residual dtype 不一致
|
||||||
|
|
||||||
|
第一版 smoke 通过了 finite / loss×2 / replay 闸门,但检查 capture metadata 时发现:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Baseline post-MLP h_l:FP32
|
||||||
|
Block aggregation partial:BF16
|
||||||
|
```
|
||||||
|
|
||||||
|
原因:
|
||||||
|
|
||||||
|
- Baseline 把 BF16 branch 加到 FP32 embedding/residual stream;
|
||||||
|
- Block 每组第一个 partial 直接引用 BF16 branch output。
|
||||||
|
|
||||||
|
这会把数值精度差异混进结构比较。正式训练前,协议与实现补充为:
|
||||||
|
|
||||||
|
```text
|
||||||
|
BF16 branch output → 显式转 FP32 → residual partial 累加
|
||||||
|
```
|
||||||
|
|
||||||
|
随后 4 个 depth×architecture 格的 smoke 全部从头重跑两次。正式输出是在这次修订之后才开始。
|
||||||
|
|
||||||
|
### 4.3 一次未启动模型的 zsh 调度错误
|
||||||
|
|
||||||
|
首个正式 Baseline 完成后,批处理脚本用 Bash 式标量切分处理 zsh 字符串,第一行就退出:
|
||||||
|
|
||||||
|
```text
|
||||||
|
argument --architecture: invalid choice: ''
|
||||||
|
```
|
||||||
|
|
||||||
|
runner 没有启动,也没有创建新结果文件。调度改为显式 `:` 分隔数组后继续。这是 orchestration
|
||||||
|
故障,不是模型运行失败,但仍在审计时间线中保留。
|
||||||
|
|
||||||
|
## 5. 运行前与复现闸门
|
||||||
|
|
||||||
|
### 5.1 activation-gradient 测量闸门
|
||||||
|
|
||||||
|
4 / 4 个 depth×architecture 格都通过:
|
||||||
|
|
||||||
|
- 捕获数量严格等于 16 / 32;
|
||||||
|
- shape 严格为 `[16,256,192]`;
|
||||||
|
- dtype 全部为 FP32;
|
||||||
|
- gradient 全部 present、finite;
|
||||||
|
- capture storage 全部不别名;
|
||||||
|
- diagnostic loss×2 后每层 gradient RMS 精确×2;
|
||||||
|
- CV、归一化谱、首尾比不变。
|
||||||
|
|
||||||
|
### 5.2 两次独立 smoke
|
||||||
|
|
||||||
|
每个格都在两个全新进程中训练 20 steps。排除 timing 后的冻结字段:
|
||||||
|
|
||||||
|
| 格 | compare SHA-256 |
|
||||||
|
|---|---|
|
||||||
|
| depth-16 Baseline | `525cdcd9…bd2a` |
|
||||||
|
| depth-16 Block | `e4330a0d…90cc` |
|
||||||
|
| depth-32 Baseline | `8ee0dc37…8e56` |
|
||||||
|
| depth-32 Block | `289073b1…34c` |
|
||||||
|
|
||||||
|
4 / 4 exact。
|
||||||
|
|
||||||
|
### 5.3 完整 formal replay
|
||||||
|
|
||||||
|
预注册格:
|
||||||
|
|
||||||
|
```text
|
||||||
|
depth-32 / Block / seed-2026073001
|
||||||
|
```
|
||||||
|
|
||||||
|
从初始化重新训练完整 8,000 steps,不加载 formal checkpoint。排除 `run_kind`、timing、
|
||||||
|
memory 与进程元数据后的全部冻结字段:
|
||||||
|
|
||||||
|
```text
|
||||||
|
formal compare SHA-256
|
||||||
|
46300a452840a9dc6a5180efe7949cf3d81cc4471ed4a942da407e9343064817
|
||||||
|
|
||||||
|
replay compare SHA-256
|
||||||
|
46300a452840a9dc6a5180efe7949cf3d81cc4471ed4a942da407e9343064817
|
||||||
|
```
|
||||||
|
|
||||||
|
最终状态:
|
||||||
|
|
||||||
|
| 对象 | formal | replay |
|
||||||
|
|---|---|---|
|
||||||
|
| model state | `3f0b97ec…2f59` | `3f0b97ec…2f59` |
|
||||||
|
| optimizer state | `ed03e6fb…4637` | `ed03e6fb…4637` |
|
||||||
|
|
||||||
|
字段级 exact。
|
||||||
|
|
||||||
|
## 6. 主结果:CV 与首尾失衡为什么方向相反
|
||||||
|
|
||||||
|
### 6.1 depth-16
|
||||||
|
|
||||||
|
| seed | Base CV | Block CV | 相对 CV 改善 | Base imbalance | Block imbalance | 相对 imbalance 改善 |
|
||||||
|
|---:|---:|---:|---:|---:|---:|---:|
|
||||||
|
| 2026073001 | 0.41719 | 0.46768 | −12.1% | 1.29458 | 0.38821 | +70.0% |
|
||||||
|
| 2026073002 | 0.43744 | 0.45175 | −3.3% | 1.35991 | 0.58958 | +56.6% |
|
||||||
|
| 2026073003 | 0.41927 | 0.48391 | −15.4% | 1.29826 | 0.56693 | +56.3% |
|
||||||
|
| 均值 | 0.42463 | 0.46778 | **−10.3%** | 1.31759 | 0.51491 | **+61.0%** |
|
||||||
|
|
||||||
|
这里“相对 CV 改善”为负,表示恶化。
|
||||||
|
|
||||||
|
原始首/末四分位比:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Baseline:3.65×, 3.90×, 3.66×
|
||||||
|
Block: 0.68×, 0.55×, 0.57×
|
||||||
|
```
|
||||||
|
|
||||||
|
Block 不只是把早层优势降到 1;它在三个 seed 中都略微“过冲”,变成末四分位平均更大。
|
||||||
|
但 `abs(log(first/last))` 仍比 Baseline 更接近 0,所以失衡改善。
|
||||||
|
|
||||||
|
三 seed 平均 normalized activation-gradient 的最高点:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Baseline:layer 3 = 1.50× mean;layer 4 = 1.42×
|
||||||
|
Block: layer 11 = 2.09× mean;layer 13 = 1.95×
|
||||||
|
```
|
||||||
|
|
||||||
|
Baseline 是宽而平滑的早层隆起;Block 是更局部的中后段尖峰。CV 对尖峰敏感,所以升高。
|
||||||
|
|
||||||
|
### 6.2 depth-32
|
||||||
|
|
||||||
|
| seed | Base CV | Block CV | 相对 CV 改善 | Base imbalance | Block imbalance | 相对 imbalance 改善 |
|
||||||
|
|---:|---:|---:|---:|---:|---:|---:|
|
||||||
|
| 2026073001 | 0.36109 | 0.64203 | −77.8% | 1.10030 | 0.20781 | +81.1% |
|
||||||
|
| 2026073002 | 0.37162 | 0.72997 | −96.4% | 1.14087 | 0.43810 | +61.6% |
|
||||||
|
| 2026073003 | 0.40320 | 0.42685 | −5.9% | 1.25169 | 0.33417 | +73.3% |
|
||||||
|
| 均值 | 0.37864 | 0.59962 | **−60.0%** | 1.16429 | 0.32669 | **+72.0%** |
|
||||||
|
|
||||||
|
原始首/末四分位比:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Baseline:3.01×, 3.13×, 3.50×
|
||||||
|
Block: 0.81×, 0.65×, 1.40×
|
||||||
|
```
|
||||||
|
|
||||||
|
三 seed 平均 normalized spectrum 的 Block 峰值:
|
||||||
|
|
||||||
|
```text
|
||||||
|
layer 21 = 3.04× mean
|
||||||
|
layer 22 = 2.41×
|
||||||
|
layer 23 = 1.91×
|
||||||
|
layer 25 = 1.77×
|
||||||
|
```
|
||||||
|
|
||||||
|
depth-32 每个 AttnRes aggregation group 含 4 个 Transformer blocks。21–24 是第 6 组,
|
||||||
|
25–28 是第 7 组。尖峰集中在这两个中后段组附近,是数据中直接可见的结构;但仅凭本实验
|
||||||
|
不能断言 pseudo-query、某个 source 或组边界是唯一因果。
|
||||||
|
|
||||||
|
### 6.3 seed-3 的中期反例
|
||||||
|
|
||||||
|
depth-32 seed-3 在 step 2,000:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Baseline CV 0.39965
|
||||||
|
Block CV 0.34914
|
||||||
|
```
|
||||||
|
|
||||||
|
此时 Block 更平;到 step 8,000 才变成:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Baseline CV 0.40320
|
||||||
|
Block CV 0.42685
|
||||||
|
```
|
||||||
|
|
||||||
|
前两个 seed 在 step 2,000 已明显恶化,seed-3 没有。网站必须保留 seed switch,不能用最终
|
||||||
|
均值倒写成“三个 seed 从头到尾都一样”。
|
||||||
|
|
||||||
|
## 7. 绝对梯度尺度:更平不等于更强
|
||||||
|
|
||||||
|
最终 activation-gradient mean:
|
||||||
|
|
||||||
|
| depth | Baseline | Block | Block / Baseline |
|
||||||
|
|---:|---:|---:|---:|
|
||||||
|
| 16 | 约 `2.02×10⁻⁴` | 约 `1.16×10⁻⁴` | **0.574×** |
|
||||||
|
| 32 | 约 `1.44×10⁻⁴` | 约 `0.78×10⁻⁴` | **0.544×** |
|
||||||
|
|
||||||
|
所以 Block 的首尾比更接近 1,并不是因为它把晚层全部抬高到 Baseline 早层的强度。更接近的
|
||||||
|
描述是:
|
||||||
|
|
||||||
|
> 整体尺度下降,早层系统性高值被削弱,同时某些中后段位置相对全层均值形成尖峰。
|
||||||
|
|
||||||
|
这也是只看 normalized curve 或只看 CV 都不够的原因。
|
||||||
|
|
||||||
|
## 8. 参数梯度没有翻转 Round 04
|
||||||
|
|
||||||
|
最终核心参数梯度 CV:
|
||||||
|
|
||||||
|
| depth | Baseline mean | Block mean | Block / Baseline |
|
||||||
|
|---:|---:|---:|---:|
|
||||||
|
| 16 | 0.41596 | 0.68287 | 1.64× |
|
||||||
|
| 32 | 0.39661 | 0.77176 | 1.95× |
|
||||||
|
|
||||||
|
6 / 6 个配对中 Block 都更高。Round 04 的反结果不是在把梯度对象改成 activation 后自动消失;
|
||||||
|
两种梯度对象在本轮 final endpoint 都显示更高的跨层 CV。
|
||||||
|
|
||||||
|
但 activation gradient 又显示首尾失衡大幅改善,这说明“均匀”至少要拆成:
|
||||||
|
|
||||||
|
```text
|
||||||
|
首尾是否平衡
|
||||||
|
全层是否有尖峰
|
||||||
|
绝对尺度多大
|
||||||
|
参数更新信号是否平衡
|
||||||
|
```
|
||||||
|
|
||||||
|
一个标量不能代替全部。
|
||||||
|
|
||||||
|
## 9. Output RMS:最接近论文叙述的正向结果
|
||||||
|
|
||||||
|
三 seed 最终 post-MLP output RMS 的最后/第一层比:
|
||||||
|
|
||||||
|
| depth | Baseline | Block |
|
||||||
|
|---:|---:|---:|
|
||||||
|
| 16 | 4.59× | 1.17× |
|
||||||
|
| 32 | 6.08× | 1.89× |
|
||||||
|
|
||||||
|
Baseline output magnitude 随深度明显累积;Block 把增长限制在 aggregation group 内并产生
|
||||||
|
周期性 reset。这个缩小实验的 output-RMS 方向与论文 Figure 5(b) 的叙述一致。
|
||||||
|
|
||||||
|
仍不能把数值直接叠到论文图上:
|
||||||
|
|
||||||
|
- 模型宽度、深度和数据不同;
|
||||||
|
- 训练 Token 相差巨大;
|
||||||
|
- 论文图的确切 output norm / reduction 也未完整公开;
|
||||||
|
- 本实验的 Block group 是 8 组固定设计。
|
||||||
|
|
||||||
|
## 10. 验证 BPC 与真实成本
|
||||||
|
|
||||||
|
最终 BPC:
|
||||||
|
|
||||||
|
### depth-16
|
||||||
|
|
||||||
|
| seed | Baseline | Block | Block − Base |
|
||||||
|
|---:|---:|---:|---:|
|
||||||
|
| 2026073001 | 1.74880 | 1.73739 | −0.01141 |
|
||||||
|
| 2026073002 | 1.73683 | 1.73118 | −0.00565 |
|
||||||
|
| 2026073003 | 1.73637 | 1.72661 | −0.00976 |
|
||||||
|
| 均值 | 1.74067 | 1.73173 | **−0.00894** |
|
||||||
|
|
||||||
|
### depth-32
|
||||||
|
|
||||||
|
| seed | Baseline | Block | Block − Base |
|
||||||
|
|---:|---:|---:|---:|
|
||||||
|
| 2026073001 | 1.71790 | 1.71235 | −0.00554 |
|
||||||
|
| 2026073002 | 1.72698 | 1.70932 | −0.01765 |
|
||||||
|
| 2026073003 | 1.70951 | 1.70310 | −0.00642 |
|
||||||
|
| 均值 | 1.71813 | 1.70826 | **−0.00987** |
|
||||||
|
|
||||||
|
6 / 6 为负。这是有价值的次要方向,但 Round 05 没有为 BPC 再预注册一个新的 support 阈值,
|
||||||
|
因此不追加事后显著性结论。
|
||||||
|
|
||||||
|
真实成本:
|
||||||
|
|
||||||
|
| depth | Base mean ms | Block mean ms | time ratio | Base peak alloc | Block peak alloc | memory ratio |
|
||||||
|
|---:|---:|---:|---:|---:|---:|---:|
|
||||||
|
| 16 | 21.47 | 54.71 | 2.55× | 3.04 GB | 6.54 GB | 2.15× |
|
||||||
|
| 32 | 42.11 | 109.38 | 2.60× | 5.94 GB | 12.82 GB | 2.16× |
|
||||||
|
|
||||||
|
这个简单 eager 实现没有论文训练系统的 kernel、并行或工程优化;成本数值不应外推到 K3。
|
||||||
|
但它足以说明本站的 BPC 对比不是同 wall time / FLOPs。
|
||||||
|
|
||||||
|
## 11. 预注册判定为何是 mixed
|
||||||
|
|
||||||
|
支持需要:
|
||||||
|
|
||||||
|
```text
|
||||||
|
CV:3 / 3 seeds 改善,平均相对改善 ≥20%
|
||||||
|
AND
|
||||||
|
imbalance:3 / 3 seeds 改善,平均相对改善 ≥20%
|
||||||
|
```
|
||||||
|
|
||||||
|
concern 需要两项都以相同规则恶化。
|
||||||
|
|
||||||
|
实际:
|
||||||
|
|
||||||
|
```text
|
||||||
|
depth-16:
|
||||||
|
CV 3 / 3 恶化,平均 10.3%
|
||||||
|
imbalance 3 / 3 改善,平均 61.0%
|
||||||
|
|
||||||
|
depth-32:
|
||||||
|
CV 3 / 3 恶化,平均 60.0%
|
||||||
|
imbalance 3 / 3 改善,平均 72.0%
|
||||||
|
```
|
||||||
|
|
||||||
|
两个指标相反,所以两个 depth 都是 `mixed / inconclusive at this depth`。这不是“数据没规律”,
|
||||||
|
而是预注册的“更均匀”概念被实验拆成了两个方向相反的组成部分。
|
||||||
|
|
||||||
|
## 12. 开放工件与校验哈希
|
||||||
|
|
||||||
|
| 工件 | SHA-256 |
|
||||||
|
|---|---|
|
||||||
|
| manifest | `080afb17…6371` |
|
||||||
|
| protocol | `f772629b…ce22` |
|
||||||
|
| definition audit | `79221c56…cc6` |
|
||||||
|
| runner | `04ae69e1…800f` |
|
||||||
|
| analyzer | `017d38d9…2fc8` |
|
||||||
|
| aggregate canonical | `69be133c…7b51` |
|
||||||
|
| compact canonical | `8cdb7180…044f` |
|
||||||
|
| reproduction canonical | `addb2e59…f68a` |
|
||||||
|
|
||||||
|
公开目录包含:
|
||||||
|
|
||||||
|
- 12 个完整 formal raw JSON;
|
||||||
|
- 8 个两套 smoke raw JSON;
|
||||||
|
- 1 个完整 replay raw JSON;
|
||||||
|
- 完整 aggregate;
|
||||||
|
- 网站 compact payload;
|
||||||
|
- manifest、runner、analyzer 与 reproduction 清单;
|
||||||
|
- 前置定义审计、本协议和本结果审计。
|
||||||
|
|
||||||
|
`experiments/k3/attnres_gradient/reproduction.json` 记录每个 raw 文件 SHA-256。
|
||||||
|
|
||||||
|
## 13. 允许和禁止的结论
|
||||||
|
|
||||||
|
允许:
|
||||||
|
|
||||||
|
> 在本轮公开定义下,Block AttnRes 一致缓解了首/末深度四分位失衡,但在中后段形成局部
|
||||||
|
> 梯度尖峰,导致全层 CV 一致升高;因此“梯度更均匀”必须拆成多个指标解释。
|
||||||
|
|
||||||
|
> 同 token / step 预算下,Block 的最终验证 BPC 在 6 / 6 个配对中更低,但实际运行成本
|
||||||
|
> 约为 Baseline 的 2.6× step time 与 2.2× peak allocated memory。
|
||||||
|
|
||||||
|
禁止:
|
||||||
|
|
||||||
|
- “复现了论文 Figure 5(c)”;
|
||||||
|
- “论文的梯度结论是错的”;
|
||||||
|
- “已测到 Kimi K3 checkpoint 的真实梯度”;
|
||||||
|
- “Block 解决了梯度消失 / 爆炸”;
|
||||||
|
- “CV 更高就代表训练一定更不稳定”;
|
||||||
|
- “BPC 改善是同 FLOPs / wall time 优势”;
|
||||||
|
- 从 3 seeds 推导总体显著性;
|
||||||
|
- 从 depth 16 / 32 外推到 48B、1T+400B Token 或 K3 2.8T 参数。
|
||||||
|
|
||||||
|
## 14. 下一步最值得问什么
|
||||||
|
|
||||||
|
这轮已经把“梯度”从一个模糊词拆成了可复查对象。下一个有价值的问题不是再换一个漂亮
|
||||||
|
汇总指标,而是追踪局部尖峰从哪里来:
|
||||||
|
|
||||||
|
1. 分开捕获 pre-attention 与 pre-MLP residual positions;
|
||||||
|
2. 把 layer 21–25 的 activation-gradient 与 mixer source weights 同步对齐;
|
||||||
|
3. 比较 aggregation-group boundary 前后;
|
||||||
|
4. 在不改变 formal 结果的前提下,对相同 raw gradient tensor 做多种公开 reduction
|
||||||
|
sensitivity analysis;
|
||||||
|
5. 若官方之后发布 Figure 5 telemetry 代码,再按其定义单独开新 protocol。
|
||||||
|
|
||||||
|
这些属于后续轮次,不能倒写进本轮预注册结论。
|
||||||
@@ -0,0 +1,459 @@
|
|||||||
|
# Kimi K3 第五轮:Attention Residuals 梯度定义与深度扩展实验协议
|
||||||
|
|
||||||
|
> 协议 ID:`llm-atlas-k3-attnres-gradient-scale-v1`
|
||||||
|
> 冻结日期:2026-07-30(Asia/Shanghai)
|
||||||
|
> 状态:正式输出前预注册
|
||||||
|
> 前置定义审计:`research/K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md`
|
||||||
|
|
||||||
|
## 0. 目标与一句话研究问题
|
||||||
|
|
||||||
|
Round 04 在缩小模型中得到两个同时成立的结果:
|
||||||
|
|
||||||
|
- Full / Block AttnRes 的 2,000-step 验证 BPC 都优于 Baseline;
|
||||||
|
- 以“每个 Transformer block 的核心**参数**梯度 RMS”定义时,跨深度 CV 反而更高。
|
||||||
|
|
||||||
|
论文 Figure 5(c) 没有公开足以唯一恢复的梯度测量合同。Round 05 不猜作者的隐藏代码,而是
|
||||||
|
冻结一个可复现、与 Figure 5 的 output/gradient 并列叙述对齐的 activation-gradient 定义,
|
||||||
|
再问:
|
||||||
|
|
||||||
|
> 当深度从 16 增至 32、训练预算从 2,000 增至 8,000 step 时,Block AttnRes 是否比
|
||||||
|
> PreNorm Baseline 更一致地降低 post-MLP block-output activation gradient 的跨深度失衡?
|
||||||
|
|
||||||
|
这不是 K3 checkpoint forward,也不是论文 Figure 5 数值复画。
|
||||||
|
|
||||||
|
## 1. 一手来源与不可补写的空白
|
||||||
|
|
||||||
|
| 工件 | 固定 revision / checksum | 用途 |
|
||||||
|
|---|---|---|
|
||||||
|
| Attention Residuals GitHub | `85e22310fe5ee860b4a023de312d791de8a5a5e6` | 公式、Figure 5 / 8、模型尺度 |
|
||||||
|
| `Attention_Residuals.pdf` | SHA-256 `e5831b0d…a26b2f` | 论文一手图文 |
|
||||||
|
| WikiText-2 raw | `Salesforce/wikitext@b08601e04326c79dfdd32d625aee71d232d685c3` | 固定公开训练语料 |
|
||||||
|
| Round 04 protocol | `llm-atlas-k3-attnres-reduced-v1` | 公共主干与数据合同来源 |
|
||||||
|
|
||||||
|
官方仓库没有模型实现、训练脚本、checkpoint 或 Figure 5 原始数据。以下字段不能归因给论文:
|
||||||
|
|
||||||
|
- Figure 5 的确切 gradient tensor;
|
||||||
|
- norm / reduction;
|
||||||
|
- diagnostic batch;
|
||||||
|
- AMP / clipping 时点;
|
||||||
|
- 单点还是时间平均。
|
||||||
|
|
||||||
|
Grok Headless 只进行一次对抗式方法检查;其建议和错误都在前置审计中公开,不是事实来源。
|
||||||
|
|
||||||
|
## 2. 设计总览
|
||||||
|
|
||||||
|
```text
|
||||||
|
2 个深度:16 / 32 Transformer blocks
|
||||||
|
× 2 个 residual graph:PreNorm Baseline / Block AttnRes
|
||||||
|
× 3 个冻结 seed
|
||||||
|
× 8,000 training steps
|
||||||
|
× 32 windows/step
|
||||||
|
× 256 target bytes/window
|
||||||
|
= 12 个正式训练格
|
||||||
|
= 786,432,000 target bytes
|
||||||
|
```
|
||||||
|
|
||||||
|
只比较 Baseline 与 Block,因为论文 Figure 5 的训练动力学面板也是这两个结构的直接对照。
|
||||||
|
Round 04 的 Full AttnRes 结果保持公开,但本轮不增加一个与主问题无关的 6-run 分支。
|
||||||
|
|
||||||
|
## 3. 数据合同
|
||||||
|
|
||||||
|
继承 Round 04 的语料和预处理:
|
||||||
|
|
||||||
|
1. 按固定 parquet 行序读取 `text`;
|
||||||
|
2. 每行追加一个 `\n`;
|
||||||
|
3. UTF-8 编码,无 normalization、strip、去空行或大小写改写;
|
||||||
|
4. byte vocabulary `0..255`;
|
||||||
|
5. 每个窗口连续取 257 bytes,前 256 预测后 256。
|
||||||
|
|
||||||
|
固定拼接后 split:
|
||||||
|
|
||||||
|
| split | bytes | SHA-256 |
|
||||||
|
|---|---:|---|
|
||||||
|
| train | 10,951,563 | `0ca7d3e74dbe44564ea5942b85232f1bbcb525c9cd481cd5d28a87ee90e7e9b4` |
|
||||||
|
| validation | 1,148,008 | `a42356f6a8ff1d25daf25ec9db49e10a537c265581b61c74604bb63231dee719` |
|
||||||
|
| test | 1,292,014 | `bfe9eb16ab9987fb88bde4ea9a30a00f2a45db01dfc14bad78d05325789c4f12` |
|
||||||
|
|
||||||
|
训练第 `step`、第 `row` 的 window 起点:
|
||||||
|
|
||||||
|
```text
|
||||||
|
z = first 8 bytes of SHA256(
|
||||||
|
protocol_id + "\0train-window\0" + seed + "\0" + step + "\0" + row
|
||||||
|
)
|
||||||
|
start = uint64_be(z) mod (len(train_bytes) - 257)
|
||||||
|
```
|
||||||
|
|
||||||
|
同一 seed 的 4 个结构格逐 step / row 使用完全相同的 token tensor。validation 64 windows、
|
||||||
|
diagnostic 16 windows,分别由标签 `validation-window` / `diagnostic-window` 与固定 index
|
||||||
|
生成,对全部结构与 seed 相同。
|
||||||
|
|
||||||
|
正式运行前 manifest 必须记录 parquet hash、split bytes/hash、全部 `3×8,000×32=768,000`
|
||||||
|
唯一训练窗口起点的 schedule hash、validation tensor hash 与 diagnostic tensor hash。
|
||||||
|
|
||||||
|
## 4. 模型合同
|
||||||
|
|
||||||
|
### 4.1 两个深度共享的结构
|
||||||
|
|
||||||
|
| 项 | 固定值 |
|
||||||
|
|---|---:|
|
||||||
|
| vocabulary | 256 bytes |
|
||||||
|
| context | 256 |
|
||||||
|
| `d_model` | 192 |
|
||||||
|
| heads / head dimension | 6 / 32 |
|
||||||
|
| `d_ff` | 768 |
|
||||||
|
| dropout | 0 |
|
||||||
|
| positional embedding | learned absolute,256 × 192 |
|
||||||
|
| norm | RMSNorm,`eps=1e-6` |
|
||||||
|
| attention | causal MHA;score 以 FP32 softmax |
|
||||||
|
| MLP | bias-free SwiGLU,`192→768`, `192→768`, `768→192` |
|
||||||
|
| embedding / readout | tied;final RMSNorm 后乘 token embedding |
|
||||||
|
|
||||||
|
全部 bias-free linear 与 embedding 初始化为 `N(0,0.02)`。attention output projection 与 MLP
|
||||||
|
down projection 的标准差为 `0.02 / sqrt(2×depth)`;普通 RMSNorm 为 1。
|
||||||
|
|
||||||
|
数值精度进一步固定为:attention / MLP 线性分支受 BF16 autocast;embedding、Baseline hidden
|
||||||
|
residual stream 与 Block aggregation partial 都以 FP32 累加。也就是说,Block 每个 BF16
|
||||||
|
branch output 在进入 `partial` 前显式转为 FP32。这样两种结构被捕获的 `h_l` 都是 FP32,
|
||||||
|
不会把 residual accumulator 精度差异混进梯度形状对比。
|
||||||
|
|
||||||
|
### 4.2 深度与 Block AttnRes 聚合
|
||||||
|
|
||||||
|
| Transformer depth | residual sublayers | aggregation groups | sublayers/group | Transformer blocks/group |
|
||||||
|
|---:|---:|---:|---:|---:|
|
||||||
|
| 16 | 32 | 8 | 4 | 2 |
|
||||||
|
| 32 | 64 | 8 | 8 | 4 |
|
||||||
|
|
||||||
|
Baseline 子层为:
|
||||||
|
|
||||||
|
```text
|
||||||
|
h ← h + f(RMSNorm(h))
|
||||||
|
```
|
||||||
|
|
||||||
|
Block AttnRes:
|
||||||
|
|
||||||
|
- embedding 永远是 source 0;
|
||||||
|
- 对已经完成的 aggregation-group sums 做跨组 softmax mixture;
|
||||||
|
- 组内 attention / MLP branch output 累加到 `partial`;
|
||||||
|
- 达到组边界时,保存完整 `partial` 为新 source,再开始下一组;
|
||||||
|
- output mixer 聚合 embedding + 8 个完整 group sums。
|
||||||
|
|
||||||
|
每个子层的 pseudo-query 为 `d_model` 向量,严格 zero-init;每个 source 的 key RMSNorm weight
|
||||||
|
严格 one-init。Baseline / Block 不要求总参数量或 residual-mixer FLOPs 相等,但同 seed /
|
||||||
|
depth 的 token/position embedding、attention、MLP、input norm、final norm 和 tied readout
|
||||||
|
必须逐 tensor SHA-256 exact。
|
||||||
|
|
||||||
|
## 5. 训练合同
|
||||||
|
|
||||||
|
| 项 | 固定值 |
|
||||||
|
|---|---:|
|
||||||
|
| seeds | `2026073001, 2026073002, 2026073003` |
|
||||||
|
| formal steps | 8,000 |
|
||||||
|
| batch | 32 |
|
||||||
|
| context | 256 |
|
||||||
|
| target bytes / run | 65,536,000 |
|
||||||
|
| optimizer | AdamW |
|
||||||
|
| betas / epsilon | `(0.9,0.95)` / `1e-8` |
|
||||||
|
| peak / min LR | `3e-4` / `3e-5` |
|
||||||
|
| warmup | 400 steps,linear |
|
||||||
|
| decay | cosine,step 400→8,000 |
|
||||||
|
| weight decay | `0.1` if `ndim>=2`,否则 `0` |
|
||||||
|
| parameter grad clip | global norm `1.0` |
|
||||||
|
| compute | BF16 autocast;FP32 optimizer state |
|
||||||
|
| compile | off / eager |
|
||||||
|
| device | one RTX 5090 |
|
||||||
|
| deterministic | deterministic algorithms;`CUBLAS_WORKSPACE_CONFIG=:4096:8` |
|
||||||
|
|
||||||
|
验证和诊断都发生在:
|
||||||
|
|
||||||
|
```text
|
||||||
|
step 0, 100, 500, 2,000, 4,000, 8,000
|
||||||
|
```
|
||||||
|
|
||||||
|
验证固定 64 windows,8 windows/eval batch,报告 token-mean CE nats 与
|
||||||
|
`bits_per_byte = CE / ln(2)`。诊断固定 16 windows,一次性输入,不改变 optimizer state。
|
||||||
|
|
||||||
|
计时合同:
|
||||||
|
|
||||||
|
- 前 20 个 training step 不计入;
|
||||||
|
- step 21–8,000 每步前后 CUDA synchronize;
|
||||||
|
- step 20 后 reset peak memory;
|
||||||
|
- 报告 mean / median / p95 step ms、peak allocated / reserved;
|
||||||
|
- timing、wall clock、hostname、GPU temperature 不进入 exact replay 字段。
|
||||||
|
|
||||||
|
## 6. 主梯度对象:精确到代码位置
|
||||||
|
|
||||||
|
### 6.1 `h_l` 定义
|
||||||
|
|
||||||
|
对 `l=1..depth`,统一在第 `l` 个 Transformer block 的 MLP branch 完成后捕获:
|
||||||
|
|
||||||
|
| 架构 | 精确定义 |
|
||||||
|
|---|---|
|
||||||
|
| Baseline | attention residual 与 MLP residual 都完成后的 hidden state |
|
||||||
|
| Block | MLP branch 已加入、本 aggregation partial 可能保存/reset **之前**的 partial |
|
||||||
|
|
||||||
|
所有 `h_l.shape = [16,256,192]`、dtype 为 FP32。实现必须为 diagnostic forward 返回独立
|
||||||
|
引用列表,不允许捕获 reset 后的零张量,不允许把 8 个 aggregation sources 当成 16 / 32 个
|
||||||
|
Transformer outputs。
|
||||||
|
|
||||||
|
### 6.2 diagnostic loss 与 gradient magnitude
|
||||||
|
|
||||||
|
```text
|
||||||
|
model.eval()
|
||||||
|
logits, h[1..depth] = forward(fixed_diagnostic_x, capture=true)
|
||||||
|
L = mean(cross_entropy(logits.float(), fixed_diagnostic_y))
|
||||||
|
backward(L) # 不使用 GradScaler,不执行 optimizer.step
|
||||||
|
m_l = sqrt(mean(float32(h_l.grad)² over batch×time×channel))
|
||||||
|
```
|
||||||
|
|
||||||
|
规则:
|
||||||
|
|
||||||
|
- forward 仍使用与训练一致的 BF16 autocast;
|
||||||
|
- logits 在 FP32 中计算 CE;
|
||||||
|
- loss 对全部 `16×256` targets 做算术平均,无 mask、无 label smoothing;
|
||||||
|
- backward 前 model / optimizer gradients 清零;
|
||||||
|
- activation gradient 在任何 parameter clipping 之前读取;
|
||||||
|
- diagnostic 不消耗训练数据,不进入 optimizer,不改变学习率或模型状态。
|
||||||
|
|
||||||
|
### 6.3 同位置 output magnitude
|
||||||
|
|
||||||
|
同一批 `h_l` 计算:
|
||||||
|
|
||||||
|
```text
|
||||||
|
o_l = sqrt(mean(float32(h_l)² over batch×time×channel))
|
||||||
|
```
|
||||||
|
|
||||||
|
它用于显示 Baseline 单调累积与 Block 组内周期,而不作为主 confirmatory endpoint。
|
||||||
|
|
||||||
|
## 7. 主指标与预注册判据
|
||||||
|
|
||||||
|
每个 diagnostic step 保存完整 `m_1..m_depth`。以下统计由未四舍五入的 float64 数组计算。
|
||||||
|
|
||||||
|
### 7.1 绝对尺度
|
||||||
|
|
||||||
|
```text
|
||||||
|
mean_grad = mean_l(m_l)
|
||||||
|
```
|
||||||
|
|
||||||
|
它防止“曲线更平只是全部梯度趋近于零”被 CV 隐藏。绝对尺度不设置优劣阈值,只公开。
|
||||||
|
|
||||||
|
### 7.2 归一化谱与 CV
|
||||||
|
|
||||||
|
```text
|
||||||
|
n_l = m_l / mean_grad
|
||||||
|
CV = population_std_l(m_l) / mean_grad
|
||||||
|
```
|
||||||
|
|
||||||
|
使用 population standard deviation(`ddof=0`)。网站必须同时显示 `m_l` 与 `n_l`,不能只显示
|
||||||
|
CV 排名。
|
||||||
|
|
||||||
|
### 7.3 前后四分位失衡
|
||||||
|
|
||||||
|
```text
|
||||||
|
q = depth / 4
|
||||||
|
Q_first = mean(m_1 .. m_q)
|
||||||
|
Q_last = mean(m_(depth-q+1) .. m_depth)
|
||||||
|
imbalance = abs(ln(Q_first / Q_last))
|
||||||
|
```
|
||||||
|
|
||||||
|
depth 16 时各取 4 层;depth 32 时各取 8 层。`imbalance=0` 才表示首尾一致;这个定义不会把
|
||||||
|
“早层偏大”和“晚层偏大”错误地都解释为越小越好。原始有符号比 `Q_first/Q_last` 仍公开。
|
||||||
|
|
||||||
|
### 7.4 seed 内配对对比
|
||||||
|
|
||||||
|
只在最终 step 8,000 做 confirmatory verdict:
|
||||||
|
|
||||||
|
```text
|
||||||
|
relative_CV_reduction
|
||||||
|
= (CV_baseline - CV_block) / CV_baseline
|
||||||
|
|
||||||
|
relative_imbalance_reduction
|
||||||
|
= (imbalance_baseline - imbalance_block) / imbalance_baseline
|
||||||
|
```
|
||||||
|
|
||||||
|
若 Baseline imbalance 精确为 0,则该 seed 的 relative imbalance reduction 定义为不可计算,
|
||||||
|
该 depth 自动不能得到“联合支持”;仍公开绝对差。
|
||||||
|
|
||||||
|
对每个 depth 分别判定:
|
||||||
|
|
||||||
|
- **joint directional support at this depth**:三个 seed 的 CV reduction 都 `>0`,其均值
|
||||||
|
`>=20%`;同时三个 seed 的 imbalance reduction 都 `>0`,其均值 `>=20%`。
|
||||||
|
- **joint directional concern at this depth**:三个 seed 的两项 reduction 都 `<0`,且两项
|
||||||
|
平均相对恶化都 `>=20%`。
|
||||||
|
- 其他:**mixed / inconclusive at this depth**。
|
||||||
|
|
||||||
|
总判定:
|
||||||
|
|
||||||
|
- 两个 depth 都 support:**scale-consistent directional support in this operationalization**;
|
||||||
|
- 两个 depth 都 concern:**scale-consistent directional concern in this operationalization**;
|
||||||
|
- 其他:**depth-dependent or inconclusive**。
|
||||||
|
|
||||||
|
不计算 p-value、population CI,不把 3 seeds 称为统计证明。
|
||||||
|
|
||||||
|
## 8. 必须公开的次要指标
|
||||||
|
|
||||||
|
### 8.1 全时间轨迹
|
||||||
|
|
||||||
|
六个预注册时点的以下数据必须全部公开,不能选择“最好看”的 checkpoint:
|
||||||
|
|
||||||
|
- validation BPC;
|
||||||
|
- absolute activation-gradient spectrum;
|
||||||
|
- normalized activation-gradient spectrum;
|
||||||
|
- activation CV;
|
||||||
|
- first/last quartile ratio 与 imbalance;
|
||||||
|
- output RMS spectrum;
|
||||||
|
- mean activation-gradient scale。
|
||||||
|
|
||||||
|
### 8.2 参数梯度
|
||||||
|
|
||||||
|
延续 Round 04 定义:每个 Transformer block 的 attention、MLP 与两个 input norm 的参数梯度
|
||||||
|
拼接后计算:
|
||||||
|
|
||||||
|
```text
|
||||||
|
parameter_grad_rms[l] = sqrt(sum(g²) / total_parameter_elements)
|
||||||
|
```
|
||||||
|
|
||||||
|
不含 embedding、final norm、LM head 与 AttnRes mixer 参数;clip 前读取。报告完整谱、CV 和
|
||||||
|
前后四分位失衡,但它们不进入 Round 05 主判定。
|
||||||
|
|
||||||
|
### 8.3 mixer / 成本
|
||||||
|
|
||||||
|
Block 同报:
|
||||||
|
|
||||||
|
- 各子层 softmax mixture 的 source-depth 分布;
|
||||||
|
- output mixer 分布;
|
||||||
|
- entropy 与 embedding / latest-complete-group mass;
|
||||||
|
- step time、peak allocated/reserved、参数量。
|
||||||
|
|
||||||
|
这些用于解释机制与成本,不改变 confirmatory verdict。
|
||||||
|
|
||||||
|
## 9. 运行前闸门
|
||||||
|
|
||||||
|
### 9.1 数据闸门
|
||||||
|
|
||||||
|
- split bytes/hash 与 Round 04 exact;
|
||||||
|
- 新 protocol 的 768,000-window schedule hash 落盘;
|
||||||
|
- validation / diagnostic tensor hash 落盘;
|
||||||
|
- 同 seed 四结构的至少 step 0 / 1 / 7,999 输入 tensor hash exact。
|
||||||
|
|
||||||
|
### 9.2 公共权重闸门
|
||||||
|
|
||||||
|
每个 depth / seed 的 Baseline 与 Block 公共参数逐 tensor exact;输出:
|
||||||
|
|
||||||
|
- 公共参数 tensor 数;
|
||||||
|
- 公共参数 element 数;
|
||||||
|
- name / shape / dtype / bytes 联合 hash;
|
||||||
|
- value bytes 联合 hash。
|
||||||
|
|
||||||
|
### 9.3 activation-gradient 闸门
|
||||||
|
|
||||||
|
四个 depth×architecture 格都必须通过:
|
||||||
|
|
||||||
|
1. 捕获数量严格等于 depth,shape 均为 `[16,256,192]`;
|
||||||
|
2. 全部 gradient 非 `None`、finite、storage 不别名;
|
||||||
|
3. 同输入把 loss 乘 2 后,每层 `m_l` 比值在 `2±1e-5`;
|
||||||
|
4. loss×2 前后 CV、normalized spectrum、quartile ratio 在 `1e-6` 绝对容差内;
|
||||||
|
5. 全新进程重复 smoke 的冻结字段 exact。
|
||||||
|
|
||||||
|
### 9.4 smoke
|
||||||
|
|
||||||
|
四个格都运行 20 training steps;冻结字段包括:
|
||||||
|
|
||||||
|
- protocol、architecture、depth、seed、device、dtype;
|
||||||
|
- data/schedule/tensor hashes;
|
||||||
|
- public-weight hashes;
|
||||||
|
- step 0 / 20 loss 与 validation;
|
||||||
|
- 全部预注册 diagnostic 数组;
|
||||||
|
- finite / alias / loss-scale checks。
|
||||||
|
|
||||||
|
smoke 不能写入 formal 目录。
|
||||||
|
|
||||||
|
## 10. 正式执行与独立 replay
|
||||||
|
|
||||||
|
12 个 formal grid 必须各自在全新进程中执行。目录键为:
|
||||||
|
|
||||||
|
```text
|
||||||
|
depth-{16|32}/{baseline|block}/seed-{2026073001|2026073002|2026073003}
|
||||||
|
```
|
||||||
|
|
||||||
|
完成后预注册 replay:
|
||||||
|
|
||||||
|
```text
|
||||||
|
depth-32 / block / seed-2026073001
|
||||||
|
```
|
||||||
|
|
||||||
|
replay 再从初始化训练完整 8,000 steps,不加载 formal checkpoint。比较时排除:
|
||||||
|
|
||||||
|
- wall time / step-time samples;
|
||||||
|
- peak memory;
|
||||||
|
- process ID / hostname;
|
||||||
|
- GPU 温度与驱动层瞬时字段;
|
||||||
|
- 文件路径和生成时间。
|
||||||
|
|
||||||
|
必须 exact 的字段:
|
||||||
|
|
||||||
|
- 数据与公共权重 hashes;
|
||||||
|
- 全部 validation CE/BPC;
|
||||||
|
- 全部 activation/output/parameter gradient 数组;
|
||||||
|
- mixer 数组与 summary;
|
||||||
|
- final model-state tensor hash;
|
||||||
|
- optimizer-state tensor hash;
|
||||||
|
- training-loss checkpoint 数组。
|
||||||
|
|
||||||
|
若 replay 不 exact,停止聚合并公开失败,不挑选另一 seed 替代。
|
||||||
|
|
||||||
|
## 11. 公开工件
|
||||||
|
|
||||||
|
正式结果完成后仓库必须包含:
|
||||||
|
|
||||||
|
```text
|
||||||
|
experiments/k3/attnres_gradient/
|
||||||
|
README.md
|
||||||
|
build_dataset.py
|
||||||
|
manifest.json
|
||||||
|
train.py
|
||||||
|
analyze.py
|
||||||
|
results/raw/*.json
|
||||||
|
results/compact.json
|
||||||
|
reproduction.json
|
||||||
|
|
||||||
|
research/
|
||||||
|
K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md
|
||||||
|
K3_ATTNRES_GRADIENT_SCALE_PROTOCOL.md
|
||||||
|
K3_ATTNRES_GRADIENT_SCALE_AUDIT.md
|
||||||
|
```
|
||||||
|
|
||||||
|
网站至少提供五个互相联动的视图:
|
||||||
|
|
||||||
|
1. 论文 Figure 5 的“已知 / 未定义”拆解;
|
||||||
|
2. activation gradient 绝对谱与 normalized spectrum;
|
||||||
|
3. depth 16 / 32、三 seed、六时间点对比;
|
||||||
|
4. output RMS 周期与 Block aggregation boundary;
|
||||||
|
5. activation gradient / parameter gradient 并排,以及判定、成本、哈希和声明边界。
|
||||||
|
|
||||||
|
图中必须能切换到所有负结果;不能只放均值、只放 final 或隐藏某个 seed。
|
||||||
|
|
||||||
|
## 12. 允许与禁止的结论
|
||||||
|
|
||||||
|
若达到支持条件,允许写:
|
||||||
|
|
||||||
|
> 在这个公开定义、两种缩小深度和 8,000-step byte-LM 合同中,Block AttnRes 方向一致地
|
||||||
|
> 降低了 post-MLP output activation gradient 的跨深度 CV 与首尾四分位失衡。
|
||||||
|
|
||||||
|
无论结果怎样,都禁止写:
|
||||||
|
|
||||||
|
- 复现了论文 Figure 5 的数值;
|
||||||
|
- 证明了论文未公开实现采用相同梯度定义;
|
||||||
|
- 证明了 Kimi K3 的真实梯度更健康;
|
||||||
|
- 证明 AttnRes 解决梯度消失、梯度爆炸或训练稳定性的全部问题;
|
||||||
|
- 从 3 seeds 推导总体显著性;
|
||||||
|
- 从 depth 16 / 32 外推至 48B、1T+400B Token 或 K3 2.8T 参数;
|
||||||
|
- 隐藏 activation 与 parameter gradient 方向不一致的结果。
|
||||||
|
|
||||||
|
## 13. 变更纪律
|
||||||
|
|
||||||
|
本文件提交后:
|
||||||
|
|
||||||
|
- 允许修复使实现符合本协议的 bug;
|
||||||
|
- 允许补充日志、注释、可视化和不改变数值的导出;
|
||||||
|
- 不允许看过 formal 结果后修改主指标、阈值、diagnostic step、seed、训练预算或 replay 格;
|
||||||
|
- 任何不得不改变实验合同的事项必须先停止、写入审计、升级 protocol ID,再重新执行全部 grid。
|
||||||
@@ -0,0 +1,231 @@
|
|||||||
|
import { writeFileSync } from "node:fs";
|
||||||
|
|
||||||
|
const cdpPort = process.env.CDP_PORT ?? "9228";
|
||||||
|
const baseUrl = process.env.SITE_URL ?? "http://127.0.0.1:4328";
|
||||||
|
const pages = await fetch(`http://127.0.0.1:${cdpPort}/json/list`).then((response) => response.json());
|
||||||
|
const page = pages.find((entry) => entry.type === "page");
|
||||||
|
if (!page) throw new Error(`CDP ${cdpPort} 没有可用页面`);
|
||||||
|
|
||||||
|
const socket = new WebSocket(page.webSocketDebuggerUrl);
|
||||||
|
await new Promise((resolve, reject) => {
|
||||||
|
socket.addEventListener("open", resolve, { once: true });
|
||||||
|
socket.addEventListener("error", reject, { once: true });
|
||||||
|
});
|
||||||
|
|
||||||
|
let nextId = 0;
|
||||||
|
const pending = new Map();
|
||||||
|
const exceptions = [];
|
||||||
|
socket.addEventListener("message", (event) => {
|
||||||
|
const message = JSON.parse(event.data);
|
||||||
|
if (message.id && pending.has(message.id)) {
|
||||||
|
const { resolve, reject } = pending.get(message.id);
|
||||||
|
pending.delete(message.id);
|
||||||
|
if (message.error) reject(new Error(message.error.message));
|
||||||
|
else resolve(message.result);
|
||||||
|
}
|
||||||
|
if (message.method === "Runtime.exceptionThrown") {
|
||||||
|
exceptions.push(message.params.exceptionDetails.exception?.description ?? message.params.exceptionDetails.text);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
const command = (method, params = {}) => new Promise((resolve, reject) => {
|
||||||
|
const id = ++nextId;
|
||||||
|
pending.set(id, { resolve, reject });
|
||||||
|
socket.send(JSON.stringify({ id, method, params }));
|
||||||
|
});
|
||||||
|
const pause = (milliseconds) => new Promise((resolve) => setTimeout(resolve, milliseconds));
|
||||||
|
const evaluate = async (expression) => {
|
||||||
|
const result = await command("Runtime.evaluate", { expression, returnByValue: true, awaitPromise: true });
|
||||||
|
if (result.exceptionDetails) throw new Error(result.exceptionDetails.exception?.description ?? result.exceptionDetails.text);
|
||||||
|
return result.result.value;
|
||||||
|
};
|
||||||
|
const navigate = async (path) => {
|
||||||
|
await command("Page.navigate", { url: `${baseUrl}${path}` });
|
||||||
|
for (let attempt = 0; attempt < 100; attempt += 1) {
|
||||||
|
await pause(100);
|
||||||
|
if (await evaluate("document.readyState === 'complete'")) return;
|
||||||
|
}
|
||||||
|
throw new Error(`${path} 加载超时`);
|
||||||
|
};
|
||||||
|
const screenshot = async (path) => {
|
||||||
|
const result = await command("Page.captureScreenshot", { format: "png", captureBeyondViewport: false });
|
||||||
|
writeFileSync(path, Buffer.from(result.data, "base64"));
|
||||||
|
};
|
||||||
|
|
||||||
|
await command("Page.enable");
|
||||||
|
await command("Runtime.enable");
|
||||||
|
await command("Emulation.setDeviceMetricsOverride", {
|
||||||
|
width: 1440,
|
||||||
|
height: 1100,
|
||||||
|
deviceScaleFactor: 1,
|
||||||
|
mobile: false,
|
||||||
|
});
|
||||||
|
await navigate("/k3/");
|
||||||
|
|
||||||
|
const desktop = await evaluate(`(() => {
|
||||||
|
const root = document.querySelector("[data-gradient-lab]");
|
||||||
|
root.scrollIntoView({ block: "start", behavior: "instant" });
|
||||||
|
window.scrollBy(0, -78);
|
||||||
|
const text = (selector) => root.querySelector(selector)?.textContent.trim();
|
||||||
|
const panel = () => root.querySelector("[data-gradient-panel]:not([hidden])")?.dataset.gradientPanel;
|
||||||
|
const points = (selector) => root.querySelector(selector)?.getAttribute("points");
|
||||||
|
const setSelect = (selector, value) => {
|
||||||
|
const node = root.querySelector(selector);
|
||||||
|
node.value = value;
|
||||||
|
node.dispatchEvent(new Event("change", { bubbles: true }));
|
||||||
|
};
|
||||||
|
|
||||||
|
const initial = {
|
||||||
|
panel: panel(),
|
||||||
|
tabs: root.querySelectorAll("[data-gradient-tab]").length,
|
||||||
|
panels: root.querySelectorAll("[data-gradient-panel]").length,
|
||||||
|
ledger: root.querySelectorAll(".gradient-ledger article").length,
|
||||||
|
definition: root.textContent.includes("论文作者就是这样算的") &&
|
||||||
|
root.textContent.includes("activation gradient") &&
|
||||||
|
root.textContent.includes("parameter gradient"),
|
||||||
|
};
|
||||||
|
|
||||||
|
root.querySelector('[data-gradient-tab="spectrum"]').click();
|
||||||
|
const spectrumInitial = {
|
||||||
|
panel: panel(),
|
||||||
|
baseCv: text("[data-spectrum-base-cv]"),
|
||||||
|
blockCv: text("[data-spectrum-block-cv]"),
|
||||||
|
baseRatio: text("[data-spectrum-base-ratio]"),
|
||||||
|
blockRatio: text("[data-spectrum-block-ratio]"),
|
||||||
|
baseLine: points('[data-chart-line="baseline"]'),
|
||||||
|
blockLine: points('[data-chart-line="block"]'),
|
||||||
|
boundaries: root.querySelectorAll("[data-chart-groups] .group-boundary").length,
|
||||||
|
};
|
||||||
|
root.querySelector('[data-spectrum-scale="normalized"]').click();
|
||||||
|
const normalized = {
|
||||||
|
title: text("[data-spectrum-title]"),
|
||||||
|
baseLine: points('[data-chart-line="baseline"]'),
|
||||||
|
};
|
||||||
|
root.querySelector('[data-spectrum-depth="16"]').click();
|
||||||
|
const depth16 = {
|
||||||
|
baseCv: text("[data-spectrum-base-cv]"),
|
||||||
|
blockCv: text("[data-spectrum-block-cv]"),
|
||||||
|
pointCount: points('[data-chart-line="baseline"]').split(" ").length,
|
||||||
|
boundaries: root.querySelectorAll("[data-chart-groups] .group-boundary").length,
|
||||||
|
};
|
||||||
|
setSelect("[data-spectrum-seed]", "2026073001");
|
||||||
|
setSelect("[data-spectrum-step]", "2000");
|
||||||
|
const seedStep = {
|
||||||
|
state: text("[data-spectrum-state]"),
|
||||||
|
baseCv: text("[data-spectrum-base-cv]"),
|
||||||
|
blockCv: text("[data-spectrum-block-cv]"),
|
||||||
|
};
|
||||||
|
|
||||||
|
root.querySelector('[data-gradient-tab="timeline"]').click();
|
||||||
|
const timelineInitial = {
|
||||||
|
panel: panel(),
|
||||||
|
title: text("[data-time-title]"),
|
||||||
|
line: points('[data-time-line="block"]'),
|
||||||
|
pointCount: root.querySelectorAll('[data-time-points="block"] circle').length,
|
||||||
|
};
|
||||||
|
root.querySelector('[data-time-metric="imbalance"]').click();
|
||||||
|
const timelineChanged = {
|
||||||
|
title: text("[data-time-title]"),
|
||||||
|
line: points('[data-time-line="block"]'),
|
||||||
|
};
|
||||||
|
|
||||||
|
root.querySelector('[data-gradient-tab="output"]').click();
|
||||||
|
const outputInitial = {
|
||||||
|
panel: panel(),
|
||||||
|
pointCount: points('[data-output-line="block"]').split(" ").length,
|
||||||
|
bars: root.querySelectorAll("[data-output-bars] i").length,
|
||||||
|
boundaries: root.querySelectorAll("[data-output-groups] .group-boundary").length,
|
||||||
|
copy: text("[data-output-copy]"),
|
||||||
|
};
|
||||||
|
root.querySelector('[data-output-depth="16"]').click();
|
||||||
|
const outputDepth16 = {
|
||||||
|
pointCount: points('[data-output-line="block"]').split(" ").length,
|
||||||
|
bars: root.querySelectorAll("[data-output-bars] i").length,
|
||||||
|
copy: text("[data-output-copy]"),
|
||||||
|
};
|
||||||
|
|
||||||
|
root.querySelector('[data-gradient-tab="verdict"]').click();
|
||||||
|
const verdict = {
|
||||||
|
panel: panel(),
|
||||||
|
rows: root.querySelectorAll(".verdict-table tbody tr").length,
|
||||||
|
metrics: root.querySelectorAll(".metric-pairs article").length,
|
||||||
|
costs: root.querySelectorAll(".cost-compare article").length,
|
||||||
|
hashes: root.querySelectorAll(".hash-ledger code").length,
|
||||||
|
exact: root.textContent.includes("model + optimizer exact") &&
|
||||||
|
root.textContent.includes("all frozen fields exact"),
|
||||||
|
mixed: root.textContent.includes("depth-dependent or inconclusive"),
|
||||||
|
};
|
||||||
|
|
||||||
|
const first = root.querySelector('[data-gradient-tab="definition"]');
|
||||||
|
first.focus();
|
||||||
|
first.dispatchEvent(new KeyboardEvent("keydown", { key: "ArrowRight", bubbles: true }));
|
||||||
|
const keyboard = {
|
||||||
|
selected: root.querySelector('[data-gradient-tab][aria-selected="true"]').dataset.gradientTab,
|
||||||
|
panel: panel(),
|
||||||
|
};
|
||||||
|
|
||||||
|
return {
|
||||||
|
initial, spectrumInitial, normalized, depth16, seedStep,
|
||||||
|
timelineInitial, timelineChanged, outputInitial, outputDepth16,
|
||||||
|
verdict, keyboard,
|
||||||
|
documentOverflow: document.documentElement.scrollWidth - document.documentElement.clientWidth,
|
||||||
|
rootOverflow: root.scrollWidth - root.clientWidth,
|
||||||
|
};
|
||||||
|
})()`);
|
||||||
|
await pause(180);
|
||||||
|
await screenshot("/tmp/llm-atlas-k3-attnres-gradient-desktop.png");
|
||||||
|
|
||||||
|
await command("Emulation.setDeviceMetricsOverride", {
|
||||||
|
width: 390,
|
||||||
|
height: 844,
|
||||||
|
deviceScaleFactor: 1,
|
||||||
|
mobile: true,
|
||||||
|
});
|
||||||
|
await navigate("/k3/");
|
||||||
|
const mobile = await evaluate(`(() => {
|
||||||
|
const root = document.querySelector("[data-gradient-lab]");
|
||||||
|
root.scrollIntoView({ block: "start", behavior: "instant" });
|
||||||
|
window.scrollBy(0, -64);
|
||||||
|
root.querySelector('[data-gradient-tab="spectrum"]').click();
|
||||||
|
root.querySelector('[data-spectrum-depth="16"]').click();
|
||||||
|
root.querySelector('[data-gradient-tab="output"]').click();
|
||||||
|
root.querySelector('[data-output-depth="16"]').click();
|
||||||
|
return {
|
||||||
|
tabs: root.querySelectorAll("[data-gradient-tab]").length,
|
||||||
|
ledger: root.querySelectorAll(".gradient-ledger article").length,
|
||||||
|
outputBars: root.querySelectorAll("[data-output-bars] i").length,
|
||||||
|
visiblePanel: root.querySelector("[data-gradient-panel]:not([hidden])")?.dataset.gradientPanel,
|
||||||
|
documentOverflow: document.documentElement.scrollWidth - document.documentElement.clientWidth,
|
||||||
|
rootOverflow: root.scrollWidth - root.clientWidth,
|
||||||
|
};
|
||||||
|
})()`);
|
||||||
|
await pause(180);
|
||||||
|
await screenshot("/tmp/llm-atlas-k3-attnres-gradient-mobile.png");
|
||||||
|
|
||||||
|
const report = { desktop, mobile, exceptions };
|
||||||
|
console.log(JSON.stringify(report, null, 2));
|
||||||
|
|
||||||
|
const numeric = (value) => Number.parseFloat(value.replace("−", "-"));
|
||||||
|
const failures = [];
|
||||||
|
if (desktop.initial.panel !== "definition" || desktop.initial.tabs !== 5 || desktop.initial.panels !== 5 || desktop.initial.ledger !== 6) failures.push("五视图初始结构异常");
|
||||||
|
if (!desktop.initial.definition) failures.push("定义与对象边界缺失");
|
||||||
|
if (desktop.spectrumInitial.panel !== "spectrum" || Math.abs(numeric(desktop.spectrumInitial.baseCv) - 0.3786) > 1e-4 || Math.abs(numeric(desktop.spectrumInitial.blockCv) - 0.5996) > 1e-4) failures.push("depth-32 final spectrum 读数异常");
|
||||||
|
if (desktop.spectrumInitial.boundaries !== 7 || desktop.spectrumInitial.baseLine === desktop.normalized.baseLine || !desktop.normalized.title.includes("NORMALIZED")) failures.push("绝对/归一化谱或组边界异常");
|
||||||
|
if (desktop.depth16.pointCount !== 16 || desktop.depth16.boundaries !== 7 || Math.abs(numeric(desktop.depth16.baseCv) - 0.4246) > 1e-4 || Math.abs(numeric(desktop.depth16.blockCv) - 0.4678) > 1e-4) failures.push("depth-16 spectrum 切换异常");
|
||||||
|
if (!desktop.seedStep.state.includes("2026073001") || !desktop.seedStep.state.includes("2,000") || Math.abs(numeric(desktop.seedStep.baseCv) - 0.4101) > 1e-4 || Math.abs(numeric(desktop.seedStep.blockCv) - 0.3656) > 1e-4) failures.push("seed/checkpoint spectrum 切换异常");
|
||||||
|
if (desktop.timelineInitial.panel !== "timeline" || desktop.timelineInitial.pointCount !== 6 || desktop.timelineInitial.line === desktop.timelineChanged.line || !desktop.timelineChanged.title.includes("FIRST / LAST")) failures.push("六时点轨迹指标切换异常");
|
||||||
|
if (desktop.outputInitial.panel !== "output" || desktop.outputInitial.pointCount !== 32 || desktop.outputInitial.bars !== 32 || desktop.outputInitial.boundaries !== 7 || desktop.outputDepth16.pointCount !== 16 || desktop.outputDepth16.bars !== 16 || !desktop.outputDepth16.copy.includes("DEPTH 16")) failures.push("Output RMS 深度/组节律切换异常");
|
||||||
|
if (desktop.verdict.panel !== "verdict" || desktop.verdict.rows !== 2 || desktop.verdict.metrics !== 4 || desktop.verdict.costs !== 3 || desktop.verdict.hashes !== 4 || !desktop.verdict.exact || !desktop.verdict.mixed) failures.push("联合判定、成本或重放视图异常");
|
||||||
|
if (desktop.keyboard.selected !== "spectrum" || desktop.keyboard.panel !== "spectrum") failures.push("键盘 tab 导航异常");
|
||||||
|
if (desktop.documentOverflow > 1 || desktop.rootOverflow > 1 || mobile.documentOverflow > 1 || mobile.rootOverflow > 1) failures.push("桌面或移动端出现文档级横向溢出");
|
||||||
|
if (mobile.tabs !== 5 || mobile.ledger !== 6 || mobile.outputBars !== 16 || mobile.visiblePanel !== "output") failures.push("移动端交互结构异常");
|
||||||
|
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||||
|
|
||||||
|
if (failures.length) {
|
||||||
|
console.error(`\nFAIL\n- ${failures.join("\n- ")}`);
|
||||||
|
process.exitCode = 1;
|
||||||
|
} else {
|
||||||
|
console.log("\nPASS K3 AttnRes gradient browser regression");
|
||||||
|
}
|
||||||
|
|
||||||
|
socket.close();
|
||||||
@@ -0,0 +1,68 @@
|
|||||||
|
import { createHash } from "node:crypto";
|
||||||
|
import { readdirSync, readFileSync } from "node:fs";
|
||||||
|
|
||||||
|
const read = (path) => {
|
||||||
|
const bytes = readFileSync(new URL(path, import.meta.url));
|
||||||
|
return {
|
||||||
|
bytes,
|
||||||
|
json: JSON.parse(bytes),
|
||||||
|
sha256: createHash("sha256").update(bytes).digest("hex"),
|
||||||
|
};
|
||||||
|
};
|
||||||
|
|
||||||
|
const raw = read("../src/data/k3-attnres-gradient.json");
|
||||||
|
const compact = read("../src/data/k3-attnres-gradient-compact.json");
|
||||||
|
const reproduction = read("../experiments/k3/attnres_gradient/reproduction.json");
|
||||||
|
const rawDirectory = new URL("../experiments/k3/attnres_gradient/results/raw/", import.meta.url);
|
||||||
|
const failures = [];
|
||||||
|
const expect = (condition, message) => {
|
||||||
|
if (!condition) failures.push(message);
|
||||||
|
};
|
||||||
|
|
||||||
|
expect(raw.sha256 === "ad461cbe74fc671f356288d37c9b628618814915116e4fbe03aaaa48367e6a8d", "raw aggregate SHA-256 changed");
|
||||||
|
expect(compact.sha256 === "5377a5e731db2fdc85a0327f05d43f1cd067d3fc34633619d3004f3fc31968e3", "compact payload SHA-256 changed");
|
||||||
|
expect(reproduction.sha256 === "aedcde6accc6eb1c24b122ffb55706ef620062e848d9fe4c5f622b084a4fdea6", "reproduction payload SHA-256 changed");
|
||||||
|
expect(compact.json.protocol_id === "llm-atlas-k3-attnres-gradient-scale-v1", "protocol identity mismatch");
|
||||||
|
expect(compact.json.study.formal_runs === 12, "formal grid is not 12 cells");
|
||||||
|
expect(compact.json.study.formal_target_bytes === 786_432_000, "formal target-byte budget mismatch");
|
||||||
|
expect(compact.json.study.replay_target_bytes === 65_536_000, "replay byte budget mismatch");
|
||||||
|
expect(compact.json.cells.length === 12, "compact cell count mismatch");
|
||||||
|
expect(Object.keys(raw.json.runs).length === 12, "raw formal cell count mismatch");
|
||||||
|
expect(readdirSync(rawDirectory).filter((name) => name.endsWith(".json")).length === 21, "public raw output count is not 21");
|
||||||
|
expect(reproduction.json.replay_exact.exact, "full formal replay is not exact");
|
||||||
|
expect(Object.values(reproduction.json.smoke_exact).every((row) => row.exact && row.gradient_gate.passed), "paired smoke or loss-scale gate failed");
|
||||||
|
|
||||||
|
for (const depth of ["16", "32"]) {
|
||||||
|
const summary = compact.json.depth_summaries[depth];
|
||||||
|
expect(summary.verdict.label === "mixed / inconclusive at this depth", `depth ${depth} verdict changed`);
|
||||||
|
expect(summary.by_seed.length === 3, `depth ${depth} seed count changed`);
|
||||||
|
expect(summary.by_seed.every((row) => row.block_minus_baseline_bpc < 0), `depth ${depth} BPC pairing changed`);
|
||||||
|
expect(summary.by_seed.every((row) => row.relative_cv_reduction < 0), `depth ${depth} CV counterevidence changed`);
|
||||||
|
expect(summary.by_seed.every((row) => row.relative_imbalance_reduction > 0), `depth ${depth} first/last improvement changed`);
|
||||||
|
}
|
||||||
|
|
||||||
|
expect(Math.abs(compact.json.depth_summaries["16"].means.relative_cv_reduction - (-0.10264135379685868)) < 1e-15, "depth-16 CV contrast changed");
|
||||||
|
expect(Math.abs(compact.json.depth_summaries["32"].means.relative_cv_reduction - (-0.600330100169428)) < 1e-15, "depth-32 CV contrast changed");
|
||||||
|
expect(Math.abs(compact.json.depth_summaries["16"].means.relative_imbalance_reduction - 0.6099652484299795) < 1e-15, "depth-16 imbalance contrast changed");
|
||||||
|
expect(Math.abs(compact.json.depth_summaries["32"].means.relative_imbalance_reduction - 0.7200517407719272) < 1e-15, "depth-32 imbalance contrast changed");
|
||||||
|
expect(compact.json.overall_verdict === "depth-dependent or inconclusive", "overall preregistered verdict changed");
|
||||||
|
|
||||||
|
if (failures.length) {
|
||||||
|
console.error(`FAIL K3 AttnRes gradient data\n- ${failures.join("\n- ")}`);
|
||||||
|
process.exit(1);
|
||||||
|
}
|
||||||
|
|
||||||
|
console.log(JSON.stringify({
|
||||||
|
protocol: compact.json.protocol_id,
|
||||||
|
formalRuns: compact.json.study.formal_runs,
|
||||||
|
formalTargetBytes: compact.json.study.formal_target_bytes,
|
||||||
|
depth16: compact.json.depth_summaries["16"].verdict.label,
|
||||||
|
depth32: compact.json.depth_summaries["32"].verdict.label,
|
||||||
|
replayExact: reproduction.json.replay_exact.exact,
|
||||||
|
hashes: {
|
||||||
|
raw: raw.sha256,
|
||||||
|
compact: compact.sha256,
|
||||||
|
reproduction: reproduction.sha256,
|
||||||
|
},
|
||||||
|
}, null, 2));
|
||||||
|
console.log("PASS K3 AttnRes gradient frozen data");
|
||||||
@@ -81,6 +81,8 @@ const overview = await evaluate(`(() => ({
|
|||||||
artifactMismatch: document.querySelector("#artifacts")?.textContent.includes("A_log [128] ≠ expected [96]"),
|
artifactMismatch: document.querySelector("#artifacts")?.textContent.includes("A_log [128] ≠ expected [96]"),
|
||||||
attnresTabs: document.querySelectorAll("[data-attnres-tab]").length,
|
attnresTabs: document.querySelectorAll("[data-attnres-tab]").length,
|
||||||
attnresPanels: document.querySelectorAll("[data-attnres-panel]").length,
|
attnresPanels: document.querySelectorAll("[data-attnres-panel]").length,
|
||||||
|
gradientTabs: document.querySelectorAll("[data-gradient-tab]").length,
|
||||||
|
gradientPanels: document.querySelectorAll("[data-gradient-panel]").length,
|
||||||
nativeVisionCorrected: document.body.textContent.includes("MoonViT‑V2 从头训练") &&
|
nativeVisionCorrected: document.body.textContent.includes("MoonViT‑V2 从头训练") &&
|
||||||
document.body.textContent.includes("同一个 next-token prediction objective"),
|
document.body.textContent.includes("同一个 next-token prediction objective"),
|
||||||
staleVisionClaim: document.body.textContent.includes("先固定语言模型训练视觉组件"),
|
staleVisionClaim: document.body.textContent.includes("先固定语言模型训练视觉组件"),
|
||||||
@@ -290,8 +292,9 @@ const mobile = await evaluate(`(() => {
|
|||||||
artifactTabs: document.querySelectorAll("[data-artifact-tab]").length,
|
artifactTabs: document.querySelectorAll("[data-artifact-tab]").length,
|
||||||
artifactLayers: document.querySelectorAll("[data-layer-cell]").length,
|
artifactLayers: document.querySelectorAll("[data-layer-cell]").length,
|
||||||
attnresTabs: document.querySelectorAll("[data-attnres-tab]").length,
|
attnresTabs: document.querySelectorAll("[data-attnres-tab]").length,
|
||||||
|
gradientTabs: document.querySelectorAll("[data-gradient-tab]").length,
|
||||||
offenders: [...document.querySelectorAll("body *")]
|
offenders: [...document.querySelectorAll("body *")]
|
||||||
.filter((node) => !node.closest(".paper-chain, .spec-table-wrap, .cache-strip, .architecture-explorer, [data-k3-lab], [data-k3-artifact-lab], [data-attnres-lab]"))
|
.filter((node) => !node.closest(".paper-chain, .spec-table-wrap, .cache-strip, .architecture-explorer, [data-k3-lab], [data-k3-artifact-lab], [data-attnres-lab], [data-gradient-lab]"))
|
||||||
.filter((node) => node.getBoundingClientRect().right > document.documentElement.clientWidth + 1)
|
.filter((node) => node.getBoundingClientRect().right > document.documentElement.clientWidth + 1)
|
||||||
.slice(0, 15)
|
.slice(0, 15)
|
||||||
.map((node) => ({
|
.map((node) => ({
|
||||||
@@ -322,12 +325,13 @@ console.log(JSON.stringify(report, null, 2));
|
|||||||
const numeric = (text) => Number.parseFloat(text.replaceAll(",", "").replace("−", "-"));
|
const numeric = (text) => Number.parseFloat(text.replaceAll(",", "").replace("−", "-"));
|
||||||
const failures = [];
|
const failures = [];
|
||||||
if (!overview.title.includes("因果环节")) failures.push("K3 二轮标题异常");
|
if (!overview.title.includes("因果环节")) failures.push("K3 二轮标题异常");
|
||||||
if (overview.sections !== 33 || overview.tocLinks !== 33) failures.push("32 个编号专题加阅读链的目录结构异常");
|
if (overview.sections !== 34 || overview.tocLinks !== 34) failures.push("33 个编号专题加阅读链的目录结构异常");
|
||||||
if (overview.ledgers !== 32 || overview.reportMap !== 9) failures.push("32 张问题账或报告地图异常");
|
if (overview.ledgers !== 32 || overview.reportMap !== 9) failures.push("32 张问题账或报告地图异常");
|
||||||
if (overview.figureAtlas !== 21 || overview.paperLinks !== 100 || overview.paperGroups < 12) failures.push("图表审计或 100 节点阅读链异常");
|
if (overview.figureAtlas !== 21 || overview.paperLinks !== 100 || overview.paperGroups < 12) failures.push("图表审计或 100 节点阅读链异常");
|
||||||
if (overview.labTabs !== 8 || overview.labPanels !== 8) failures.push("八联实验结构异常");
|
if (overview.labTabs !== 8 || overview.labPanels !== 8) failures.push("八联实验结构异常");
|
||||||
if (overview.artifactTabs !== 4 || overview.artifactPanels !== 4 || overview.artifactLayers !== 93 || !overview.artifactMismatch) failures.push("开放工件四视图、93 层条带或形状冲突异常");
|
if (overview.artifactTabs !== 4 || overview.artifactPanels !== 4 || overview.artifactLayers !== 93 || !overview.artifactMismatch) failures.push("开放工件四视图、93 层条带或形状冲突异常");
|
||||||
if (overview.attnresTabs !== 5 || overview.attnresPanels !== 5) failures.push("AttnRes 独立实验五视图异常");
|
if (overview.attnresTabs !== 5 || overview.attnresPanels !== 5) failures.push("AttnRes 独立实验五视图异常");
|
||||||
|
if (overview.gradientTabs !== 5 || overview.gradientPanels !== 5) failures.push("AttnRes 梯度定义扩展五视图异常");
|
||||||
if (!overview.nativeVisionCorrected || overview.staleVisionClaim) failures.push("原生多模态纠错未生效或旧错误残留");
|
if (!overview.nativeVisionCorrected || overview.staleVisionClaim) failures.push("原生多模态纠错未生效或旧错误残留");
|
||||||
if (overview.documentOverflow > 1 || mobile.documentOverflow > 1) failures.push("桌面或移动端存在文档级横向溢出");
|
if (overview.documentOverflow > 1 || mobile.documentOverflow > 1) failures.push("桌面或移动端存在文档级横向溢出");
|
||||||
if (labs.memoryInitial.panel !== "memory" || numeric(labs.memoryInitial.additiveError) <= numeric(labs.memoryInitial.deltaError)) failures.push("Delta memory 初始递推异常");
|
if (labs.memoryInitial.panel !== "memory" || numeric(labs.memoryInitial.additiveError) <= numeric(labs.memoryInitial.deltaError)) failures.push("Delta memory 初始递推异常");
|
||||||
@@ -352,7 +356,7 @@ if (artifacts.parameterChanged.shape !== "[96,128] F32" || !artifacts.parameterC
|
|||||||
if (artifacts.reproductionInitial.panel !== "reproduction" || numeric(artifacts.reproductionInitial.speedup) !== 1.85 || numeric(artifacts.reproductionInitial.localMean) < 2.6 || !artifacts.reproductionInitial.exactSuite || numeric(artifacts.reproductionInitial.cv) < 2) failures.push("FlashKDA H20、本机 exact suite 或 router 初始探针异常");
|
if (artifacts.reproductionInitial.panel !== "reproduction" || numeric(artifacts.reproductionInitial.speedup) !== 1.85 || numeric(artifacts.reproductionInitial.localMean) < 2.6 || !artifacts.reproductionInitial.exactSuite || numeric(artifacts.reproductionInitial.cv) < 2) failures.push("FlashKDA H20、本机 exact suite 或 router 初始探针异常");
|
||||||
if (numeric(artifacts.reproductionChanged.speedup) !== 3.27 || numeric(artifacts.reproductionChanged.flash) !== 0.7064 || numeric(artifacts.reproductionChanged.localMean) >= numeric(artifacts.reproductionInitial.localMean) || !artifacts.reproductionChanged.localMode.includes("FP32 state") || numeric(artifacts.reproductionChanged.cv) <= numeric(artifacts.reproductionInitial.cv) || numeric(artifacts.reproductionChanged.zero) <= numeric(artifacts.reproductionInitial.zero)) failures.push("GB200 benchmark、本机 varlen/state 或 synthetic router counterexample 未更新");
|
if (numeric(artifacts.reproductionChanged.speedup) !== 3.27 || numeric(artifacts.reproductionChanged.flash) !== 0.7064 || numeric(artifacts.reproductionChanged.localMean) >= numeric(artifacts.reproductionInitial.localMean) || !artifacts.reproductionChanged.localMode.includes("FP32 state") || numeric(artifacts.reproductionChanged.cv) <= numeric(artifacts.reproductionInitial.cv) || numeric(artifacts.reproductionChanged.zero) <= numeric(artifacts.reproductionInitial.zero)) failures.push("GB200 benchmark、本机 varlen/state 或 synthetic router counterexample 未更新");
|
||||||
if (artifacts.keyboardSelected !== "tensors" || artifacts.keyboardVisible !== "tensors") failures.push("开放工件键盘 tab 导航异常");
|
if (artifacts.keyboardSelected !== "tensors" || artifacts.keyboardVisible !== "tensors") failures.push("开放工件键盘 tab 导航异常");
|
||||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 8 || mobile.artifactTabs !== 4 || mobile.artifactLayers !== 93 || mobile.attnresTabs !== 5) failures.push("移动端导航或实验异常");
|
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 8 || mobile.artifactTabs !== 4 || mobile.artifactLayers !== 93 || mobile.attnresTabs !== 5 || mobile.gradientTabs !== 5) failures.push("移动端导航或实验异常");
|
||||||
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
||||||
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,813 @@
|
|||||||
|
---
|
||||||
|
import rawLab from "@/data/k3-attnres-gradient-compact.json";
|
||||||
|
|
||||||
|
const lab = rawLab as any;
|
||||||
|
const json = JSON.stringify(lab).replaceAll("<", "\\u003c");
|
||||||
|
const depth16 = lab.depth_summaries["16"];
|
||||||
|
const depth32 = lab.depth_summaries["32"];
|
||||||
|
const formatSigned = (value: number, digits = 4) =>
|
||||||
|
`${value < 0 ? "−" : value > 0 ? "+" : ""}${Math.abs(value).toFixed(digits)}`;
|
||||||
|
const pct = (value: number, digits = 1) => `${(value * 100).toFixed(digits)}%`;
|
||||||
|
const shortHash = (value: string) => `${value.slice(0, 10)}…${value.slice(-8)}`;
|
||||||
|
---
|
||||||
|
|
||||||
|
<figure class="gradient-lab" data-gradient-lab>
|
||||||
|
<header class="gradient-head">
|
||||||
|
<div>
|
||||||
|
<p>ROUND 05 / GRADIENT DEFINITION × DEPTH SCALE</p>
|
||||||
|
<h3>“早层不再过大”与“整条谱更均匀”不是同一件事</h3>
|
||||||
|
</div>
|
||||||
|
<p>
|
||||||
|
16 / 32 Transformer blocks · Baseline / Block · 3 seeds · 8,000 steps。
|
||||||
|
被测对象是 post-MLP output activation gradient,不冒充论文未公开的 Figure 5 实现。
|
||||||
|
</p>
|
||||||
|
</header>
|
||||||
|
|
||||||
|
<div class="gradient-ledger">
|
||||||
|
<article><span>FORMAL GRID</span><b>2 × 2 × 3</b><p>12 个独立训练格</p></article>
|
||||||
|
<article><span>TARGET BYTES</span><b>786.432M</b><p>每格 65,536,000</p></article>
|
||||||
|
<article><span>DEPTH</span><b>16 → 32</b><p>宽度固定 192</p></article>
|
||||||
|
<article class="split"><span>CV</span><b>6 / 6 更差</b><p>中后段局部尖峰</p></article>
|
||||||
|
<article class="pass"><span>FIRST ↔ LAST</span><b>6 / 6 更近</b><p>失衡改善 56%–81%</p></article>
|
||||||
|
<article class="pass"><span>FULL REPLAY</span><b>exact</b><p>model + optimizer state</p></article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="gradient-tabs" role="tablist" aria-label="选择 AttnRes 梯度定义与深度实验视图">
|
||||||
|
<button type="button" role="tab" data-gradient-tab="definition" aria-selected="true">
|
||||||
|
<span>01</span><b>论文到底定义了什么</b><small>known · underdefined · operationalization</small>
|
||||||
|
</button>
|
||||||
|
<button type="button" role="tab" data-gradient-tab="spectrum" aria-selected="false">
|
||||||
|
<span>02</span><b>绝对谱与归一化谱</b><small>depth · seed · checkpoint</small>
|
||||||
|
</button>
|
||||||
|
<button type="button" role="tab" data-gradient-tab="timeline" aria-selected="false">
|
||||||
|
<span>03</span><b>六个时点怎样演化</b><small>CV · imbalance · mean scale</small>
|
||||||
|
</button>
|
||||||
|
<button type="button" role="tab" data-gradient-tab="output" aria-selected="false">
|
||||||
|
<span>04</span><b>Output RMS 与块节律</b><small>growth · reset · group boundary</small>
|
||||||
|
</button>
|
||||||
|
<button type="button" role="tab" data-gradient-tab="verdict" aria-selected="false">
|
||||||
|
<span>05</span><b>联合判定、成本与重放</b><small>activation ≠ parameter</small>
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<section class="gradient-panel" data-gradient-panel="definition">
|
||||||
|
<div class="panel-lead">
|
||||||
|
<div><span>I / DEFINITION AUDIT</span><h4>论文画出一条“梯度曲线”,却没有给出足以唯一重算的测量合同</h4></div>
|
||||||
|
<p>
|
||||||
|
Figure 5(c) 只写 “Each transformer block’s gradient magnitude”。
|
||||||
|
官方仓库没有训练代码、checkpoint、统计脚本或原始数组。
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="known-grid">
|
||||||
|
<article class="known">
|
||||||
|
<span>OFFICIAL / KNOWN</span>
|
||||||
|
<h5>图与文字能确认</h5>
|
||||||
|
<ul>
|
||||||
|
<li>横轴是 Transformer block index。</li>
|
||||||
|
<li>Figure 5(b) 同时画 block output magnitude。</li>
|
||||||
|
<li>正文说 Baseline 早层梯度过大,Block 更均匀。</li>
|
||||||
|
<li>最终模型约 27 blocks / 54 residual layers。</li>
|
||||||
|
</ul>
|
||||||
|
</article>
|
||||||
|
<article class="unknown">
|
||||||
|
<span>OFFICIAL / UNDERDEFINED</span>
|
||||||
|
<h5>无法从公开工件唯一恢复</h5>
|
||||||
|
<ul>
|
||||||
|
<li>activation、branch 还是 parameter gradient?</li>
|
||||||
|
<li>L2、RMS、mean absolute 还是别的 norm?</li>
|
||||||
|
<li>batch / token / channel 怎样 reduction?</li>
|
||||||
|
<li>哪个 checkpoint、AMP / clip 前还是后?</li>
|
||||||
|
</ul>
|
||||||
|
</article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="object-chain" aria-label="Round 05 activation gradient 测量对象">
|
||||||
|
<div><span>FIXED INPUT</span><b>16 × 256 bytes</b><p>同一 diagnostic tensor</p></div>
|
||||||
|
<i>→</i>
|
||||||
|
<div><span>POST-MLP OUTPUT</span><b>h₁ … h<sub>L</sub></b><p>每个 Transformer block 一个</p></div>
|
||||||
|
<i>→</i>
|
||||||
|
<div><span>TOKEN-MEAN CE</span><b>ℒ</b><p>FP32 cross entropy</p></div>
|
||||||
|
<i>→</i>
|
||||||
|
<div class="accent"><span>MEASURE</span><b>RMS(∂ℒ/∂h<sub>l</sub>)</b><p>B × T × C 联合 RMS</p></div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="object-compare">
|
||||||
|
<article><span>ROUND 04</span><b>∇<sub>θl</sub>ℒ</b><p>核心参数梯度:权重收到多大更新信号。</p></article>
|
||||||
|
<i>≠</i>
|
||||||
|
<article><span>ROUND 05</span><b>∂ℒ/∂h<sub>l</sub></b><p>activation gradient:损失对这一深度表示有多敏感。</p></article>
|
||||||
|
<p>两种对象都公开;新指标不会覆盖上一轮反结果。</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="definition-boundary">
|
||||||
|
<b>能说</b><p>“这是与 Figure 5 叙述对齐的一种公开 operationalization。”</p>
|
||||||
|
<b>不能说</b><p>“论文作者就是这样算的,或本站复画了 Figure 5(c)。”</p>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="gradient-panel" data-gradient-panel="spectrum" hidden>
|
||||||
|
<div class="panel-lead">
|
||||||
|
<div><span>II / ALIGNED DEPTH SPECTRUM</span><h4>同一条梯度谱,绝对值与归一化形状要一起看</h4></div>
|
||||||
|
<p>竖线是 8 个 AttnRes aggregation groups 的边界;Baseline 也画同位置,方便逐层配对。</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="lab-controls">
|
||||||
|
<div role="group" aria-label="选择深度">
|
||||||
|
<button type="button" data-spectrum-depth="16" aria-pressed="false">DEPTH 16</button>
|
||||||
|
<button type="button" data-spectrum-depth="32" aria-pressed="true">DEPTH 32</button>
|
||||||
|
</div>
|
||||||
|
<label>SEED
|
||||||
|
<select data-spectrum-seed>
|
||||||
|
<option value="mean">3-SEED MEAN</option>
|
||||||
|
{lab.study.seeds.map((seed: number) => <option value={String(seed)}>{seed}</option>)}
|
||||||
|
</select>
|
||||||
|
</label>
|
||||||
|
<label>CHECKPOINT
|
||||||
|
<select data-spectrum-step>
|
||||||
|
{lab.study.diagnostic_steps.map((step: number) => <option value={String(step)} selected={step === 8000}>STEP {step.toLocaleString("en-US")}</option>)}
|
||||||
|
</select>
|
||||||
|
</label>
|
||||||
|
<div role="group" aria-label="选择绝对或归一化梯度谱">
|
||||||
|
<button type="button" data-spectrum-scale="absolute" aria-pressed="true">ABSOLUTE RMS</button>
|
||||||
|
<button type="button" data-spectrum-scale="normalized" aria-pressed="false">÷ LAYER MEAN</button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="spectrum-layout">
|
||||||
|
<div class="chart-shell">
|
||||||
|
<header><span data-spectrum-title>ACTIVATION GRADIENT RMS</span><b>POST-MLP BLOCK OUTPUT</b></header>
|
||||||
|
<svg data-spectrum-chart viewBox="0 0 920 360" role="img" aria-label="Baseline 与 Block 的逐层 activation gradient 谱">
|
||||||
|
<g data-chart-groups></g>
|
||||||
|
<g data-chart-grid></g>
|
||||||
|
<polyline data-chart-line="baseline" class="series baseline"></polyline>
|
||||||
|
<polyline data-chart-line="block" class="series block"></polyline>
|
||||||
|
<g data-chart-points="baseline"></g>
|
||||||
|
<g data-chart-points="block"></g>
|
||||||
|
<text x="460" y="350" class="axis-title">TRANSFORMER BLOCK INDEX</text>
|
||||||
|
</svg>
|
||||||
|
<div class="chart-legend">
|
||||||
|
<span><i class="baseline"></i>Baseline</span>
|
||||||
|
<span><i class="block"></i>Block AttnRes</span>
|
||||||
|
<span><i class="boundary"></i>aggregation boundary</span>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<div class="spectrum-readout">
|
||||||
|
<span data-spectrum-state>DEPTH 32 · 3-SEED MEAN · STEP 8,000</span>
|
||||||
|
<article><b>POPULATION CV</b><div><span>BASE</span><strong data-spectrum-base-cv>0.3786</strong></div><div><span>BLOCK</span><strong data-spectrum-block-cv>0.5996</strong></div></article>
|
||||||
|
<article><b>FIRST / LAST QUARTILE</b><div><span>BASE</span><strong data-spectrum-base-ratio>3.21×</strong></div><div><span>BLOCK</span><strong data-spectrum-block-ratio>0.95×</strong></div></article>
|
||||||
|
<article><b>MEAN ABSOLUTE SCALE</b><div><span>BASE</span><strong data-spectrum-base-mean>1.44e−4</strong></div><div><span>BLOCK</span><strong data-spectrum-block-mean>0.78e−4</strong></div></article>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="split-result">
|
||||||
|
<article><span>FIRST ↔ LAST</span><b>更接近</b><p>Baseline 的早层整体隆起被削弱。</p></article>
|
||||||
|
<i>但</i>
|
||||||
|
<article><span>ALL-LAYER CV</span><b>反而更高</b><p>中后段少数位置形成更尖的峰。</p></article>
|
||||||
|
<i>所以</i>
|
||||||
|
<article class="accent"><span>PRE-REGISTERED</span><b>mixed</b><p>“更均匀”必须拆成至少两个指标。</p></article>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="gradient-panel" data-gradient-panel="timeline" hidden>
|
||||||
|
<div class="panel-lead">
|
||||||
|
<div><span>III / SIX FROZEN CHECKPOINTS</span><h4>Block 的中期优势会反转;不能挑一个 checkpoint 讲故事</h4></div>
|
||||||
|
<p>横轴按六个预注册诊断时点等距排列;标签保留真实 step,不暗示实际时间等距。</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="lab-controls">
|
||||||
|
<div role="group" aria-label="选择时间轨迹深度">
|
||||||
|
<button type="button" data-time-depth="16" aria-pressed="false">DEPTH 16</button>
|
||||||
|
<button type="button" data-time-depth="32" aria-pressed="true">DEPTH 32</button>
|
||||||
|
</div>
|
||||||
|
<label>SEED
|
||||||
|
<select data-time-seed>
|
||||||
|
<option value="mean">3-SEED MEAN</option>
|
||||||
|
{lab.study.seeds.map((seed: number) => <option value={String(seed)}>{seed}</option>)}
|
||||||
|
</select>
|
||||||
|
</label>
|
||||||
|
<div role="group" aria-label="选择时间轨迹指标">
|
||||||
|
<button type="button" data-time-metric="cv" aria-pressed="true">CV</button>
|
||||||
|
<button type="button" data-time-metric="imbalance" aria-pressed="false">FIRST/LAST IMBALANCE</button>
|
||||||
|
<button type="button" data-time-metric="mean" aria-pressed="false">ABSOLUTE MEAN</button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="chart-shell timeline-chart">
|
||||||
|
<header><span data-time-title>POPULATION CV</span><b data-time-copy>DEPTH 32 · 3-SEED MEAN</b></header>
|
||||||
|
<svg data-time-chart viewBox="0 0 920 360" role="img" aria-label="六个固定训练时点的 activation gradient 指标轨迹">
|
||||||
|
<g data-time-grid></g>
|
||||||
|
<polyline data-time-line="baseline" class="series baseline"></polyline>
|
||||||
|
<polyline data-time-line="block" class="series block"></polyline>
|
||||||
|
<g data-time-points="baseline"></g>
|
||||||
|
<g data-time-points="block"></g>
|
||||||
|
<text x="460" y="350" class="axis-title">PREREGISTERED DIAGNOSTIC STEP</text>
|
||||||
|
</svg>
|
||||||
|
<div class="chart-legend"><span><i class="baseline"></i>Baseline</span><span><i class="block"></i>Block AttnRes</span></div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="timeline-notes">
|
||||||
|
<article><span>SEED 01 / DEPTH 32</span><b>0.344 → 0.538</b><p>step 2,000:Block 已显著更尖。</p></article>
|
||||||
|
<article><span>SEED 02 / DEPTH 32</span><b>0.352 → 0.618</b><p>同一时点复现恶化方向。</p></article>
|
||||||
|
<article class="counter"><span>SEED 03 / DEPTH 32</span><b>0.400 → 0.349</b><p>中期反例:Block 此时反而更平。</p></article>
|
||||||
|
<article><span>SEED 03 / FINAL</span><b>0.403 → 0.427</b><p>到 8,000 step 才轻微反转。</p></article>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="gradient-panel" data-gradient-panel="output" hidden>
|
||||||
|
<div class="panel-lead">
|
||||||
|
<div><span>IV / OUTPUT MAGNITUDE</span><h4>梯度结论 mixed,不代表论文所有训练动力学叙述都没有出现</h4></div>
|
||||||
|
<p>同一个 post-MLP 位置计算 output RMS;这里不做 backward,也不改变主判定。</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="lab-controls">
|
||||||
|
<div role="group" aria-label="选择 output RMS 深度">
|
||||||
|
<button type="button" data-output-depth="16" aria-pressed="false">DEPTH 16</button>
|
||||||
|
<button type="button" data-output-depth="32" aria-pressed="true">DEPTH 32</button>
|
||||||
|
</div>
|
||||||
|
<label>SEED
|
||||||
|
<select data-output-seed>
|
||||||
|
<option value="mean">3-SEED MEAN</option>
|
||||||
|
{lab.study.seeds.map((seed: number) => <option value={String(seed)}>{seed}</option>)}
|
||||||
|
</select>
|
||||||
|
</label>
|
||||||
|
<label>CHECKPOINT
|
||||||
|
<select data-output-step>
|
||||||
|
{lab.study.diagnostic_steps.map((step: number) => <option value={String(step)} selected={step === 8000}>STEP {step.toLocaleString("en-US")}</option>)}
|
||||||
|
</select>
|
||||||
|
</label>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="chart-shell">
|
||||||
|
<header><span>POST-MLP OUTPUT RMS</span><b data-output-copy>DEPTH 32 · 3-SEED MEAN · STEP 8,000</b></header>
|
||||||
|
<svg data-output-chart viewBox="0 0 920 360" role="img" aria-label="Baseline 与 Block 的逐层 output RMS">
|
||||||
|
<g data-output-groups></g>
|
||||||
|
<g data-output-grid></g>
|
||||||
|
<polyline data-output-line="baseline" class="series baseline"></polyline>
|
||||||
|
<polyline data-output-line="block" class="series block"></polyline>
|
||||||
|
<g data-output-points="baseline"></g>
|
||||||
|
<g data-output-points="block"></g>
|
||||||
|
<text x="460" y="350" class="axis-title">TRANSFORMER BLOCK INDEX</text>
|
||||||
|
</svg>
|
||||||
|
<div class="chart-legend"><span><i class="baseline"></i>Baseline</span><span><i class="block"></i>Block AttnRes</span><span><i class="boundary"></i>aggregation boundary</span></div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="output-ratios">
|
||||||
|
<article><span>DEPTH 16 / LAST ÷ FIRST</span><div><b>Baseline</b><strong>4.59×</strong></div><div><b>Block</b><strong>1.17×</strong></div></article>
|
||||||
|
<article><span>DEPTH 32 / LAST ÷ FIRST</span><div><b>Baseline</b><strong>6.08×</strong></div><div><b>Block</b><strong>1.89×</strong></div></article>
|
||||||
|
<article class="accent"><span>WHAT THE CURVE SAYS</span><b>全局累积 → 组内锯齿</b><p>Block 限制 output magnitude 持续跨深度增长;这一方向与 Figure 5(b) 叙述一致。</p></article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="group-rhythm">
|
||||||
|
<header><span data-rhythm-title>BLOCK ATTNRES · DEPTH 32</span><b>8 AGGREGATION GROUPS</b></header>
|
||||||
|
<div data-output-bars aria-label="Block AttnRes output RMS 组内节律"></div>
|
||||||
|
<p>粗分隔线是 group boundary;柱高来自当前选择的 checkpoint / seed,不是示意动画。</p>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="gradient-panel" data-gradient-panel="verdict" hidden>
|
||||||
|
<div class="panel-lead">
|
||||||
|
<div><span>V / JOINT VERDICT</span><h4>一个正结果、两个反结果和一张成本账,要同时摆在桌面上</h4></div>
|
||||||
|
<p>Round 05 的主判定只读 activation CV + imbalance;BPC、参数梯度与成本是必须公开的次要结果。</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="verdict-table-wrap">
|
||||||
|
<table class="verdict-table">
|
||||||
|
<thead><tr><th>Depth</th><th>Δ BPC</th><th>Activation CV</th><th>First/last imbalance</th><th>Mean grad scale</th><th>Parameter CV</th><th>Verdict</th></tr></thead>
|
||||||
|
<tbody>
|
||||||
|
<tr>
|
||||||
|
<th>16</th>
|
||||||
|
<td>{formatSigned(depth16.means.block_minus_baseline_bpc, 5)}</td>
|
||||||
|
<td class="bad">{pct(depth16.means.relative_cv_reduction)} reduction</td>
|
||||||
|
<td class="good">+{pct(depth16.means.relative_imbalance_reduction)}</td>
|
||||||
|
<td>{pct(depth16.means.block_to_baseline_activation_grad_mean)} of Base</td>
|
||||||
|
<td>{depth16.means.baseline_parameter_grad_cv.toFixed(3)} → {depth16.means.block_parameter_grad_cv.toFixed(3)}</td>
|
||||||
|
<td>mixed</td>
|
||||||
|
</tr>
|
||||||
|
<tr>
|
||||||
|
<th>32</th>
|
||||||
|
<td>{formatSigned(depth32.means.block_minus_baseline_bpc, 5)}</td>
|
||||||
|
<td class="bad">{pct(depth32.means.relative_cv_reduction)} reduction</td>
|
||||||
|
<td class="good">+{pct(depth32.means.relative_imbalance_reduction)}</td>
|
||||||
|
<td>{pct(depth32.means.block_to_baseline_activation_grad_mean)} of Base</td>
|
||||||
|
<td>{depth32.means.baseline_parameter_grad_cv.toFixed(3)} → {depth32.means.block_parameter_grad_cv.toFixed(3)}</td>
|
||||||
|
<td>mixed</td>
|
||||||
|
</tr>
|
||||||
|
</tbody>
|
||||||
|
</table>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="metric-pairs">
|
||||||
|
<article><span>ACTIVATION / FIRST-LAST</span><b>Block 改善</b><p>6 / 6 配对更接近 1。</p></article>
|
||||||
|
<article class="warn"><span>ACTIVATION / ALL-LAYER CV</span><b>Block 恶化</b><p>6 / 6 配对 CV 更高。</p></article>
|
||||||
|
<article class="warn"><span>PARAMETER / ALL-LAYER CV</span><b>Block 恶化</b><p>0.416→0.683;0.397→0.772。</p></article>
|
||||||
|
<article class="pass"><span>VALIDATION BPC</span><b>Block 更低</b><p>6 / 6 配对为负;非同算力。</p></article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="cost-compare">
|
||||||
|
<article><span>DEPTH 16</span><div><b>STEP TIME</b><strong>21.47 → 54.71 ms</strong></div><div><b>PEAK ALLOC</b><strong>3.04 → 6.54 GB</strong></div><p>约 2.55× time · 2.15× memory</p></article>
|
||||||
|
<article><span>DEPTH 32</span><div><b>STEP TIME</b><strong>42.11 → 109.38 ms</strong></div><div><b>PEAK ALLOC</b><strong>5.94 → 12.82 GB</strong></div><p>约 2.60× time · 2.16× memory</p></article>
|
||||||
|
<article class="boundary"><span>CLAIM BOUNDARY</span><b>同 token / step</b><p>不能写成同 FLOPs、同 wall time,或 K3 生产成本。</p></article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="replay-ledger">
|
||||||
|
<article><span>FORMAL</span><b>depth-32 · Block · seed-01</b><p>8,000 steps from initialization</p></article>
|
||||||
|
<i>≡</i>
|
||||||
|
<article><span>FRESH REPLAY</span><b>all frozen fields exact</b><p>额外 65,536,000 target bytes</p></article>
|
||||||
|
<i>→</i>
|
||||||
|
<article class="pass"><span>STATE HASH</span><b>model + optimizer exact</b><p>{shortHash(lab.cells.find((cell: any) => cell.depth === 32 && cell.architecture === "block" && cell.seed === 2026073001).hashes.final_model_state)}</p></article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="hash-ledger">
|
||||||
|
<article><span>MANIFEST</span><code>{lab.manifest_summary.file_sha256}</code></article>
|
||||||
|
<article><span>SCHEDULE</span><code>{lab.manifest_summary.formal_schedule_sha256}</code></article>
|
||||||
|
<article><span>COMPACT CANONICAL</span><code>{lab.canonical_sha256_without_self}</code></article>
|
||||||
|
<article><span>OVERALL</span><code>{lab.overall_verdict}</code></article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="claim-grid">
|
||||||
|
<article class="yes"><span>THIS STUDY SUPPORTS</span><ul><li>在公开定义下,Block 改写了梯度失衡的形态。</li><li>早/晚深度更接近,但中后段局部峰更尖。</li><li>Output RMS 的跨深度增长显著受限。</li></ul></article>
|
||||||
|
<article class="no"><span>THIS STUDY DOES NOT SUPPORT</span><ul><li>复现论文 Figure 5(c) 的数值或隐藏实现。</li><li>测到 K3 checkpoint 的真实梯度。</li><li>证明 AttnRes 一般更稳定或更省算力。</li></ul></article>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<script is:inline type="application/json" data-gradient-payload set:html={json}></script>
|
||||||
|
</figure>
|
||||||
|
|
||||||
|
<script>
|
||||||
|
const initializeGradientLab = (root: HTMLElement) => {
|
||||||
|
if (root.dataset.gradientReady === "true") return;
|
||||||
|
root.dataset.gradientReady = "true";
|
||||||
|
const payload = root.querySelector<HTMLScriptElement>("[data-gradient-payload]");
|
||||||
|
if (!payload?.textContent) return;
|
||||||
|
const data = JSON.parse(payload.textContent);
|
||||||
|
const architectures = ["baseline", "block"];
|
||||||
|
const steps = data.study.diagnostic_steps;
|
||||||
|
const seeds = data.study.seeds;
|
||||||
|
const ns = "http://www.w3.org/2000/svg";
|
||||||
|
const colors: Record<string, string> = { baseline: "#77746b", block: "#ba603b" };
|
||||||
|
|
||||||
|
const tabs = [...root.querySelectorAll<HTMLButtonElement>("[data-gradient-tab]")];
|
||||||
|
const panels = [...root.querySelectorAll<HTMLElement>("[data-gradient-panel]")];
|
||||||
|
const selectTab = (id: string) => {
|
||||||
|
tabs.forEach((tab) => tab.setAttribute("aria-selected", String(tab.dataset.gradientTab === id)));
|
||||||
|
panels.forEach((panel) => { panel.hidden = panel.dataset.gradientPanel !== id; });
|
||||||
|
};
|
||||||
|
tabs.forEach((tab, index) => {
|
||||||
|
tab.addEventListener("click", () => selectTab(tab.dataset.gradientTab || "definition"));
|
||||||
|
tab.addEventListener("keydown", (event) => {
|
||||||
|
if (!["ArrowLeft", "ArrowRight", "Home", "End"].includes(event.key)) return;
|
||||||
|
event.preventDefault();
|
||||||
|
const nextIndex = event.key === "Home" ? 0
|
||||||
|
: event.key === "End" ? tabs.length - 1
|
||||||
|
: (index + (event.key === "ArrowRight" ? 1 : -1) + tabs.length) % tabs.length;
|
||||||
|
tabs[nextIndex].focus();
|
||||||
|
selectTab(tabs[nextIndex].dataset.gradientTab || "definition");
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
const findCell = (depth: number, architecture: string, seed: number) =>
|
||||||
|
data.cells.find((cell: any) => cell.depth === depth && cell.architecture === architecture && cell.seed === seed);
|
||||||
|
const findDiagnostic = (cell: any, step: number) =>
|
||||||
|
cell.diagnostics.find((row: any) => row.step === step);
|
||||||
|
const seedList = (seedKey: string) => seedKey === "mean" ? seeds : [Number(seedKey)];
|
||||||
|
const average = (values: number[]) => values.reduce((sum, value) => sum + value, 0) / values.length;
|
||||||
|
const arrayFor = (depth: number, architecture: string, seedKey: string, step: number, field: string) => {
|
||||||
|
const rows = seedList(seedKey).map((seed: number) => findDiagnostic(findCell(depth, architecture, seed), step)[field]);
|
||||||
|
return rows[0].map((_: number, index: number) => average(rows.map((row: number[]) => row[index])));
|
||||||
|
};
|
||||||
|
const statFor = (depth: number, architecture: string, seedKey: string, step: number, group: string, field: string) =>
|
||||||
|
average(seedList(seedKey).map((seed: number) => findDiagnostic(findCell(depth, architecture, seed), step)[group][field]));
|
||||||
|
const formatStep = (value: number) => value.toLocaleString("en-US");
|
||||||
|
const formatScientific = (value: number) => value.toExponential(2).replace("e-", "e−");
|
||||||
|
const svgNode = (name: string, attrs: Record<string, string | number>) => {
|
||||||
|
const node = document.createElementNS(ns, name);
|
||||||
|
Object.entries(attrs).forEach(([key, value]) => node.setAttribute(key, String(value)));
|
||||||
|
return node;
|
||||||
|
};
|
||||||
|
|
||||||
|
const drawGroups = (group: SVGGElement, depth: number) => {
|
||||||
|
group.replaceChildren();
|
||||||
|
for (let index = 1; index < 8; index += 1) {
|
||||||
|
const block = depth / 8 * index;
|
||||||
|
const x = 58 + block / (depth - 1) * 822;
|
||||||
|
group.append(svgNode("line", { x1: x, x2: x, y1: 24, y2: 304, class: "group-boundary" }));
|
||||||
|
}
|
||||||
|
};
|
||||||
|
const drawChart = (
|
||||||
|
svg: SVGSVGElement,
|
||||||
|
groups: SVGGElement | null,
|
||||||
|
grid: SVGGElement,
|
||||||
|
lines: Record<string, SVGPolylineElement | null>,
|
||||||
|
pointGroups: Record<string, SVGGElement | null>,
|
||||||
|
series: Record<string, number[]>,
|
||||||
|
labels: string[],
|
||||||
|
options: { depth?: number; formatter?: (value: number) => string; zero?: boolean } = {},
|
||||||
|
) => {
|
||||||
|
grid.replaceChildren();
|
||||||
|
const values = Object.values(series).flat();
|
||||||
|
const rawMin = options.zero === false ? Math.min(...values) : 0;
|
||||||
|
const rawMax = Math.max(...values);
|
||||||
|
const padding = rawMax === rawMin ? 1 : (rawMax - rawMin) * 0.08;
|
||||||
|
const minimum = Math.max(0, rawMin - padding);
|
||||||
|
const maximum = rawMax + padding;
|
||||||
|
const xAt = (index: number) => 58 + index / Math.max(1, labels.length - 1) * 822;
|
||||||
|
const yAt = (value: number) => 24 + (maximum - value) / Math.max(1e-30, maximum - minimum) * 280;
|
||||||
|
const formatter = options.formatter || ((value: number) => value.toFixed(2));
|
||||||
|
for (let index = 0; index < 5; index += 1) {
|
||||||
|
const value = minimum + (maximum - minimum) * (4 - index) / 4;
|
||||||
|
const y = 24 + index / 4 * 280;
|
||||||
|
grid.append(svgNode("line", { x1: 58, x2: 880, y1: y, y2: y, class: "grid-line" }));
|
||||||
|
const text = svgNode("text", { x: 48, y: y + 4, class: "axis-label y" });
|
||||||
|
text.textContent = formatter(value);
|
||||||
|
grid.append(text);
|
||||||
|
}
|
||||||
|
labels.forEach((label, index) => {
|
||||||
|
const text = svgNode("text", { x: xAt(index), y: 326, class: "axis-label x" });
|
||||||
|
text.textContent = label;
|
||||||
|
grid.append(text);
|
||||||
|
});
|
||||||
|
if (groups && options.depth) drawGroups(groups, options.depth);
|
||||||
|
architectures.forEach((architecture) => {
|
||||||
|
const points = series[architecture].map((value, index) => `${xAt(index)},${yAt(value)}`).join(" ");
|
||||||
|
lines[architecture]?.setAttribute("points", points);
|
||||||
|
pointGroups[architecture]?.replaceChildren();
|
||||||
|
series[architecture].forEach((value, index) => {
|
||||||
|
const point = svgNode("circle", {
|
||||||
|
cx: xAt(index), cy: yAt(value), r: 3.6, fill: colors[architecture],
|
||||||
|
});
|
||||||
|
const title = svgNode("title", {});
|
||||||
|
title.textContent = `${architecture === "baseline" ? "Baseline" : "Block"} · ${labels[index]} · ${formatter(value)}`;
|
||||||
|
point.append(title);
|
||||||
|
pointGroups[architecture]?.append(point);
|
||||||
|
});
|
||||||
|
});
|
||||||
|
svg.dataset.maximum = String(maximum);
|
||||||
|
};
|
||||||
|
|
||||||
|
let spectrumDepth = 32;
|
||||||
|
let spectrumScale = "absolute";
|
||||||
|
const spectrumSeed = root.querySelector<HTMLSelectElement>("[data-spectrum-seed]")!;
|
||||||
|
const spectrumStep = root.querySelector<HTMLSelectElement>("[data-spectrum-step]")!;
|
||||||
|
const spectrumChart = root.querySelector<SVGSVGElement>("[data-spectrum-chart]")!;
|
||||||
|
const updateSpectrum = () => {
|
||||||
|
const seedKey = spectrumSeed.value;
|
||||||
|
const step = Number(spectrumStep.value);
|
||||||
|
const absolute = Object.fromEntries(architectures.map((architecture) => [
|
||||||
|
architecture,
|
||||||
|
arrayFor(spectrumDepth, architecture, seedKey, step, "activation_grad_rms_by_block"),
|
||||||
|
]));
|
||||||
|
const series = spectrumScale === "normalized"
|
||||||
|
? Object.fromEntries(architectures.map((architecture) => {
|
||||||
|
const mean = average(absolute[architecture]);
|
||||||
|
return [architecture, absolute[architecture].map((value: number) => value / mean)];
|
||||||
|
}))
|
||||||
|
: absolute;
|
||||||
|
drawChart(
|
||||||
|
spectrumChart,
|
||||||
|
spectrumChart.querySelector("[data-chart-groups]"),
|
||||||
|
spectrumChart.querySelector("[data-chart-grid]")!,
|
||||||
|
Object.fromEntries(architectures.map((architecture) => [architecture, spectrumChart.querySelector(`[data-chart-line="${architecture}"]`)])),
|
||||||
|
Object.fromEntries(architectures.map((architecture) => [architecture, spectrumChart.querySelector(`[data-chart-points="${architecture}"]`)])),
|
||||||
|
series,
|
||||||
|
Array.from({ length: spectrumDepth }, (_, index) => String(index + 1)),
|
||||||
|
{
|
||||||
|
depth: spectrumDepth,
|
||||||
|
formatter: spectrumScale === "absolute" ? formatScientific : (value) => `${value.toFixed(1)}×`,
|
||||||
|
},
|
||||||
|
);
|
||||||
|
const title = root.querySelector<HTMLElement>("[data-spectrum-title]");
|
||||||
|
if (title) title.textContent = spectrumScale === "absolute" ? "ACTIVATION GRADIENT RMS" : "NORMALIZED BY LAYER MEAN";
|
||||||
|
const state = root.querySelector<HTMLElement>("[data-spectrum-state]");
|
||||||
|
if (state) state.textContent = `DEPTH ${spectrumDepth} · ${seedKey === "mean" ? "3-SEED MEAN" : seedKey} · STEP ${formatStep(step)}`;
|
||||||
|
architectures.forEach((architecture) => {
|
||||||
|
const cv = statFor(spectrumDepth, architecture, seedKey, step, "activation_grad_statistics", "population_cv");
|
||||||
|
const ratio = statFor(spectrumDepth, architecture, seedKey, step, "activation_grad_statistics", "first_to_last_ratio");
|
||||||
|
const mean = statFor(spectrumDepth, architecture, seedKey, step, "activation_grad_statistics", "mean");
|
||||||
|
const prefix = architecture === "baseline" ? "base" : "block";
|
||||||
|
root.querySelector<HTMLElement>(`[data-spectrum-${prefix}-cv]`)!.textContent = cv.toFixed(4);
|
||||||
|
root.querySelector<HTMLElement>(`[data-spectrum-${prefix}-ratio]`)!.textContent = `${ratio.toFixed(2)}×`;
|
||||||
|
root.querySelector<HTMLElement>(`[data-spectrum-${prefix}-mean]`)!.textContent = formatScientific(mean);
|
||||||
|
});
|
||||||
|
};
|
||||||
|
root.querySelectorAll<HTMLButtonElement>("[data-spectrum-depth]").forEach((button) => button.addEventListener("click", () => {
|
||||||
|
spectrumDepth = Number(button.dataset.spectrumDepth);
|
||||||
|
root.querySelectorAll<HTMLButtonElement>("[data-spectrum-depth]").forEach((peer) => peer.setAttribute("aria-pressed", String(peer === button)));
|
||||||
|
updateSpectrum();
|
||||||
|
}));
|
||||||
|
root.querySelectorAll<HTMLButtonElement>("[data-spectrum-scale]").forEach((button) => button.addEventListener("click", () => {
|
||||||
|
spectrumScale = button.dataset.spectrumScale || "absolute";
|
||||||
|
root.querySelectorAll<HTMLButtonElement>("[data-spectrum-scale]").forEach((peer) => peer.setAttribute("aria-pressed", String(peer === button)));
|
||||||
|
updateSpectrum();
|
||||||
|
}));
|
||||||
|
spectrumSeed.addEventListener("change", updateSpectrum);
|
||||||
|
spectrumStep.addEventListener("change", updateSpectrum);
|
||||||
|
updateSpectrum();
|
||||||
|
|
||||||
|
let timeDepth = 32;
|
||||||
|
let timeMetric = "cv";
|
||||||
|
const timeSeed = root.querySelector<HTMLSelectElement>("[data-time-seed]")!;
|
||||||
|
const timeChart = root.querySelector<SVGSVGElement>("[data-time-chart]")!;
|
||||||
|
const timeMetricMap: Record<string, [string, string, (value: number) => string]> = {
|
||||||
|
cv: ["activation_grad_statistics", "population_cv", (value) => value.toFixed(2)],
|
||||||
|
imbalance: ["activation_grad_statistics", "imbalance_abs_log_ratio", (value) => value.toFixed(2)],
|
||||||
|
mean: ["activation_grad_statistics", "mean", formatScientific],
|
||||||
|
};
|
||||||
|
const updateTimeline = () => {
|
||||||
|
const seedKey = timeSeed.value;
|
||||||
|
const [group, field, formatter] = timeMetricMap[timeMetric];
|
||||||
|
const series = Object.fromEntries(architectures.map((architecture) => [
|
||||||
|
architecture,
|
||||||
|
steps.map((step: number) => statFor(timeDepth, architecture, seedKey, step, group, field)),
|
||||||
|
]));
|
||||||
|
drawChart(
|
||||||
|
timeChart,
|
||||||
|
null,
|
||||||
|
timeChart.querySelector("[data-time-grid]")!,
|
||||||
|
Object.fromEntries(architectures.map((architecture) => [architecture, timeChart.querySelector(`[data-time-line="${architecture}"]`)])),
|
||||||
|
Object.fromEntries(architectures.map((architecture) => [architecture, timeChart.querySelector(`[data-time-points="${architecture}"]`)])),
|
||||||
|
series,
|
||||||
|
steps.map((step: number) => step >= 1000 ? `${step / 1000}K` : String(step)),
|
||||||
|
{ formatter },
|
||||||
|
);
|
||||||
|
const titles: Record<string, string> = {
|
||||||
|
cv: "POPULATION CV",
|
||||||
|
imbalance: "ABS(LOG(FIRST / LAST QUARTILE))",
|
||||||
|
mean: "MEAN ACTIVATION GRADIENT RMS",
|
||||||
|
};
|
||||||
|
root.querySelector<HTMLElement>("[data-time-title]")!.textContent = titles[timeMetric];
|
||||||
|
root.querySelector<HTMLElement>("[data-time-copy]")!.textContent = `DEPTH ${timeDepth} · ${seedKey === "mean" ? "3-SEED MEAN" : seedKey}`;
|
||||||
|
};
|
||||||
|
root.querySelectorAll<HTMLButtonElement>("[data-time-depth]").forEach((button) => button.addEventListener("click", () => {
|
||||||
|
timeDepth = Number(button.dataset.timeDepth);
|
||||||
|
root.querySelectorAll<HTMLButtonElement>("[data-time-depth]").forEach((peer) => peer.setAttribute("aria-pressed", String(peer === button)));
|
||||||
|
updateTimeline();
|
||||||
|
}));
|
||||||
|
root.querySelectorAll<HTMLButtonElement>("[data-time-metric]").forEach((button) => button.addEventListener("click", () => {
|
||||||
|
timeMetric = button.dataset.timeMetric || "cv";
|
||||||
|
root.querySelectorAll<HTMLButtonElement>("[data-time-metric]").forEach((peer) => peer.setAttribute("aria-pressed", String(peer === button)));
|
||||||
|
updateTimeline();
|
||||||
|
}));
|
||||||
|
timeSeed.addEventListener("change", updateTimeline);
|
||||||
|
updateTimeline();
|
||||||
|
|
||||||
|
let outputDepth = 32;
|
||||||
|
const outputSeed = root.querySelector<HTMLSelectElement>("[data-output-seed]")!;
|
||||||
|
const outputStep = root.querySelector<HTMLSelectElement>("[data-output-step]")!;
|
||||||
|
const outputChart = root.querySelector<SVGSVGElement>("[data-output-chart]")!;
|
||||||
|
const updateOutput = () => {
|
||||||
|
const seedKey = outputSeed.value;
|
||||||
|
const step = Number(outputStep.value);
|
||||||
|
const series = Object.fromEntries(architectures.map((architecture) => [
|
||||||
|
architecture,
|
||||||
|
arrayFor(outputDepth, architecture, seedKey, step, "activation_output_rms_by_block"),
|
||||||
|
]));
|
||||||
|
drawChart(
|
||||||
|
outputChart,
|
||||||
|
outputChart.querySelector("[data-output-groups]"),
|
||||||
|
outputChart.querySelector("[data-output-grid]")!,
|
||||||
|
Object.fromEntries(architectures.map((architecture) => [architecture, outputChart.querySelector(`[data-output-line="${architecture}"]`)])),
|
||||||
|
Object.fromEntries(architectures.map((architecture) => [architecture, outputChart.querySelector(`[data-output-points="${architecture}"]`)])),
|
||||||
|
series,
|
||||||
|
Array.from({ length: outputDepth }, (_, index) => String(index + 1)),
|
||||||
|
{ depth: outputDepth, formatter: (value) => value.toFixed(2) },
|
||||||
|
);
|
||||||
|
root.querySelector<HTMLElement>("[data-output-copy]")!.textContent = `DEPTH ${outputDepth} · ${seedKey === "mean" ? "3-SEED MEAN" : seedKey} · STEP ${formatStep(step)}`;
|
||||||
|
root.querySelector<HTMLElement>("[data-rhythm-title]")!.textContent = `BLOCK ATTNRES · DEPTH ${outputDepth}`;
|
||||||
|
const bars = root.querySelector<HTMLElement>("[data-output-bars]")!;
|
||||||
|
const values: number[] = series.block;
|
||||||
|
const maximum = Math.max(...values);
|
||||||
|
const blocksPerGroup = outputDepth / 8;
|
||||||
|
bars.replaceChildren();
|
||||||
|
values.forEach((value, index) => {
|
||||||
|
const bar = document.createElement("i");
|
||||||
|
bar.style.setProperty("--bar", `${value / maximum * 100}%`);
|
||||||
|
if ((index + 1) % blocksPerGroup === 0) bar.classList.add("boundary");
|
||||||
|
const label = document.createElement("span");
|
||||||
|
label.textContent = String(index + 1);
|
||||||
|
bar.append(label);
|
||||||
|
bars.append(bar);
|
||||||
|
});
|
||||||
|
};
|
||||||
|
root.querySelectorAll<HTMLButtonElement>("[data-output-depth]").forEach((button) => button.addEventListener("click", () => {
|
||||||
|
outputDepth = Number(button.dataset.outputDepth);
|
||||||
|
root.querySelectorAll<HTMLButtonElement>("[data-output-depth]").forEach((peer) => peer.setAttribute("aria-pressed", String(peer === button)));
|
||||||
|
updateOutput();
|
||||||
|
}));
|
||||||
|
outputSeed.addEventListener("change", updateOutput);
|
||||||
|
outputStep.addEventListener("change", updateOutput);
|
||||||
|
updateOutput();
|
||||||
|
};
|
||||||
|
|
||||||
|
document.querySelectorAll<HTMLElement>("[data-gradient-lab]").forEach(initializeGradientLab);
|
||||||
|
document.addEventListener("astro:page-load", () => {
|
||||||
|
document.querySelectorAll<HTMLElement>("[data-gradient-lab]").forEach(initializeGradientLab);
|
||||||
|
});
|
||||||
|
</script>
|
||||||
|
|
||||||
|
<style>
|
||||||
|
.gradient-lab {
|
||||||
|
--g-ink: #1c201e;
|
||||||
|
--g-muted: #77746b;
|
||||||
|
--g-line: rgba(28, 32, 30, .16);
|
||||||
|
--g-paper: #f4f0e7;
|
||||||
|
--g-raised: #faf7ef;
|
||||||
|
--g-copper: #ba603b;
|
||||||
|
--g-green: #163f3b;
|
||||||
|
width: min(1120px, 100%);
|
||||||
|
margin: 42px 0;
|
||||||
|
color: var(--g-ink);
|
||||||
|
border: 1px solid var(--g-line);
|
||||||
|
background: var(--g-paper);
|
||||||
|
box-shadow: 0 30px 80px rgba(28, 32, 30, .09);
|
||||||
|
}
|
||||||
|
.gradient-head {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: minmax(0, 1.45fr) minmax(260px, .7fr);
|
||||||
|
gap: 44px;
|
||||||
|
padding: 30px;
|
||||||
|
color: #f5efe4;
|
||||||
|
background: var(--g-green);
|
||||||
|
}
|
||||||
|
.gradient-head p { margin: 0; color: rgba(245,239,228,.7); font: .65rem/1.7 var(--mono); }
|
||||||
|
.gradient-head div > p { color: #d58a68; letter-spacing: .08em; }
|
||||||
|
.gradient-head h3 { max-width: 720px; margin: 14px 0 0; color: inherit; font-size: clamp(1.15rem, 2.2vw, 1.75rem); line-height: 1.35; }
|
||||||
|
.gradient-ledger { display: grid; grid-template-columns: repeat(6, 1fr); border-bottom: 1px solid var(--g-line); }
|
||||||
|
.gradient-ledger article { min-height: 126px; padding: 18px 15px; border-right: 1px solid var(--g-line); }
|
||||||
|
.gradient-ledger article:last-child { border-right: 0; }
|
||||||
|
.gradient-ledger span, .panel-lead span { color: var(--g-muted); font: .56rem/1.2 var(--mono); letter-spacing: .08em; }
|
||||||
|
.gradient-ledger b { display: block; margin-top: 23px; font: 700 .95rem/1 var(--mono); }
|
||||||
|
.gradient-ledger p { margin: 8px 0 0; color: var(--g-muted); font-size: .6rem; line-height: 1.45; }
|
||||||
|
.gradient-ledger .pass { color: #f7f0e6; background: var(--g-green); }
|
||||||
|
.gradient-ledger .split { color: #f7f0e6; background: var(--g-copper); }
|
||||||
|
.gradient-ledger .pass span, .gradient-ledger .pass p, .gradient-ledger .split span, .gradient-ledger .split p { color: rgba(247,240,230,.72); }
|
||||||
|
.gradient-tabs { display: grid; grid-template-columns: repeat(5, 1fr); border-bottom: 1px solid var(--g-line); background: #e9e4da; }
|
||||||
|
.gradient-tabs button { min-height: 116px; padding: 16px; text-align: left; color: inherit; border: 0; border-right: 1px solid var(--g-line); background: transparent; cursor: pointer; }
|
||||||
|
.gradient-tabs button:last-child { border-right: 0; }
|
||||||
|
.gradient-tabs button[aria-selected="true"] { color: #f7f0e6; background: var(--g-copper); }
|
||||||
|
.gradient-tabs span, .gradient-tabs small { display: block; color: var(--g-muted); font: .54rem/1.25 var(--mono); }
|
||||||
|
.gradient-tabs b { display: block; margin: 15px 0 8px; font-size: .69rem; line-height: 1.35; }
|
||||||
|
.gradient-tabs button[aria-selected="true"] span, .gradient-tabs button[aria-selected="true"] small { color: rgba(247,240,230,.72); }
|
||||||
|
.gradient-panel { padding: 30px; }
|
||||||
|
.panel-lead { display: grid; grid-template-columns: 1.05fr .95fr; gap: 48px; align-items: end; margin-bottom: 28px; }
|
||||||
|
.panel-lead h4 { max-width: 680px; margin: 10px 0 0; font-size: 1.2rem; line-height: 1.4; }
|
||||||
|
.panel-lead p { margin: 0; color: var(--g-muted); font-size: .7rem; line-height: 1.7; }
|
||||||
|
.known-grid { display: grid; grid-template-columns: repeat(2, 1fr); border: 1px solid var(--g-line); }
|
||||||
|
.known-grid article { min-height: 270px; padding: 24px; }
|
||||||
|
.known-grid article + article { border-left: 1px solid var(--g-line); background: #ece2d6; }
|
||||||
|
.known-grid span, .object-chain span, .object-compare span, .split-result span, .timeline-notes span, .output-ratios span, .metric-pairs span, .cost-compare > article > span, .replay-ledger span, .hash-ledger span, .claim-grid span {
|
||||||
|
color: var(--g-copper); font: .56rem/1 var(--mono); letter-spacing: .06em;
|
||||||
|
}
|
||||||
|
.known-grid h5 { margin: 26px 0 16px; font-size: .95rem; }
|
||||||
|
.known-grid ul { margin: 0; padding-left: 18px; }
|
||||||
|
.known-grid li { margin-top: 11px; color: var(--g-muted); font-size: .68rem; line-height: 1.55; }
|
||||||
|
.object-chain { display: grid; grid-template-columns: 1fr 30px 1fr 30px .8fr 30px 1.2fr; gap: 6px; align-items: center; margin-top: 20px; }
|
||||||
|
.object-chain div { min-height: 145px; padding: 18px; border: 1px solid var(--g-line); background: var(--g-raised); }
|
||||||
|
.object-chain .accent { color: #f7f0e6; background: var(--g-green); }
|
||||||
|
.object-chain .accent span, .object-chain .accent p { color: rgba(247,240,230,.68); }
|
||||||
|
.object-chain b { display: block; margin-top: 25px; font: 700 .77rem/1.35 var(--mono); }
|
||||||
|
.object-chain p { color: var(--g-muted); font-size: .61rem; line-height: 1.45; }
|
||||||
|
.object-chain > i, .object-compare > i, .split-result > i, .replay-ledger > i { color: var(--g-copper); font-style: normal; text-align: center; }
|
||||||
|
.object-compare { display: grid; grid-template-columns: 1fr 50px 1fr; gap: 12px; align-items: center; margin-top: 20px; padding: 20px; background: #e9e4da; }
|
||||||
|
.object-compare article { padding: 12px; }
|
||||||
|
.object-compare b { display: block; margin-top: 18px; font: 700 1.1rem/1 var(--mono); }
|
||||||
|
.object-compare p { color: var(--g-muted); font-size: .65rem; line-height: 1.55; }
|
||||||
|
.object-compare > p { grid-column: 1/-1; margin: 0; padding-top: 16px; border-top: 1px solid var(--g-line); }
|
||||||
|
.definition-boundary { display: grid; grid-template-columns: 80px 1fr; margin-top: 20px; border-top: 1px solid var(--g-line); }
|
||||||
|
.definition-boundary > * { margin: 0; padding: 15px; border-right: 1px solid var(--g-line); border-bottom: 1px solid var(--g-line); }
|
||||||
|
.definition-boundary b { color: var(--g-copper); font: .6rem/1.4 var(--mono); }
|
||||||
|
.definition-boundary p { color: var(--g-muted); font-size: .66rem; line-height: 1.55; }
|
||||||
|
.lab-controls { display: flex; flex-wrap: wrap; gap: 10px 18px; align-items: end; margin-bottom: 20px; }
|
||||||
|
.lab-controls > div { display: flex; }
|
||||||
|
.lab-controls button, .lab-controls select { min-height: 38px; padding: 10px 12px; color: var(--g-muted); font: 700 .56rem/1 var(--mono); border: 1px solid var(--g-line); background: var(--g-raised); }
|
||||||
|
.lab-controls button { cursor: pointer; }
|
||||||
|
.lab-controls button + button { border-left: 0; }
|
||||||
|
.lab-controls button[aria-pressed="true"] { color: #fff9ef; background: var(--g-green); }
|
||||||
|
.lab-controls label { display: grid; gap: 6px; color: var(--g-muted); font: .52rem/1 var(--mono); }
|
||||||
|
.spectrum-layout { display: grid; grid-template-columns: minmax(0, 1fr) 235px; border: 1px solid var(--g-line); background: var(--g-raised); }
|
||||||
|
.chart-shell { min-width: 0; padding: 18px; border: 1px solid var(--g-line); background: var(--g-raised); }
|
||||||
|
.spectrum-layout .chart-shell { border: 0; border-right: 1px solid var(--g-line); }
|
||||||
|
.chart-shell header { display: flex; justify-content: space-between; gap: 12px; color: var(--g-muted); font: .55rem/1 var(--mono); }
|
||||||
|
.chart-shell svg { display: block; width: 100%; height: auto; margin-top: 10px; overflow: visible; }
|
||||||
|
.series { fill: none; stroke-width: 3; stroke-linejoin: round; stroke-linecap: round; }
|
||||||
|
.series.baseline { stroke: #77746b; }
|
||||||
|
.series.block { stroke: #ba603b; }
|
||||||
|
.grid-line { stroke: rgba(28,32,30,.1); stroke-width: 1; }
|
||||||
|
.group-boundary { stroke: rgba(22,63,59,.22); stroke-width: 1.5; stroke-dasharray: 4 4; }
|
||||||
|
.axis-label { fill: #8b867c; font: 11px var(--mono); }
|
||||||
|
.axis-label.y { text-anchor: end; }
|
||||||
|
.axis-label.x { text-anchor: middle; }
|
||||||
|
.axis-title { fill: #8b867c; font: 11px var(--mono); text-anchor: middle; }
|
||||||
|
.chart-legend { display: flex; flex-wrap: wrap; gap: 18px; margin-top: 4px; color: var(--g-muted); font: .56rem/1 var(--mono); }
|
||||||
|
.chart-legend span { display: inline-flex; gap: 7px; align-items: center; }
|
||||||
|
.chart-legend i { width: 22px; height: 3px; }
|
||||||
|
.chart-legend i.baseline { background: #77746b; }
|
||||||
|
.chart-legend i.block { background: #ba603b; }
|
||||||
|
.chart-legend i.boundary { height: 0; border-top: 2px dashed var(--g-green); background: transparent; }
|
||||||
|
.spectrum-readout { padding: 20px 18px; }
|
||||||
|
.spectrum-readout > span { color: var(--g-copper); font: .54rem/1.4 var(--mono); }
|
||||||
|
.spectrum-readout article { padding: 18px 0; border-bottom: 1px solid var(--g-line); }
|
||||||
|
.spectrum-readout article b { color: var(--g-muted); font: .52rem/1 var(--mono); }
|
||||||
|
.spectrum-readout article div { display: flex; justify-content: space-between; margin-top: 13px; }
|
||||||
|
.spectrum-readout article span { color: var(--g-muted); font: .5rem/1 var(--mono); }
|
||||||
|
.spectrum-readout article strong { font: 700 .7rem/1 var(--mono); }
|
||||||
|
.split-result { display: grid; grid-template-columns: 1fr 45px 1fr 45px 1fr; gap: 8px; align-items: center; margin-top: 20px; }
|
||||||
|
.split-result article { min-height: 145px; padding: 20px; border: 1px solid var(--g-line); }
|
||||||
|
.split-result .accent { color: #f7f0e6; background: var(--g-green); }
|
||||||
|
.split-result .accent span, .split-result .accent p { color: rgba(247,240,230,.68); }
|
||||||
|
.split-result b { display: block; margin-top: 25px; font-size: .83rem; }
|
||||||
|
.split-result p { color: var(--g-muted); font-size: .62rem; line-height: 1.5; }
|
||||||
|
.timeline-chart { margin-top: 0; }
|
||||||
|
.timeline-notes { display: grid; grid-template-columns: repeat(4, 1fr); margin-top: 20px; border-top: 1px solid var(--g-line); border-left: 1px solid var(--g-line); }
|
||||||
|
.timeline-notes article { min-height: 145px; padding: 18px; border-right: 1px solid var(--g-line); border-bottom: 1px solid var(--g-line); background: var(--g-raised); }
|
||||||
|
.timeline-notes .counter { background: #e9e4da; }
|
||||||
|
.timeline-notes b { display: block; margin-top: 25px; font: 700 .78rem/1 var(--mono); }
|
||||||
|
.timeline-notes p { color: var(--g-muted); font-size: .61rem; line-height: 1.5; }
|
||||||
|
.output-ratios { display: grid; grid-template-columns: 1fr 1fr 1.25fr; margin-top: 20px; border-top: 1px solid var(--g-line); border-left: 1px solid var(--g-line); }
|
||||||
|
.output-ratios article { min-height: 170px; padding: 20px; border-right: 1px solid var(--g-line); border-bottom: 1px solid var(--g-line); }
|
||||||
|
.output-ratios article > div { display: flex; justify-content: space-between; margin-top: 24px; }
|
||||||
|
.output-ratios article > div b { color: var(--g-muted); font: .55rem/1 var(--mono); }
|
||||||
|
.output-ratios article > div strong { font: 700 .75rem/1 var(--mono); }
|
||||||
|
.output-ratios .accent { color: #f7f0e6; background: var(--g-green); }
|
||||||
|
.output-ratios .accent span, .output-ratios .accent p { color: rgba(247,240,230,.68); }
|
||||||
|
.output-ratios .accent b { display: block; margin-top: 25px; font-size: .8rem; }
|
||||||
|
.output-ratios p { color: var(--g-muted); font-size: .62rem; line-height: 1.5; }
|
||||||
|
.group-rhythm { margin-top: 20px; padding: 20px; color: #f7f0e6; background: var(--g-green); overflow: hidden; }
|
||||||
|
.group-rhythm header { display: flex; justify-content: space-between; color: rgba(247,240,230,.68); font: .55rem/1 var(--mono); }
|
||||||
|
.group-rhythm > div { display: grid; grid-template-columns: repeat(auto-fit, minmax(8px, 1fr)); align-items: end; height: 170px; margin-top: 18px; border-bottom: 1px solid rgba(255,255,255,.3); }
|
||||||
|
.group-rhythm i { position: relative; display: block; height: max(5px, var(--bar)); margin-right: 2px; background: #d4835d; }
|
||||||
|
.group-rhythm i.boundary { margin-right: 8px; border-right: 2px solid rgba(255,255,255,.75); }
|
||||||
|
.group-rhythm i span { position: absolute; bottom: -18px; left: 50%; color: rgba(255,255,255,.55); font: .43rem/1 var(--mono); transform: translateX(-50%); }
|
||||||
|
.group-rhythm p { margin: 35px 0 0; color: rgba(247,240,230,.72); font-size: .64rem; }
|
||||||
|
.verdict-table-wrap { overflow-x: auto; }
|
||||||
|
.verdict-table { width: 100%; min-width: 850px; border-collapse: collapse; font-size: .62rem; }
|
||||||
|
.verdict-table th, .verdict-table td { padding: 15px 12px; border-bottom: 1px solid var(--g-line); text-align: left; }
|
||||||
|
.verdict-table thead th { color: var(--g-muted); font: .53rem/1.3 var(--mono); }
|
||||||
|
.verdict-table tbody th, .verdict-table tbody td:last-child { font: 700 .66rem/1 var(--mono); }
|
||||||
|
.verdict-table .good { color: var(--g-green); }
|
||||||
|
.verdict-table .bad { color: var(--g-copper); }
|
||||||
|
.metric-pairs { display: grid; grid-template-columns: repeat(4, 1fr); margin-top: 20px; border-top: 1px solid var(--g-line); border-left: 1px solid var(--g-line); }
|
||||||
|
.metric-pairs article { min-height: 150px; padding: 18px; border-right: 1px solid var(--g-line); border-bottom: 1px solid var(--g-line); background: var(--g-raised); }
|
||||||
|
.metric-pairs article.warn { border-top: 4px solid var(--g-copper); }
|
||||||
|
.metric-pairs article.pass { border-top: 4px solid var(--g-green); }
|
||||||
|
.metric-pairs b { display: block; margin-top: 25px; font-size: .77rem; }
|
||||||
|
.metric-pairs p { color: var(--g-muted); font-size: .61rem; line-height: 1.5; }
|
||||||
|
.cost-compare { display: grid; grid-template-columns: repeat(3, 1fr); margin-top: 20px; border-top: 1px solid var(--g-line); border-left: 1px solid var(--g-line); }
|
||||||
|
.cost-compare article { min-height: 205px; padding: 20px; border-right: 1px solid var(--g-line); border-bottom: 1px solid var(--g-line); }
|
||||||
|
.cost-compare article > div { display: flex; justify-content: space-between; gap: 12px; margin-top: 25px; }
|
||||||
|
.cost-compare article > div b { color: var(--g-muted); font: .52rem/1 var(--mono); }
|
||||||
|
.cost-compare article > div strong { font: 700 .62rem/1 var(--mono); text-align: right; }
|
||||||
|
.cost-compare p { color: var(--g-copper); font: .56rem/1.5 var(--mono); }
|
||||||
|
.cost-compare .boundary { color: #f7f0e6; background: var(--g-green); }
|
||||||
|
.cost-compare .boundary span, .cost-compare .boundary p { color: rgba(247,240,230,.68); }
|
||||||
|
.cost-compare .boundary b { display: block; margin-top: 30px; font-size: .85rem; }
|
||||||
|
.replay-ledger { display: grid; grid-template-columns: 1fr 40px 1fr 40px 1fr; gap: 8px; align-items: center; margin-top: 20px; }
|
||||||
|
.replay-ledger article { min-height: 145px; padding: 18px; border: 1px solid var(--g-line); background: var(--g-raised); }
|
||||||
|
.replay-ledger article.pass { color: #f7f0e6; background: var(--g-green); }
|
||||||
|
.replay-ledger article.pass span, .replay-ledger article.pass p { color: rgba(247,240,230,.68); }
|
||||||
|
.replay-ledger b { display: block; margin-top: 25px; font-size: .7rem; }
|
||||||
|
.replay-ledger p { color: var(--g-muted); font-size: .58rem; line-height: 1.5; overflow-wrap: anywhere; }
|
||||||
|
.hash-ledger { display: grid; grid-template-columns: repeat(2, 1fr); margin-top: 20px; border-top: 1px solid var(--g-line); border-left: 1px solid var(--g-line); }
|
||||||
|
.hash-ledger article { min-width: 0; padding: 16px; border-right: 1px solid var(--g-line); border-bottom: 1px solid var(--g-line); background: #e9e4da; }
|
||||||
|
.hash-ledger code { display: block; margin-top: 12px; overflow: hidden; color: var(--g-green); font: .54rem/1.3 var(--mono); text-overflow: ellipsis; }
|
||||||
|
.claim-grid { display: grid; grid-template-columns: repeat(2, 1fr); margin-top: 20px; }
|
||||||
|
.claim-grid article { padding: 22px; }
|
||||||
|
.claim-grid .yes { color: #f7f0e6; background: var(--g-green); }
|
||||||
|
.claim-grid .no { background: #e7d8ca; }
|
||||||
|
.claim-grid ul { margin: 18px 0 0; padding-left: 17px; }
|
||||||
|
.claim-grid li { margin-top: 10px; font-size: .65rem; line-height: 1.55; }
|
||||||
|
.claim-grid .yes span, .claim-grid .yes li { color: rgba(247,240,230,.78); }
|
||||||
|
@media (max-width: 920px) {
|
||||||
|
.gradient-ledger { grid-template-columns: repeat(3, 1fr); }
|
||||||
|
.gradient-ledger article:nth-child(3) { border-right: 0; }
|
||||||
|
.gradient-tabs { grid-template-columns: repeat(3, 1fr); }
|
||||||
|
.spectrum-layout { grid-template-columns: 1fr; }
|
||||||
|
.spectrum-layout .chart-shell { border-right: 0; border-bottom: 1px solid var(--g-line); }
|
||||||
|
.timeline-notes, .metric-pairs { grid-template-columns: repeat(2, 1fr); }
|
||||||
|
.object-chain { grid-template-columns: 1fr 24px 1fr; }
|
||||||
|
.object-chain > i:nth-of-type(n+3) { display: none; }
|
||||||
|
}
|
||||||
|
@media (max-width: 680px) {
|
||||||
|
.gradient-head, .panel-lead, .known-grid { grid-template-columns: 1fr; gap: 20px; }
|
||||||
|
.gradient-head, .gradient-panel { padding: 20px; }
|
||||||
|
.known-grid article + article { border-left: 0; border-top: 1px solid var(--g-line); }
|
||||||
|
.gradient-ledger { grid-template-columns: repeat(2, 1fr); }
|
||||||
|
.gradient-ledger article:nth-child(3) { border-right: 1px solid var(--g-line); }
|
||||||
|
.gradient-ledger article:nth-child(even) { border-right: 0; }
|
||||||
|
.gradient-tabs { display: flex; overflow-x: auto; }
|
||||||
|
.gradient-tabs button { flex: 0 0 190px; }
|
||||||
|
.lab-controls { align-items: stretch; }
|
||||||
|
.lab-controls > div, .lab-controls label { flex: 1 0 100%; }
|
||||||
|
.lab-controls button { flex: 1; }
|
||||||
|
.lab-controls select { width: 100%; }
|
||||||
|
.object-chain, .object-compare, .split-result, .replay-ledger { grid-template-columns: 1fr; }
|
||||||
|
.object-chain > i, .object-compare > i, .split-result > i, .replay-ledger > i { display: block !important; transform: rotate(90deg); }
|
||||||
|
.object-compare > p { grid-column: 1; }
|
||||||
|
.definition-boundary { grid-template-columns: 70px 1fr; }
|
||||||
|
.timeline-notes, .output-ratios, .metric-pairs, .cost-compare, .hash-ledger, .claim-grid { grid-template-columns: 1fr; }
|
||||||
|
.chart-shell { padding: 12px 8px; }
|
||||||
|
.chart-shell header { padding: 0 8px; }
|
||||||
|
.axis-label { font-size: 9px; }
|
||||||
|
.group-rhythm > div { min-width: 620px; }
|
||||||
|
.group-rhythm { overflow-x: auto; }
|
||||||
|
}
|
||||||
|
</style>
|
||||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -2,6 +2,7 @@
|
|||||||
import BaseLayout from "@/layouts/BaseLayout.astro";
|
import BaseLayout from "@/layouts/BaseLayout.astro";
|
||||||
import ArchitectureExplorer from "@/components/ArchitectureExplorer.astro";
|
import ArchitectureExplorer from "@/components/ArchitectureExplorer.astro";
|
||||||
import K3ArtifactLab from "@/components/K3ArtifactLab.astro";
|
import K3ArtifactLab from "@/components/K3ArtifactLab.astro";
|
||||||
|
import K3AttnResGradientLab from "@/components/K3AttnResGradientLab.astro";
|
||||||
import K3AttnResTraceLab from "@/components/K3AttnResTraceLab.astro";
|
import K3AttnResTraceLab from "@/components/K3AttnResTraceLab.astro";
|
||||||
import K3ReportLab from "@/components/K3ReportLab.astro";
|
import K3ReportLab from "@/components/K3ReportLab.astro";
|
||||||
import { k3FigureAtlas, k3Ledgers, k3PaperChain, k3ReportMap } from "@/data/k3";
|
import { k3FigureAtlas, k3Ledgers, k3PaperChain, k3ReportMap } from "@/data/k3";
|
||||||
@@ -38,7 +39,8 @@ const toc = [
|
|||||||
["28", "lab", "八联交互实验"],
|
["28", "lab", "八联交互实验"],
|
||||||
["29", "artifacts", "开放权重工件审计"],
|
["29", "artifacts", "开放权重工件审计"],
|
||||||
["30", "attnres-reduced", "AttnRes 缩小机制实验"],
|
["30", "attnres-reduced", "AttnRes 缩小机制实验"],
|
||||||
["31", "audit", "21 张图表审计"],
|
["31", "attnres-gradient", "梯度定义与深度扩展"],
|
||||||
|
["32", "audit", "21 张图表审计"],
|
||||||
["↳", "papers", "100 节点阅读链"],
|
["↳", "papers", "100 节点阅读链"],
|
||||||
];
|
];
|
||||||
|
|
||||||
@@ -107,13 +109,13 @@ const paperGroups = [
|
|||||||
|
|
||||||
<BaseLayout
|
<BaseLayout
|
||||||
title="Kimi K3 技术报告完整深读:架构、训练、RL、系统与评测"
|
title="Kimi K3 技术报告完整深读:架构、训练、RL、系统与评测"
|
||||||
description="用三十二张问题账、二十一张图表审计、八个机制实验、四个开放工件视图、五个 AttnRes 独立实验视图与一百个一手阅读节点,逐节读懂 Kimi K3。"
|
description="用三十二张问题账、二十一张图表审计、八个机制实验、四个开放工件视图、两轮十个 AttnRes 独立实验视图与一百个一手阅读节点,逐节读懂 Kimi K3。"
|
||||||
section="k3"
|
section="k3"
|
||||||
>
|
>
|
||||||
<header class="page-hero k3-hero">
|
<header class="page-hero k3-hero">
|
||||||
<div class="page-hero-inner">
|
<div class="page-hero-inner">
|
||||||
<div>
|
<div>
|
||||||
<p class="eyebrow"><span>ANCHOR REPORT / ROUND 04</span> KIMI K3 · REPORT → ARTIFACTS → INDEPENDENT PROBE</p>
|
<p class="eyebrow"><span>ANCHOR REPORT / ROUND 05</span> KIMI K3 · REPORT → ARTIFACTS → INDEPENDENT PROBE</p>
|
||||||
<h1>不把报告压成摘要<br />把每个因果环节<br />重新展开</h1>
|
<h1>不把报告压成摘要<br />把每个因果环节<br />重新展开</h1>
|
||||||
<p class="lead">
|
<p class="lead">
|
||||||
K3 同时扩展序列、深度、宽度、视觉与 Agent 轨迹。真正值得读的不是 2.8T 这个最大数字,
|
K3 同时扩展序列、深度、宽度、视觉与 Agent 轨迹。真正值得读的不是 2.8T 这个最大数字,
|
||||||
@@ -123,11 +125,11 @@ const paperGroups = [
|
|||||||
<dl class="page-facts">
|
<dl class="page-facts">
|
||||||
<div><dt>QUESTIONS</dt><dd>32 张问题账</dd></div>
|
<div><dt>QUESTIONS</dt><dd>32 张问题账</dd></div>
|
||||||
<div><dt>REPORT</dt><dd>16 Figures · 5 Tables</dd></div>
|
<div><dt>REPORT</dt><dd>16 Figures · 5 Tables</dd></div>
|
||||||
<div><dt>LABS</dt><dd>8 + 4 + 5 个交互视图</dd></div>
|
<div><dt>LABS</dt><dd>8 + 4 + 5 + 5 个交互视图</dd></div>
|
||||||
<div><dt>READING</dt><dd>100 个一手 / 官方节点</dd></div>
|
<div><dt>READING</dt><dd>100 个一手 / 官方节点</dd></div>
|
||||||
<div><dt>MODEL</dt><dd>2.78T total / 104.2B active</dd></div>
|
<div><dt>MODEL</dt><dd>2.78T total / 104.2B active</dd></div>
|
||||||
<div><dt>ARTIFACTS</dt><dd>96 shards · 497,220 tensors</dd></div>
|
<div><dt>ARTIFACTS</dt><dd>96 shards · 497,220 tensors</dd></div>
|
||||||
<div><dt>STATUS</dt><dd>K3 四轮 · AttnRes 实验</dd></div>
|
<div><dt>STATUS</dt><dd>K3 五轮 · 梯度定义闭环</dd></div>
|
||||||
</dl>
|
</dl>
|
||||||
</div>
|
</div>
|
||||||
</header>
|
</header>
|
||||||
@@ -892,12 +894,36 @@ const paperGroups = [
|
|||||||
<a class="button primary" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/research/K3_ATTNRES_REDUCED_AUDIT.md">阅读完整研究审计</a>
|
<a class="button primary" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/research/K3_ATTNRES_REDUCED_AUDIT.md">阅读完整研究审计</a>
|
||||||
<a class="button" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/experiments/k3/attnres">复跑公开实验代码</a>
|
<a class="button" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/experiments/k3/attnres">复跑公开实验代码</a>
|
||||||
<a class="button" href="https://arxiv.org/abs/2603.15031">Attention Residuals 原论文</a>
|
<a class="button" href="https://arxiv.org/abs/2603.15031">Attention Residuals 原论文</a>
|
||||||
<a class="button" href="https://github.com/MoonshotAI/Attention-Residuals">官方实现</a>
|
<a class="button" href="https://github.com/MoonshotAI/Attention-Residuals">官方论文工件</a>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="article-section" id="attnres-gradient">
|
||||||
|
<p class="eyebrow"><span>31</span> GRADIENT DEFINITION × DEPTH SCALE</p>
|
||||||
|
<h2>“论文说梯度更均匀”,和上一轮参数梯度反结果,测的是同一件事吗?</h2>
|
||||||
|
<p class="lede">
|
||||||
|
第五轮先审计 Attention Residuals 官方论文与仓库:Figure 5(c) 没有公开 gradient tensor、
|
||||||
|
norm、reduction、diagnostic batch、AMP / clipping 时点或统计代码。本站因此冻结一个可复现的
|
||||||
|
post-MLP output activation-gradient 定义,把深度扩到 16 / 32 blocks、预算扩到 8,000 steps,
|
||||||
|
再用三 seed 检查“首尾平衡”和“全层离散度”是否真的同方向。
|
||||||
|
</p>
|
||||||
|
<div class="artifact-callout">
|
||||||
|
<article><span>F / FROZEN</span><b>12 × 8,000 steps</b><p>786,432,000 formal target bytes;两深度、两结构、三 seed。</p></article>
|
||||||
|
<article><span>X / OBSERVED</span><b>first/last 6 / 6 改善</b><p>depth-16 平均 61.0%;depth-32 平均 72.0%。</p></article>
|
||||||
|
<article class="warning"><span>X / COUNTEREVIDENCE</span><b>CV 6 / 6 恶化</b><p>局部尖峰让 depth-16 / 32 平均相对恶化 10.3% / 60.0%。</p></article>
|
||||||
|
<article><span>R / REPLAY</span><b>model + optimizer exact</b><p>指定 32-layer Block 格从零重训 8,000 steps,冻结字段逐项一致。</p></article>
|
||||||
|
</div>
|
||||||
|
<K3AttnResGradientLab />
|
||||||
|
<div class="hero-actions">
|
||||||
|
<a class="button primary" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/research/K3_ATTNRES_GRADIENT_SCALE_AUDIT.md">阅读完整结果审计</a>
|
||||||
|
<a class="button" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/research/K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md">核对论文定义边界</a>
|
||||||
|
<a class="button" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/experiments/k3/attnres_gradient">复跑 12 格实验</a>
|
||||||
|
<a class="button" href="https://github.com/MoonshotAI/Attention-Residuals">官方一手工件</a>
|
||||||
</div>
|
</div>
|
||||||
</section>
|
</section>
|
||||||
|
|
||||||
<section class="article-section" id="audit">
|
<section class="article-section" id="audit">
|
||||||
<p class="eyebrow"><span>31</span> FIGURE & TABLE AUDIT</p>
|
<p class="eyebrow"><span>32</span> FIGURE & TABLE AUDIT</p>
|
||||||
<h2>Figure 1–16、Table 1–5:每张图究竟支持什么,不能支持什么</h2>
|
<h2>Figure 1–16、Table 1–5:每张图究竟支持什么,不能支持什么</h2>
|
||||||
<div class="figure-atlas">
|
<div class="figure-atlas">
|
||||||
{k3FigureAtlas.map(([id, report, title, contract]) => (
|
{k3FigureAtlas.map(([id, report, title, contract]) => (
|
||||||
|
|||||||
@@ -9,7 +9,7 @@ const researching = chapters.filter((chapter) => ["researching", "drafting"].inc
|
|||||||
const workstreams = [
|
const workstreams = [
|
||||||
{ label: "研究框架与规范", value: 83, next: "给 Scaling 与推理专题补逐篇图表/实验精读层级" },
|
{ label: "研究框架与规范", value: 83, next: "给 Scaling 与推理专题补逐篇图表/实验精读层级" },
|
||||||
{ label: "网站设计系统", value: 89, next: "打印样式与更多通用可视化组件" },
|
{ label: "网站设计系统", value: 89, next: "打印样式与更多通用可视化组件" },
|
||||||
{ label: "Kimi K3 深读", value: 96, next: "对齐 AttnRes 梯度定义并扩展深度/预算;等待 A_log 官方转换合同" },
|
{ label: "Kimi K3 深读", value: 98, next: "对齐 layer 21–25 梯度尖峰与 mixer weights;等待 A_log 官方转换合同" },
|
||||||
{ label: "语言模型前史", value: 78, next: "逐图精读 Kneser–Ney、LSTM 与 Bahdanau,并加入真实小语料复现" },
|
{ label: "语言模型前史", value: 78, next: "逐图精读 Kneser–Ney、LSTM 与 Bahdanau,并加入真实小语料复现" },
|
||||||
{ label: "Transformer 基础", value: 79, next: "逐图精读多头电路、Pre/Post-LN 与真实 kernel / KV 配置" },
|
{ label: "Transformer 基础", value: 79, next: "逐图精读多头电路、Pre/Post-LN 与真实 kernel / KV 配置" },
|
||||||
{ label: "表示、位置与残差高速公路", value: 81, next: "加入真实 hidden-state / norm traces、长上下文位置外推复现与更多深层稳定性消融" },
|
{ label: "表示、位置与残差高速公路", value: 81, next: "加入真实 hidden-state / norm traces、长上下文位置外推复现与更多深层稳定性消融" },
|
||||||
@@ -50,7 +50,7 @@ const workstreams = [
|
|||||||
<div><dt>OVERALL</dt><dd>专题平均 {average}%</dd></div>
|
<div><dt>OVERALL</dt><dd>专题平均 {average}%</dd></div>
|
||||||
<div><dt>READABLE</dt><dd>{published} 个首版可读专题</dd></div>
|
<div><dt>READABLE</dt><dd>{published} 个首版可读专题</dd></div>
|
||||||
<div><dt>ACTIVE</dt><dd>{researching} 个研究/写作中</dd></div>
|
<div><dt>ACTIVE</dt><dd>{researching} 个研究/写作中</dd></div>
|
||||||
<div><dt>UPDATED</dt><dd>2026-07-30 07:30 CST</dd></div>
|
<div><dt>UPDATED</dt><dd>2026-07-30 10:05 CST</dd></div>
|
||||||
<div><dt>MODE</dt><dd>持续迭代,不锁死版本</dd></div>
|
<div><dt>MODE</dt><dd>持续迭代,不锁死版本</dd></div>
|
||||||
</dl>
|
</dl>
|
||||||
</div>
|
</div>
|
||||||
@@ -97,7 +97,7 @@ const workstreams = [
|
|||||||
<article><span>✓</span><h3>K3 报告已结构化拆解</h3><p>47 页报告目录、151 条参考来源和架构/后训练/系统主线已经提取。</p></article>
|
<article><span>✓</span><h3>K3 报告已结构化拆解</h3><p>47 页报告目录、151 条参考来源和架构/后训练/系统主线已经提取。</p></article>
|
||||||
<article><span>✓</span><h3>17 专题知识图</h3><p>从语言模型基础到评测安全,包含先修依赖和三条贯穿案例。</p></article>
|
<article><span>✓</span><h3>17 专题知识图</h3><p>从语言模型基础到评测安全,包含先修依赖和三条贯穿案例。</p></article>
|
||||||
<article><span>✓</span><h3>编辑式网站系统</h3><p>响应式导航、章节模板、侧栏、进度、论文链和证据提示组件。</p></article>
|
<article><span>✓</span><h3>编辑式网站系统</h3><p>响应式导航、章节模板、侧栏、进度、论文链和证据提示组件。</p></article>
|
||||||
<article><span>✓</span><h3>九十四个原创交互视图</h3><p>K3 三轴图、八联报告实验、四联开放工件实验与五联 AttnRes 独立实验,DeepSeek 四联公式实验、十三联 Base 工件实验、Chat 行为、completion/full-depth、multi-seed、cross-source 与 task-bootstrap CRN 五轮实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
|
<article><span>✓</span><h3>九十九个原创交互视图</h3><p>K3 三轴图、八联报告实验、四联开放工件实验与两轮十联 AttnRes 独立实验,DeepSeek 四联公式实验、十三联 Base 工件实验、Chat 行为、completion/full-depth、multi-seed、cross-source 与 task-bootstrap CRN 五轮实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
|
||||||
<article><span>✓</span><h3>十七篇首版长文</h3><p>K3、语言模型前史、Transformer、表示/位置/残差、DeepSeek、Scaling、数据工程、长上下文、MoE、后训练、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全专题。</p></article>
|
<article><span>✓</span><h3>十七篇首版长文</h3><p>K3、语言模型前史、Transformer、表示/位置/残差、DeepSeek、Scaling、数据工程、长上下文、MoE、后训练、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全专题。</p></article>
|
||||||
<article><span>✓</span><h3>语言模型前史深度专题</h3><p>八张独立问题账、33 个正式节点、20 段长文与概率—向量—记忆—对齐四联实验。</p></article>
|
<article><span>✓</span><h3>语言模型前史深度专题</h3><p>八张独立问题账、33 个正式节点、20 段长文与概率—向量—记忆—对齐四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>Transformer 深度专题</h3><p>十张独立问题账、40 个正式节点、21 段正文与 QKV—Mask—多头位置—Block 成本四联实验。</p></article>
|
<article><span>✓</span><h3>Transformer 深度专题</h3><p>十张独立问题账、40 个正式节点、21 段正文与 QKV—Mask—多头位置—Block 成本四联实验。</p></article>
|
||||||
@@ -106,6 +106,7 @@ const workstreams = [
|
|||||||
<article><span>✓</span><h3>Kimi K3 技术报告二轮深读</h3><p>三十二张问题账、Figure 1–16 / Table 1–5 审计、100 节点阅读链,以及 Delta—Decay—AttnRes—LatentMoE—SiTU—QB—MOPD—Cache 八联实验。</p></article>
|
<article><span>✓</span><h3>Kimi K3 技术报告二轮深读</h3><p>三十二张问题账、Figure 1–16 / Table 1–5 审计、100 节点阅读链,以及 Delta—Decay—AttnRes—LatentMoE—SiTU—QB—MOPD—Cache 八联实验。</p></article>
|
||||||
<article><span>✓</span><h3>Kimi K3 三轮开放工件里程碑</h3><p>固定官方 revisions,审计 96 个 shards、497,220 个 tensor entries 与真实 KDA / MLA / MoE / MoonViT shapes;四联实验分开显示层型、tensor anatomy、参数范围和复现边界。</p></article>
|
<article><span>✓</span><h3>Kimi K3 三轮开放工件里程碑</h3><p>固定官方 revisions,审计 96 个 shards、497,220 个 tensor entries 与真实 KDA / MLA / MoE / MoonViT shapes;四联实验分开显示层型、tensor anatomy、参数范围和复现边界。</p></article>
|
||||||
<article><span>✓</span><h3>Kimi K3 四轮 AttnRes 独立实验</h3><p>冻结三结构 × 三 seed 的 9 个 2,000-step 格;Full / Block 相对 Baseline 的平均 paired delta 为 −0.01457 / −0.04247 BPC,但核心参数梯度 CV 没有复现论文叙述。指定正式格全新进程八字段 exact,五视图同时展示结果、反证、成本与 claim boundary。</p></article>
|
<article><span>✓</span><h3>Kimi K3 四轮 AttnRes 独立实验</h3><p>冻结三结构 × 三 seed 的 9 个 2,000-step 格;Full / Block 相对 Baseline 的平均 paired delta 为 −0.01457 / −0.04247 BPC,但核心参数梯度 CV 没有复现论文叙述。指定正式格全新进程八字段 exact,五视图同时展示结果、反证、成本与 claim boundary。</p></article>
|
||||||
|
<article><span>✓</span><h3>Kimi K3 五轮梯度定义与深度扩展</h3><p>先确认 Figure 5 没有公开唯一 gradient telemetry 合同,再冻结 16/32 blocks × Baseline/Block × 3 seeds 的 12 个 8,000-step 格。Block 的首尾失衡 6/6 改善但全层 CV 6/6 恶化,两个深度都判为 mixed;指定 32 层格完整重训的模型、优化器与全部冻结字段 exact。</p></article>
|
||||||
<article><span>✓</span><h3>FlashKDA RTX 5090 执行闸门</h3><p>隔离 CUDA 13.0 / glibc 2.39 编译 sm_120a wheel;6/6 官方参考逐元素相等,并完成 fixed / varlen、三种 state mode 的 1,800 个 CUDA Event samples。</p></article>
|
<article><span>✓</span><h3>FlashKDA RTX 5090 执行闸门</h3><p>隔离 CUDA 13.0 / glibc 2.39 编译 sm_120a wheel;6/6 官方参考逐元素相等,并完成 fixed / varlen、三种 state mode 的 1,800 个 CUDA Event samples。</p></article>
|
||||||
<article><span>✓</span><h3>Scaling Laws 深度专题</h3><p>九张账、29 个一手节点、DeepSeek/Kimi 双谱系与曲面—部署—复用—涌现四联实验。</p></article>
|
<article><span>✓</span><h3>Scaling Laws 深度专题</h3><p>九张账、29 个一手节点、DeepSeek/Kimi 双谱系与曲面—部署—复用—涌现四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>数据工程深度专题</h3><p>十二张账、31 个一手节点、DeepSeek/Kimi 双谱系与流水线—去重—混合—改写四联实验。</p></article>
|
<article><span>✓</span><h3>数据工程深度专题</h3><p>十二张账、31 个一手节点、DeepSeek/Kimi 双谱系与流水线—去重—混合—改写四联实验。</p></article>
|
||||||
@@ -134,7 +135,7 @@ const workstreams = [
|
|||||||
</div>
|
</div>
|
||||||
<div class="queue-table">
|
<div class="queue-table">
|
||||||
<div class="head"><b>优先级</b><b>专题</b><b>本轮交付</b><b>完成闸门</b></div>
|
<div class="head"><b>优先级</b><b>专题</b><b>本轮交付</b><b>完成闸门</b></div>
|
||||||
<div><span>P0</span><strong>K3 四轮后续</strong><p>对齐论文梯度定义 → 增加 depth / budget → 等待 A_log 官方合同后进入真实 checkpoint forward</p><em>尺度复查 + 工件边界</em></div>
|
<div><span>P0</span><strong>K3 五轮后续</strong><p>对齐 layer 21–25 尖峰、pre-attention / pre-MLP 与 mixer source weights → 等待 A_log 官方合同后进入真实 checkpoint forward</p><em>局部机制 + 工件边界</em></div>
|
||||||
<div><span>P0</span><strong>DeepSeek 八轮后续</strong><p>干预式 mediation → SM90 FlashMLA / FP8 / pipeline traces → R1-like RL 小模型复现</p><em>运行证据 + 独立复现</em></div>
|
<div><span>P0</span><strong>DeepSeek 八轮后续</strong><p>干预式 mediation → SM90 FlashMLA / FP8 / pipeline traces → R1-like RL 小模型复现</p><em>运行证据 + 独立复现</em></div>
|
||||||
<div><span>P0</span><strong>Transformer 二轮</strong><p>多头电路逐图 → Pre/Post-LN 真实 traces → Flash/KV 配置与 kernel 对照</p><em>逐图笔记 + 实测边界</em></div>
|
<div><span>P0</span><strong>Transformer 二轮</strong><p>多头电路逐图 → Pre/Post-LN 真实 traces → Flash/KV 配置与 kernel 对照</p><em>逐图笔记 + 实测边界</em></div>
|
||||||
<div><span>P0</span><strong>表示、位置与残差二轮</strong><p>真实 hidden-state / norm traces → 长上下文位置外推 → mHC / AttnRes 深层稳定性消融</p><em>可复现实验 + 逐图笔记</em></div>
|
<div><span>P0</span><strong>表示、位置与残差二轮</strong><p>真实 hidden-state / norm traces → 长上下文位置外推 → mHC / AttnRes 深层稳定性消融</p><em>可复现实验 + 逐图笔记</em></div>
|
||||||
@@ -220,6 +221,10 @@ const workstreams = [
|
|||||||
<div><time>2026-07-30</time><b>AttnRes 缩小实验先冻结、后运行</b><p>三结构共享公共主干、初始化、窗口与优化器;只按三个 paired seed 和预注册 −0.010 BPC 阈值给出本协议内方向判断。</p></div>
|
<div><time>2026-07-30</time><b>AttnRes 缩小实验先冻结、后运行</b><p>三结构共享公共主干、初始化、窗口与优化器;只按三个 paired seed 和预注册 −0.010 BPC 阈值给出本协议内方向判断。</p></div>
|
||||||
<div><time>2026-07-30</time><b>支持结果与梯度反结果同时进入主视区</b><p>Full / Block 的最终 BPC 同向改善;核心参数 gradient RMS CV 却高于 Baseline,不换指标掩盖。</p></div>
|
<div><time>2026-07-30</time><b>支持结果与梯度反结果同时进入主视区</b><p>Full / Block 的最终 BPC 同向改善;核心参数 gradient RMS CV 却高于 Baseline,不换指标掩盖。</p></div>
|
||||||
<div><time>2026-07-30</time><b>正式重放不把 wall time 纳入 exact</b><p>Block / seed-1 的模型、优化器、曲线、历史、诊断和环境八字段 exact;计时受调度影响,单独报告。</p></div>
|
<div><time>2026-07-30</time><b>正式重放不把 wall time 纳入 exact</b><p>Block / seed-1 的模型、优化器、曲线、历史、诊断和环境八字段 exact;计时受调度影响,单独报告。</p></div>
|
||||||
|
<div><time>2026-07-30</time><b>Figure 5 的“梯度”不再靠猜测补合同</b><p>官方未公开 gradient tensor、norm、reduction 与统计代码;本站 activation-gradient 定义只叫 operationalization,不叫论文复画。</p></div>
|
||||||
|
<div><time>2026-07-30</time><b>首尾平衡与全层 CV 永久分账</b><p>Block 在 6/6 配对中改善 first/last,却因中后段局部尖峰让 CV 在 6/6 配对中恶化;联合判定保持 mixed。</p></div>
|
||||||
|
<div><time>2026-07-30</time><b>绝对梯度尺度必须与归一化谱同屏</b><p>Block mean gradient 约为 Baseline 的 54%–57%;更接近 1 的首尾比不能偷换成各层信号更强。</p></div>
|
||||||
|
<div><time>2026-07-30</time><b>32 层完整重放扩到状态哈希</b><p>8,000-step fresh replay 的全部冻结字段以及 model / optimizer state hashes exact;额外 replay bytes 单列,不混入 formal 预算。</p></div>
|
||||||
<div><time>2026-07-29</time><b>32-token 对照改为同源 16→24</b><p>TNEWS 只有 105/10,000 条达到 32 tokens,强行统一会落入约 1% 极端长尾;24-token eligibility 仍保留 1,609 条中文候选。</p></div>
|
<div><time>2026-07-29</time><b>32-token 对照改为同源 16→24</b><p>TNEWS 只有 105/10,000 条达到 32 tokens,强行统一会落入约 1% 极端长尾;24-token eligibility 仍保留 1,609 条中文候选。</p></div>
|
||||||
<div><time>2026-07-29</time><b>长度敏感性必须成对重采样</b><p>16-token 输入严格是 24-token 输入前缀,2,000 次 bootstrap 共用 prompt indices;结果只描述固定 cohort 的长度敏感性。</p></div>
|
<div><time>2026-07-29</time><b>长度敏感性必须成对重采样</b><p>16-token 输入严格是 24-token 输入前缀,2,000 次 bootstrap 共用 prompt indices;结果只描述固定 cohort 的长度敏感性。</p></div>
|
||||||
<div><time>2026-07-29</time><b>三类 cohort 永久分身份</b><p>自然长度回答本批样本如何路由;matched-16 / 24 回答同一 prompt 多看 8 tokens 后如何变化,不把二者混成内容因果。</p></div>
|
<div><time>2026-07-29</time><b>三类 cohort 永久分身份</b><p>自然长度回答本批样本如何路由;matched-16 / 24 回答同一 prompt 多看 8 tokens 后如何变化,不把二者混成内容因果。</p></div>
|
||||||
|
|||||||
Reference in New Issue
Block a user