Compare commits

...

5 Commits

Author SHA1 Message Date
wuyang f8f0712103 docs: record K3 Round 08 release 2026-07-30 18:46:12 +08:00
wuyang f177fa676d research: publish AttnRes forward training study 2026-07-30 18:41:18 +08:00
wuyang 7ea91caabb experiment: implement AttnRes forward training runner 2026-07-30 15:14:15 +08:00
wuyang d81aacfc68 research: preregister AttnRes forward training study 2026-07-30 15:03:31 +08:00
wuyang badf498598 docs: record K3 round seven release 2026-07-30 14:43:16 +08:00
54 changed files with 210594 additions and 37 deletions
+16 -3
View File
@@ -8,7 +8,7 @@
|---|---:|---:|---|
| 研究框架与规范 | 进行中 | 83% | Scaling Laws 二轮拟合复现与逐图精读 |
| 网站设计系统 | 进行中 | 89% | 打印样式与更多通用可视化组件 |
| Kimi K3 深读 | 七轮实证进行中 | 99% | 设计前向训练变体,并等待 `A_log` 社区候选的官方裁决 |
| Kimi K3 深读 | 八轮实证已收敛 | 100% | 稳定维护;真实 forward 等待 `A_log` 官方裁决 |
| 语言模型前史 | 完成首版 | 78% | Kneser–Ney、LSTM、Bahdanau 逐图精读与真实小语料复现 |
| Transformer 基础 | 完成首版 | 79% | 多头电路、归一化 traces 与真实 kernel / KV 配置 |
| 表示、位置与残差高速公路 | 完成首版 | 81% | 真实 hidden-state / norm traces、长上下文位置外推与深层稳定性消融 |
@@ -41,7 +41,7 @@
- [x] 完成 486 篇关键论文索引,覆盖 16 个标签专题与 Kimi/DeepSeek 聚光主线。
- [x] 完成可检索、可按专题筛选的论文库页面。
- [x] 完成 K3、语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全十七篇首版长文。
- [x] 完成 K3 三轴架构、八联报告实验、四联开放工件实验、Round 04 / 05 / 06 / 07 各五联 AttnRes 独立实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 二十二联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等一百零九个原创交互视图。
- [x] 完成 K3 三轴架构、八联报告实验、四联开放工件实验、Round 04 / 05 / 06 / 07 / 08 各五联 AttnRes 独立实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 二十二联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等一百一十四个原创交互视图。
- [x] 完成长上下文首版:五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。
- [x] 核验 FlashAttention、DeepSeek-V2/V3.2/V4、Kimi Linear/K3 等六份论文原文,并建立长上下文研究账本。
- [x] 核验 Switch、ST-MoE、DeepSeekMoE、Loss-Free、V3、LatentMoE 与 K3 原文,并建立 MoE 研究账本。
@@ -304,10 +304,17 @@
- [x] 次级控制显示 group 7 MLP-only branch 在 sufficiency 方向 6 / 6 通过 material + margin 门;group 6 MLP-only 为 5 / 6,不能宣布 dominance。output-only mean sufficiency 仅 `.131 / .187`、0 / 6 过 50%;all-depth 为 `.949 / 1.145`。所有 score 都是非加性 log-gap 诊断,不写成贡献率。
- [x] K3 Round 07 五视图实验室完成:65-node 路径图、14-mask 全矩阵、双向主门、branch/output controls 与 32 层原始谱/replay 审计可交互;protocol、scoping、Grok 结果前审阅、audit、runner、analyzer、packager、四个 raw JSON、aggregate、compact 与 reproduction 全部进入公开树。
- [x] Round 07 本地闸门通过:100 个 Astro 文件零诊断,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;Round 04/05/06/07 四套冻结数据、四套 K3 专项、K3 全量与全站 23 套真实 Chrome 回归通过。动态生成矩阵的 scoped CSS 退化由截图复查发现并修复;桌面/390px 移动端零文档级溢出、零 offender、零运行时异常。
- [x] K3 Round 07 以功能源提交 `dc8ec30`、不可变镜像 `20260730T063811Z-dc8ec30` 发布;OCI index digest `sha256:98d441412408774628195ce23f431d25034bb1590210c11230d884513898e71d`。NAS `12010→8080` healthy / 0 次重启、Compose Manager 标签、VPS→NAS、NPM host 31 / cert 41、DNS、HTTPS/2、gzip、immutable asset、首页/K3 公网内容、门户 `LLM ATLAS / projects / 180` 与全站 23 套生产 Chrome 回归全部通过;保留 Round 06 `20260730T042651Z-b393780` 回滚点。
- [x] K3 Round 08 在正式输出前冻结 `llm-atlas-k3-attnres-forward-training-v1`:四个 train-time forward variants × 三 seed × 8,000 steps、同 seed Round 05 historical pairing、step 8,000 的 layers 21–25 contrast / peak 双指标 20% 主门、BPC 每 seed / mean 护栏、描述性 `I67` 与完整 primary replay;不把架构消融写成 pure-forward 因果实验。
- [x] 12 formal + 1 replay 全部一次完成,每格 65,536,000 target bytes;新处理总量 851,968,000,历史 references 196,608,000 单列。联合 groups 6+7 的 contrast / peak 六格降幅为 62.1%–77.3% / 32.3%–62.0%,6 / 6 通过;BPC 三 seed 最大 `+.009598`、均值 `+.006578`,质量门 4 / 4 通过。
- [x] primary seed-1 replay scientific payload exact,SHA-256 为 `b85563ca…c051`;冻结主状态为 `forward_training_attenuation_established_within_reduced_protocol`。`I67` 的 step-8,000 mean 为 contrast `−.3670`、peak `−.1704`,只保留为跨独立训练的描述性 log residual。
- [x] 结果前 Grok 实现审阅指出 smoke-only empty selector 的空 census;矩阵结束后删除 early return、加入 `forward_calls > 0`,修补后的 learned wrapper 实际执行 39 次 forward,父/包装器 15 组科学字段仍 exact。结果后 Grok 只读复算报告 `blocking_errors=0`、status / replay 均确认。
- [x] K3 Round 08 五视图实验室完成:forward contract、六 checkpoint 训练轨迹、attenuation+BPC 主门、non-additivity map 与 32 层谱/replay audit 可交互;protocol、scoping、两阶段 Grok 审阅、runner、analyzer、13 raw、aggregate、compact、reproduction 与结果审计全部进入公开树。
- [x] Round 08 本地闸门通过:101 个 Astro 文件零诊断,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;Round 04–08 五套冻结数据、五套 K3 专项、K3 全量与全站 24 套真实 Chrome 回归通过。截图复查修复黑色实验室标题对比度;桌面/390px 移动端零文档级溢出、零 offender、零运行时异常。
- [x] K3 Round 08 以功能源提交 `f177fa6`、不可变镜像 `20260730T104228Z-f177fa6` 发布;OCI index digest `sha256:0b5b9d9cbf1b4538a37b01727ba1d16be13f991408d3d60e20bc76023b9008bb`。NAS `12010→8080` healthy / 0 次重启、Compose Manager、VPS→NAS、NPM host 31 / cert 41、DNS、HTTPS/2、gzip、首页/K3/进度页、门户 `LLM ATLAS / projects / 180` 与全站 24 套生产 Chrome 回归全部通过;保留 Round 07 `20260730T064331Z-badf498` 回滚点。
## 正在进行
- [ ] K3 七轮下一闸门:设计前向训练变体与非加性局部交互地图;真实 K3 forward 继续等待 `A_log [128]↔[96]` 社区候选的官方裁决或权重修订。
- [ ] DeepSeek 八轮下一闸门:推进干预式 mediation、SM90 FlashMLA、FP8 / pipeline traces 与 R1-like RL 小模型复现。
- [ ] 表示、位置与残差二轮:真实 hidden-state / norm traces、长上下文位置外推复现与 mHC / AttnRes 深层稳定性消融。
- [ ] 评测安全二轮:真实 cross-harness / pass@k 复跑、Judge 元评测、动态污染与过拒案例。
@@ -515,6 +522,12 @@
| 2026-07-30 | 局部路径必须同时通过 sufficiency 与 restoration | groups 6+7 的 sufficiency 6/6 过 `.50`,restoration 仅 contrast 3/3 通过、peak 0/3 通过;强单侧证据不升级为 localization |
| 2026-07-30 | 局部 score 不写成可加贡献率 | 同一 scope 在 learned 与 uniform 背景的响应不同,`S_peak > 1` 与负 interaction residual 都是非加性诊断,不是 170% 贡献或方差分解 |
| 2026-07-30 | branch 与 output 控制保持次级证据身份 | group 7 MLP-only 的 6/6 只属于 sufficiency branch gate;group 6 为 5/6,output-only 为 0/6,都不能补救失败的双向主门 |
| 2026-07-30 | K3 Round 07 局部路径里程碑发布 | 功能源 `dc8ec30`、镜像 `20260730T063811Z-dc8ec30`、OCI `sha256:98d44141…e71d`;21/21 公网页面链路与全站 23 套生产 Chrome 通过,保留 Round 06 回滚点 |
| 2026-07-30 | 训练期前向干预属于架构消融 | selected uniform mixer 同时改变 train/eval forward、natural backward 与后续 updates;不能写成只改 forward 的路径因果 |
| 2026-07-30 | Round 08 主门与质量门同时成立 | groups 6+7 六个 attenuation 格 6/6 ≥20%;三 seed ΔBPC mean `+.006578`,4/4 过闸;结论只限固定缩小协议 |
| 2026-07-30 | non-additivity 永久保留描述身份 | `I67` 来自三套独立训练,只是 cross-run log residual,不是因果 interaction、Shapley 或贡献率 |
| 2026-07-30 | K3 研究线在 Round 08 主动收敛 | 不启动 Round 09;公开 `A_log [128]↔[96]` 冲突继续等待官方裁决,现有五轮实证停在可复现、可回滚的稳定边界 |
| 2026-07-30 | K3 Round 08 训练期前向里程碑发布 | 功能源 `f177fa6`、镜像 `20260730T104228Z-f177fa6`、OCI `sha256:0b5b9d9c…008bb`;21/21 公网页面链路与全站 24 套生产 Chrome 通过,保留 Round 07 回滚点 |
## 未决问题
+70
View File
@@ -0,0 +1,70 @@
# Attention Residuals train-time forward intervention
This directory implements preregistered protocol
`llm-atlas-k3-attnres-forward-training-v1`:
- `research/K3_ATTNRES_FORWARD_TRAINING_SCOPING.md`
- `research/K3_ATTNRES_FORWARD_TRAINING_PROTOCOL.md`
- `research/K3_ATTNRES_FORWARD_TRAINING_GROK_REVIEW.md`
- `research/K3_ATTNRES_FORWARD_TRAINING_IMPLEMENTATION_REVIEW.md`
- `research/K3_ATTNRES_FORWARD_TRAINING_AUDIT.md`
It is a depth-32 reduced Block AttnRes architecture ablation. It is not a
Kimi-K3 checkpoint forward pass and does not claim to recover unpublished
Figure 5 telemetry.
## Frozen environment
```text
Python /home/wuyang/.pyenv/versions/3.10.14/envs/navi-router-cu128/bin/python
PyTorch 2.11.0+cu128
GPU NVIDIA GeForce RTX 5090
CUBLAS_WORKSPACE_CONFIG=:4096:8
maximum concurrency 2
```
## Pre-result gates
The checked-in gate artifacts must pass before formal output:
```bash
CUBLAS_WORKSPACE_CONFIG=:4096:8 \
/home/wuyang/.pyenv/versions/3.10.14/envs/navi-router-cu128/bin/python \
experiments/k3/attnres_forward/verify.py step-zero \
--cache-dir /home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1 \
--parent-manifest experiments/k3/attnres_gradient/manifest.json \
--study-manifest experiments/k3/attnres_forward/manifest.json \
--output experiments/k3/attnres_forward/results/gates/step-zero.json
```
`learned_reference` is smoke-only. Its 20-step result is compared with a
fresh parent Round 05 smoke using `verify.py smoke-compare`.
## Formal matrix
```bash
/home/wuyang/.pyenv/versions/3.10.14/envs/navi-router-cu128/bin/python \
experiments/k3/attnres_forward/run_matrix.py \
--python /home/wuyang/.pyenv/versions/3.10.14/envs/navi-router-cu128/bin/python \
--cache-dir /home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1 \
--parent-manifest experiments/k3/attnres_gradient/manifest.json \
--study-manifest experiments/k3/attnres_forward/manifest.json \
--output-dir experiments/k3/attnres_forward/results/raw \
--phase all \
--concurrency 2
```
This runs 12 formal cells and one full replay. The analyzer reads all cells,
the frozen historical paired references, and generates the only authoritative
status, interaction map, and website compact artifact.
The checked-in Round 08 release contains:
- 13 raw results under `results/raw/`;
- `reproduction.json` with the exact primary scientific-payload hash;
- aggregate / compact website data under `src/data/`;
- a frozen-data checker and real-Chrome five-view regression in `scripts/`.
The established status is deliberately scoped to this reduced protocol. It is
not a real Kimi-K3 checkpoint result or a reproduction of unpublished Figure
5(c) telemetry.
+699
View File
@@ -0,0 +1,699 @@
#!/usr/bin/env python3
"""Aggregate and gate preregistered Round 08 forward-training results."""
from __future__ import annotations
import argparse
import copy
import hashlib
import json
import math
import statistics
from pathlib import Path
from typing import Any, Iterable
PROTOCOL_ID = "llm-atlas-k3-attnres-forward-training-v1"
PARENT_PROTOCOL_ID = "llm-atlas-k3-attnres-gradient-scale-v1"
METRICS = ("spike_contrast", "peak_normalized")
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser()
parser.add_argument("--manifest", type=Path, required=True)
parser.add_argument("--formal", type=Path, action="append", required=True)
parser.add_argument("--replay", type=Path, required=True)
parser.add_argument("--reference-dir", type=Path, required=True)
parser.add_argument("--aggregate-output", type=Path, required=True)
parser.add_argument("--compact-output", type=Path, required=True)
parser.add_argument("--reproduction-output", type=Path, required=True)
return parser.parse_args()
def canonical_sha256(value: Any) -> str:
return hashlib.sha256(
json.dumps(
value, ensure_ascii=False, sort_keys=True, separators=(",", ":")
).encode()
).hexdigest()
def file_sha256(path: Path) -> str:
digest = hashlib.sha256()
with path.open("rb") as handle:
for chunk in iter(lambda: handle.read(1024 * 1024), b""):
digest.update(chunk)
return digest.hexdigest()
def mean(values: Iterable[float]) -> float:
return statistics.fmean(values)
def read_result(path: Path, expected_protocol: str) -> dict[str, Any]:
value = json.loads(path.read_text())
if value.get("protocol_id") != expected_protocol:
raise RuntimeError(f"protocol mismatch: {path}")
expected = value.get("canonical_sha256_without_self")
payload = {
key: item
for key, item in value.items()
if key != "canonical_sha256_without_self"
}
if not isinstance(expected, str) or canonical_sha256(payload) != expected:
raise RuntimeError(f"canonical self-hash mismatch: {path}")
return value
def exactly_one(values: list[dict[str, Any]], step: int) -> dict[str, Any]:
matches = [value for value in values if value["step"] == step]
if len(matches) != 1:
raise RuntimeError(f"step {step} missing or duplicated")
return matches[0]
def spectrum_metrics(
diagnostic: dict[str, Any],
spike_layers: tuple[int, ...],
epsilon: float,
) -> dict[str, Any]:
values = [
float(value)
for value in diagnostic["activation_grad_rms_by_block"]
]
if len(values) != 32:
raise RuntimeError("activation-gradient spectrum must have 32 layers")
if any(not math.isfinite(value) or value <= epsilon for value in values):
raise RuntimeError("activation-gradient spectrum is non-finite/non-positive")
spike_indices = {layer - 1 for layer in spike_layers}
spike_values = [
value for index, value in enumerate(values) if index in spike_indices
]
reference_values = [
value for index, value in enumerate(values) if index not in spike_indices
]
spike_mean = mean(spike_values)
reference_mean = mean(reference_values)
global_mean = mean(values)
contrast = spike_mean / reference_mean
peak = max(values) / global_mean
if any(
not math.isfinite(value) or value <= epsilon
for value in (spike_mean, reference_mean, contrast, peak)
):
raise RuntimeError("derived spike metric is non-finite/non-positive")
ordered = sorted(range(32), key=lambda index: (-values[index], index))
return {
"values": values,
"normalized": [value / global_mean for value in values],
"spike_mean": spike_mean,
"reference_mean": reference_mean,
"global_mean": global_mean,
"spike_contrast": contrast,
"peak_normalized": peak,
"peak_layer_1based": ordered[0] + 1,
"top_five_layers_1based": [index + 1 for index in ordered[:5]],
}
def final_bpc(value: dict[str, Any], step: int) -> float:
result = float(exactly_one(value["evaluations"], step)["bits_per_byte"])
if not math.isfinite(result):
raise RuntimeError("final validation BPC is non-finite")
return result
def stable_environment(value: dict[str, Any]) -> dict[str, Any]:
keys = (
"cublas_workspace_config",
"deterministic_algorithms",
"autocast",
"compile",
)
return {key: value["environment"][key] for key in keys}
def pairing_checks(
run: dict[str, Any], reference: dict[str, Any]
) -> dict[str, bool]:
manifest_fields = (
"formal_schedule_sha256",
"validation_tensor_sha256",
"diagnostic_tensor_sha256",
"input_gate_tensor_hashes",
)
checks = {
"seed": run["seed"] == reference["seed"],
"architecture": (
run["architecture"] == reference["architecture"] == "block"
),
"depth": run["depth"] == reference["depth"] == 32,
"steps": run["steps"] == reference["steps"] == 8000,
"batch_size": run["batch_size"] == reference["batch_size"] == 32,
"initial_public_parameters": (
run["hashes"]["initial_public_parameters"]
== reference["hashes"]["initial_public_parameters"]
),
"initial_mixer_parameters": (
run["hashes"]["initial_mixer_parameters"]
== reference["hashes"]["initial_mixer_parameters"]
),
"model_topology": run["model"] == reference["model"],
"optimizer_hyperparameters": (
run["optimizer"] == reference["optimizer"]
),
"scientific_environment": (
stable_environment(run) == stable_environment(reference)
),
}
for field in manifest_fields:
checks[f"manifest.{field}"] = (
run["manifest"][field] == reference["manifest"][field]
)
return checks
def scientific_replay_payload(value: dict[str, Any]) -> dict[str, Any]:
payload = copy.deepcopy(value)
for key in (
"run_kind",
"timing",
"canonical_sha256_without_self",
"parent_runner_canonical_sha256",
):
payload.pop(key, None)
payload["manifest"].pop("path", None)
payload["study_manifest"].pop("path", None)
payload["environment"] = stable_environment(value)
return payload
def quality_gate(
variant_runs: dict[int, dict[str, Any]],
references: dict[int, dict[str, Any]],
*,
step: int,
per_seed_maximum: float,
mean_maximum: float,
) -> dict[str, Any]:
per_seed = {}
for seed, run in sorted(variant_runs.items()):
variant_bpc = final_bpc(run, step)
reference_bpc = final_bpc(references[seed], step)
delta = variant_bpc - reference_bpc
per_seed[str(seed)] = {
"variant_bpc": variant_bpc,
"reference_bpc": reference_bpc,
"delta_bpc": delta,
"passed": delta <= per_seed_maximum,
}
mean_delta = mean(item["delta_bpc"] for item in per_seed.values())
per_seed_passed = all(item["passed"] for item in per_seed.values())
mean_passed = mean_delta <= mean_maximum
return {
"passed": per_seed_passed and mean_passed,
"passed_checks": (
sum(item["passed"] for item in per_seed.values())
+ int(mean_passed)
),
"required_checks": 4,
"per_seed_maximum": per_seed_maximum,
"mean_maximum": mean_maximum,
"mean_delta_bpc": mean_delta,
"mean_passed": mean_passed,
"per_seed": per_seed,
}
def variant_effect(
variant: str,
runs: dict[int, dict[str, Any]],
references: dict[int, dict[str, Any]],
metrics_by_cell: dict[tuple[str, int, int], dict[str, Any]],
*,
step: int,
threshold: float,
quality: dict[str, Any],
) -> dict[str, Any]:
cells = []
for seed in sorted(runs):
candidate = metrics_by_cell[(variant, seed, step)]
reference = metrics_by_cell[("learned_reference", seed, step)]
for metric in METRICS:
reference_value = reference[metric]
candidate_value = candidate[metric]
relative_drop = (
reference_value - candidate_value
) / reference_value
cells.append(
{
"seed": seed,
"metric": metric,
"reference": reference_value,
"variant": candidate_value,
"relative_drop": relative_drop,
"passed": relative_drop >= threshold,
}
)
attenuation_passed = all(cell["passed"] for cell in cells)
return {
"variant": variant,
"threshold": threshold,
"passed_cells": sum(cell["passed"] for cell in cells),
"required_cells": len(cells),
"attenuation_passed": attenuation_passed,
"quality": quality,
"material_response_passed": (
attenuation_passed and quality["passed"]
),
"cells": cells,
}
def interaction_map(
metrics_by_cell: dict[tuple[str, int, int], dict[str, Any]],
seeds: tuple[int, ...],
steps: tuple[int, ...],
) -> dict[str, Any]:
cells = []
for step in steps:
for seed in seeds:
reference = metrics_by_cell[
("learned_reference", seed, step)
]
group6 = metrics_by_cell[
("uniform_group_6_forward", seed, step)
]
group7 = metrics_by_cell[
("uniform_group_7_forward", seed, step)
]
joint = metrics_by_cell[
("uniform_groups_6_7_forward", seed, step)
]
for metric in METRICS:
ref = reference[metric]
effects = {
"group6": math.log(ref / group6[metric]),
"group7": math.log(ref / group7[metric]),
"groups6_7": math.log(ref / joint[metric]),
}
residual = (
effects["groups6_7"]
- effects["group6"]
- effects["group7"]
)
cells.append(
{
"step": step,
"seed": seed,
"metric": metric,
"log_effects": effects,
"interaction_residual": residual,
"relative_drops": {
"group6": (ref - group6[metric]) / ref,
"group7": (ref - group7[metric]) / ref,
"groups6_7": (ref - joint[metric]) / ref,
},
}
)
summaries = []
for step in steps:
for metric in METRICS:
selected = [
cell
for cell in cells
if cell["step"] == step and cell["metric"] == metric
]
residuals = [
cell["interaction_residual"] for cell in selected
]
summaries.append(
{
"step": step,
"metric": metric,
"mean_interaction_residual": mean(residuals),
"minimum": min(residuals),
"maximum": max(residuals),
}
)
return {
"definition": "I67=ln(Xref/X67)-ln(Xref/X6)-ln(Xref/X7)",
"interpretation": (
"descriptive cross-run log-attenuation residual from three "
"independently trained variants; not a causal interaction"
),
"cells": cells,
"summaries": summaries,
}
def environment_metadata(value: dict[str, Any]) -> dict[str, Any]:
return {
key: value["environment"].get(key)
for key in ("gpu", "torch", "cuda", "compute_capability")
}
def write_hashed(path: Path, value: dict[str, Any]) -> None:
value["canonical_sha256_without_self"] = canonical_sha256(value)
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(
json.dumps(value, ensure_ascii=False, indent=2, sort_keys=True) + "\n"
)
def main() -> None:
args = parse_args()
manifest = json.loads(args.manifest.read_text())
if (
manifest["protocol_id"] != PROTOCOL_ID
or manifest["status"] != "frozen-before-model-output"
):
raise RuntimeError("manifest is not the frozen Round 08 contract")
variants = tuple(manifest["variants"].keys())
seeds = tuple(manifest["formal_seeds"])
steps = tuple(manifest["diagnostic_steps"])
primary_step = manifest["primary_step"]
epsilon = manifest["thresholds"]["positive_denominator_epsilon"]
spike_layers = tuple(manifest["spike_layers_1based"])
expected_cells = {(variant, seed) for variant in variants for seed in seeds}
if len(args.formal) != len(expected_cells):
raise RuntimeError("formal path count does not match the 4×3 matrix")
runs: dict[tuple[str, int], dict[str, Any]] = {}
run_paths: dict[tuple[str, int], Path] = {}
pairing: dict[str, Any] = {}
references: dict[int, dict[str, Any]] = {}
reference_paths: dict[int, Path] = {}
for seed in seeds:
path = args.reference_dir / (
f"formal-depth-32-block-seed-{seed}.json"
)
references[seed] = read_result(path, PARENT_PROTOCOL_ID)
reference_paths[seed] = path
for path in args.formal:
value = read_result(path, PROTOCOL_ID)
identity = (value["variant"], value["seed"])
if identity in runs:
raise RuntimeError(f"duplicate formal cell: {identity}")
if (
value["run_kind"] != "formal"
or value["steps"] != manifest["formal_steps"]
or not value["forward_intervention"]["passed"]
):
raise RuntimeError(f"invalid formal cell: {path}")
runs[identity] = value
run_paths[identity] = path
if set(runs) != expected_cells:
raise RuntimeError("formal matrix identities do not match manifest")
for (variant, seed), value in sorted(runs.items()):
checks = pairing_checks(value, references[seed])
if not all(checks.values()):
raise RuntimeError(
f"historical reference pairing failed: "
f"{variant}/{seed}: {checks}"
)
pairing[f"{variant}:{seed}"] = {
"passed": True,
"checks": checks,
"run_environment": environment_metadata(value),
"reference_environment": environment_metadata(references[seed]),
"metadata_equal": (
environment_metadata(value)
== environment_metadata(references[seed])
),
}
metadata_warnings = [
{
"cell": cell,
"message": (
"GPU/version metadata differs from the historical paired "
"reference; frozen scientific-environment fields still match"
),
"run_environment": item["run_environment"],
"reference_environment": item["reference_environment"],
}
for cell, item in pairing.items()
if not item["metadata_equal"]
]
replay = read_result(args.replay, PROTOCOL_ID)
replay_contract = manifest["replay"]
if (
replay["run_kind"] != "replay"
or replay["variant"] != replay_contract["variant"]
or replay["seed"] != replay_contract["seed"]
or replay["steps"] != manifest["formal_steps"]
or not replay["forward_intervention"]["passed"]
):
raise RuntimeError("invalid replay identity/audit")
formal_primary = runs[
(replay_contract["variant"], replay_contract["seed"])
]
formal_payload = scientific_replay_payload(formal_primary)
replay_payload = scientific_replay_payload(replay)
replay_exact = formal_payload == replay_payload
if not replay_exact:
raise RuntimeError("primary formal/replay scientific payload mismatch")
metrics_by_cell: dict[tuple[str, int, int], dict[str, Any]] = {}
for seed, reference in references.items():
for step in steps:
metrics_by_cell[("learned_reference", seed, step)] = (
spectrum_metrics(
exactly_one(reference["diagnostics"], step),
spike_layers,
epsilon,
)
)
for (variant, seed), value in runs.items():
if tuple(item["step"] for item in value["diagnostics"]) != steps:
raise RuntimeError(f"diagnostic schedule drift: {variant}/{seed}")
for step in steps:
metrics_by_cell[(variant, seed, step)] = spectrum_metrics(
exactly_one(value["diagnostics"], step),
spike_layers,
epsilon,
)
runs_by_variant = {
variant: {seed: runs[(variant, seed)] for seed in seeds}
for variant in variants
}
qualities = {
variant: quality_gate(
variant_runs,
references,
step=primary_step,
per_seed_maximum=manifest["thresholds"][
"final_bpc_delta_per_seed_maximum"
],
mean_maximum=manifest["thresholds"][
"final_bpc_delta_mean_maximum"
],
)
for variant, variant_runs in runs_by_variant.items()
}
effects = {
variant: variant_effect(
variant,
variant_runs,
references,
metrics_by_cell,
step=primary_step,
threshold=manifest["thresholds"]["material_relative_drop"],
quality=qualities[variant],
)
for variant, variant_runs in runs_by_variant.items()
}
primary = effects[manifest["primary_variant"]]
if primary["attenuation_passed"] and primary["quality"]["passed"]:
status = (
"forward_training_attenuation_established_within_reduced_protocol"
)
elif primary["attenuation_passed"]:
status = "quality_guard_failed"
elif primary["quality"]["passed"]:
status = "attenuation_not_established"
else:
status = "attenuation_and_quality_failed"
secondary = {
variant: (
"secondary_material_response"
if effect["material_response_passed"]
else "secondary_response_not_established"
)
for variant, effect in effects.items()
if variant != manifest["primary_variant"]
}
interaction = interaction_map(metrics_by_cell, seeds, steps)
trajectories = []
final_spectra = []
for variant in ("learned_reference",) + variants:
for seed in seeds:
for step in steps:
record = metrics_by_cell[(variant, seed, step)]
reference = metrics_by_cell[
("learned_reference", seed, step)
]
trajectories.append(
{
"variant": variant,
"seed": seed,
"step": step,
"spike_mean": record["spike_mean"],
"reference_mean": record["reference_mean"],
"spike_contrast": record["spike_contrast"],
"peak_normalized": record["peak_normalized"],
"relative_drop": {
metric: (
reference[metric] - record[metric]
)
/ reference[metric]
for metric in METRICS
},
}
)
final = metrics_by_cell[(variant, seed, primary_step)]
final_spectra.append(
{
"variant": variant,
"seed": seed,
**final,
}
)
input_files = {
"manifest": {
"path": str(args.manifest),
"sha256": file_sha256(args.manifest),
},
"formal": [
{
"variant": variant,
"seed": seed,
"path": str(run_paths[(variant, seed)]),
"sha256": file_sha256(run_paths[(variant, seed)]),
}
for variant, seed in sorted(runs)
],
"references": [
{
"seed": seed,
"path": str(reference_paths[seed]),
"sha256": file_sha256(reference_paths[seed]),
}
for seed in seeds
],
"replay": {
"path": str(args.replay),
"sha256": file_sha256(args.replay),
},
}
aggregate = {
"schema_version": 1,
"protocol_id": PROTOCOL_ID,
"status": status,
"scope": (
"depth-32 reduced Block AttnRes train-time architecture "
"ablation; not a real Kimi-K3 checkpoint result"
),
"primary_step": primary_step,
"spike_layers_1based": list(spike_layers),
"thresholds": manifest["thresholds"],
"input_files": input_files,
"historical_pairing": pairing,
"metadata_warnings": metadata_warnings,
"replay": {
"passed": replay_exact,
"scientific_payload_sha256": canonical_sha256(formal_payload),
"excluded": [
"run_kind",
"timing",
"self hashes",
"manifest path strings",
"GPU/version metadata",
],
},
"primary": primary,
"secondary_status": secondary,
"effects": effects,
"interaction": interaction,
"trajectories": trajectories,
"final_spectra": final_spectra,
"processed_target_bytes": manifest["new_target_bytes"],
"historical_reference_target_bytes": (
manifest["historical_reference_target_bytes"]
),
"reporting_boundary": (
"C can change through spike-window numerator and the 27-layer "
"reference denominator; layers 26-28 are intervened but belong "
"to the denominator."
),
}
write_hashed(args.aggregate_output, aggregate)
compact = {
"schema_version": 1,
"protocol_id": PROTOCOL_ID,
"status": status,
"primary_step": primary_step,
"spike_layers_1based": list(spike_layers),
"thresholds": manifest["thresholds"],
"primary": primary,
"secondary_status": secondary,
"effects": effects,
"interaction": interaction,
"trajectories": trajectories,
"final_spectra": final_spectra,
"replay": aggregate["replay"],
"metadata_warnings": metadata_warnings,
"processed_target_bytes": manifest["new_target_bytes"],
"reporting_boundary": aggregate["reporting_boundary"],
"aggregate_sha256": aggregate["canonical_sha256_without_self"],
}
write_hashed(args.compact_output, compact)
reproduction = {
"schema_version": 1,
"protocol_id": PROTOCOL_ID,
"passed": replay_exact,
"formal_variant": replay_contract["variant"],
"seed": replay_contract["seed"],
"formal_file_sha256": file_sha256(
run_paths[
(replay_contract["variant"], replay_contract["seed"])
]
),
"replay_file_sha256": file_sha256(args.replay),
"scientific_payload_sha256": canonical_sha256(formal_payload),
"excluded_fields": aggregate["replay"]["excluded"],
}
write_hashed(args.reproduction_output, reproduction)
print(
json.dumps(
{
"status": status,
"primary_attenuation": {
"passed_cells": primary["passed_cells"],
"required_cells": primary["required_cells"],
},
"primary_quality": {
"passed_checks": primary["quality"]["passed_checks"],
"required_checks": primary["quality"]["required_checks"],
},
"replay_exact": replay_exact,
"aggregate": str(args.aggregate_output),
"compact": str(args.compact_output),
},
ensure_ascii=False,
indent=2,
)
)
if __name__ == "__main__":
main()
@@ -0,0 +1,89 @@
{
"schema_version": 1,
"protocol_id": "llm-atlas-k3-attnres-forward-training-v1",
"status": "frozen-before-model-output",
"parent_protocol_id": "llm-atlas-k3-attnres-gradient-scale-v1",
"parent_manifest": "experiments/k3/attnres_gradient/manifest.json",
"architecture": "block",
"depth": 32,
"formal_steps": 8000,
"smoke_steps": 20,
"batch_size": 32,
"formal_seeds": [
2026073001,
2026073002,
2026073003
],
"diagnostic_steps": [
0,
100,
500,
2000,
4000,
8000
],
"primary_step": 8000,
"spike_layers_1based": [21, 22, 23, 24, 25],
"historical_reference": {
"directory": "experiments/k3/attnres_gradient/results/raw",
"filename_template": "formal-depth-32-block-seed-{seed}.json",
"identity": "historical-paired-reference-not-contemporaneous-randomized-control"
},
"variants": {
"uniform_group_6_forward": {
"selected_depth_indices": [40, 41, 42, 43, 44, 45, 46, 47]
},
"uniform_group_7_forward": {
"selected_depth_indices": [48, 49, 50, 51, 52, 53, 54, 55]
},
"uniform_groups_6_7_forward": {
"selected_depth_indices": [40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55]
},
"uniform_group_7_mlp_forward": {
"selected_depth_indices": [49, 51, 53, 55]
}
},
"smoke_only_variants": [
"learned_reference"
],
"primary_variant": "uniform_groups_6_7_forward",
"replay": {
"variant": "uniform_groups_6_7_forward",
"seed": 2026073001
},
"selected_source_counts": {
"40": 6,
"41": 7,
"42": 7,
"43": 7,
"44": 7,
"45": 7,
"46": 7,
"47": 7,
"48": 7,
"49": 8,
"50": 8,
"51": 8,
"52": 8,
"53": 8,
"54": 8,
"55": 8
},
"thresholds": {
"positive_denominator_epsilon": 1e-30,
"material_relative_drop": 0.2,
"final_bpc_delta_per_seed_maximum": 0.05,
"final_bpc_delta_mean_maximum": 0.03,
"uniform_weight_max_abs_error": 1e-12,
"loss_scale_ratio_abs_error": 1e-5,
"loss_scale_shape_abs_error": 1e-6
},
"new_target_bytes": {
"per_cell": 65536000,
"formal_12_cells": 786432000,
"primary_replay": 65536000,
"total": 851968000
},
"historical_reference_target_bytes": 196608000,
"concurrency_maximum": 2
}
@@ -0,0 +1,25 @@
{
"canonical_sha256_without_self": "57346df80c0d76bd1d306d5fa213feed74ac2094237c22ebcb49f16ea16437e0",
"excluded_fields": [
"run_kind",
"timing",
"self hashes",
"manifest path strings",
"GPU/version metadata"
],
"formal_file_sha256": "0962ebd1a00a11e61ac795282bfa99412166c8752f2a3731f137030d7f134dc1",
"formal_variant": "uniform_groups_6_7_forward",
"passed": true,
"post_result_grok_review": {
"blocking_errors": 0,
"claim_boundary_confirmed": true,
"replay_confirmed": true,
"session_id": "019fb28b-a9e1-7643-8e43-06f5e16a2077",
"status_confirmed": true
},
"protocol_id": "llm-atlas-k3-attnres-forward-training-v1",
"replay_file_sha256": "b85ac8062b2b0b8b3f2305d4b22a7c0212466fbfb28a9d63099cadd0a6917c8e",
"schema_version": 1,
"scientific_payload_sha256": "b85563ca5cb53e60b39c3801d372376206105b8a089a8633b3e81973a7f0c051",
"seed": 2026073001
}
@@ -0,0 +1,39 @@
{
"canonical_sha256_without_self": "99aeefc0ba35c199725ed7af377450ecbb5cc5f1aabdda9ac2345a30f03b444f",
"excluded_fields": [
"protocol wrapper fields",
"timing",
"self hash",
"parent runner self hash",
"study manifest"
],
"field_checks": {
"architecture": true,
"batch_size": true,
"depth": true,
"diagnostics": true,
"environment": true,
"evaluations": true,
"gradient_gate": true,
"hashes": true,
"manifest": true,
"model": true,
"optimizer": true,
"seed": true,
"steps": true,
"target_bytes_seen": true,
"training_history": true
},
"gate": "empty-selector-parent-equivalence",
"parent_file": "/home/wuyang/Code/K3/experiments/k3/attnres_forward/results/gates/parent-smoke.json",
"passed": true,
"protocol_id": "llm-atlas-k3-attnres-forward-training-v1",
"schema_version": 1,
"wrapper_file": "/home/wuyang/Code/K3/experiments/k3/attnres_forward/results/gates/wrapper-smoke.json",
"wrapper_identity": {
"forward_audit": true,
"parent_protocol": true,
"protocol": true,
"variant": true
}
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,220 @@
#!/usr/bin/env python3
"""Run the frozen Round 08 matrix with at most two isolated processes."""
from __future__ import annotations
import argparse
import json
import os
import subprocess
import time
from pathlib import Path
from typing import Any
VARIANTS = (
"uniform_group_6_forward",
"uniform_group_7_forward",
"uniform_groups_6_7_forward",
"uniform_group_7_mlp_forward",
)
SEEDS = (2026073001, 2026073002, 2026073003)
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser()
parser.add_argument("--python", type=Path, required=True)
parser.add_argument("--cache-dir", type=Path, required=True)
parser.add_argument("--parent-manifest", type=Path, required=True)
parser.add_argument("--study-manifest", type=Path, required=True)
parser.add_argument("--output-dir", type=Path, required=True)
parser.add_argument(
"--phase", choices=("formal", "replay", "all"), default="all"
)
parser.add_argument("--concurrency", type=int, default=2)
return parser.parse_args()
def cell_output(
output_dir: Path, variant: str, seed: int, run_kind: str
) -> Path:
return output_dir / (
f"{run_kind}-{variant}-seed-{seed}.json"
)
def command_for(
args: argparse.Namespace, variant: str, seed: int, run_kind: str
) -> list[str]:
runner = Path(__file__).resolve().parent / "train.py"
return [
str(args.python),
str(runner),
"--variant",
variant,
"--study-manifest",
str(args.study_manifest),
"--run-kind",
run_kind,
"--architecture",
"block",
"--depth",
"32",
"--seed",
str(seed),
"--cache-dir",
str(args.cache_dir),
"--manifest",
str(args.parent_manifest),
"--output",
str(cell_output(args.output_dir, variant, seed, run_kind)),
]
def validate_manifest(args: argparse.Namespace) -> None:
manifest = json.loads(args.study_manifest.read_text())
if (
manifest["status"] != "frozen-before-model-output"
or tuple(manifest["variants"]) != VARIANTS
or tuple(manifest["formal_seeds"]) != SEEDS
or manifest["concurrency_maximum"] != 2
):
raise RuntimeError("study manifest matrix/concurrency drift")
if args.concurrency < 1 or args.concurrency > 2:
raise ValueError("the frozen protocol permits one or two processes")
def stop_processes(items: list[dict[str, Any]]) -> None:
for item in items:
if item["process"].poll() is None:
item["process"].terminate()
for item in items:
process = item["process"]
if process.poll() is not None:
continue
try:
process.wait(timeout=10)
except subprocess.TimeoutExpired:
process.kill()
process.wait()
def quarantine_failed_output(
output_dir: Path, variant: str, seed: int, run_kind: str
) -> str | None:
output = cell_output(output_dir, variant, seed, run_kind)
if not output.exists():
return None
failed = output.with_suffix(".failed.json")
if failed.exists():
failed = output.with_suffix(f".failed-{time.time_ns()}.json")
output.replace(failed)
return str(failed)
def run_cells(
args: argparse.Namespace,
cells: list[tuple[str, int, str]],
) -> None:
args.output_dir.mkdir(parents=True, exist_ok=True)
for variant, seed, run_kind in cells:
output = cell_output(args.output_dir, variant, seed, run_kind)
if output.exists():
raise FileExistsError(
f"refusing to overwrite existing result: {output}"
)
environment = dict(os.environ)
environment["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
pending = list(cells)
running: list[dict[str, Any]] = []
completed = 0
while pending or running:
while pending and len(running) < args.concurrency:
variant, seed, run_kind = pending.pop(0)
command = command_for(args, variant, seed, run_kind)
process = subprocess.Popen(command, env=environment)
running.append(
{
"identity": (variant, seed, run_kind),
"process": process,
"started": time.monotonic(),
}
)
print(
json.dumps(
{
"event": "cell_started",
"variant": variant,
"seed": seed,
"run_kind": run_kind,
"pid": process.pid,
"active": len(running),
"remaining": len(pending),
},
sort_keys=True,
),
flush=True,
)
time.sleep(1)
survivors = []
for item in running:
return_code = item["process"].poll()
if return_code is None:
survivors.append(item)
continue
variant, seed, run_kind = item["identity"]
elapsed = time.monotonic() - item["started"]
if return_code != 0:
stop_processes(
[candidate for candidate in running if candidate is not item]
)
quarantined = quarantine_failed_output(
args.output_dir, variant, seed, run_kind
)
raise RuntimeError(
f"cell failed: {variant}/{seed}/{run_kind}: {return_code}; "
f"quarantined_output={quarantined}"
)
completed += 1
print(
json.dumps(
{
"event": "cell_completed",
"variant": variant,
"seed": seed,
"run_kind": run_kind,
"elapsed_seconds": elapsed,
"completed": completed,
"total": len(cells),
},
sort_keys=True,
),
flush=True,
)
running = survivors
def main() -> None:
args = parse_args()
validate_manifest(args)
formal = [
(variant, seed, "formal")
for variant in VARIANTS
for seed in SEEDS
]
replay = [
("uniform_groups_6_7_forward", 2026073001, "replay")
]
cells = (
formal
if args.phase == "formal"
else replay
if args.phase == "replay"
else formal + replay
)
run_cells(args, cells)
if __name__ == "__main__":
main()
+449
View File
@@ -0,0 +1,449 @@
#!/usr/bin/env python3
"""Run one preregistered Round 08 train-time uniform-forward cell."""
from __future__ import annotations
import importlib.util
import json
import math
import os
import sys
from pathlib import Path
from typing import Any
import torch
import torch.nn.functional as F
PROTOCOL_ID = "llm-atlas-k3-attnres-forward-training-v1"
PARENT_PROTOCOL_ID = "llm-atlas-k3-attnres-gradient-scale-v1"
VARIANTS = {
"learned_reference": (),
"uniform_group_6_forward": tuple(range(40, 48)),
"uniform_group_7_forward": tuple(range(48, 56)),
"uniform_groups_6_7_forward": tuple(range(40, 56)),
"uniform_group_7_mlp_forward": (49, 51, 53, 55),
}
FORMAL_VARIANTS = tuple(name for name in VARIANTS if name != "learned_reference")
EXPECTED_SOURCE_COUNTS = {
**{40: 6},
**{index: 7 for index in range(41, 49)},
**{index: 8 for index in range(49, 56)},
}
def load_parent_module() -> Any:
path = Path(__file__).resolve().parents[1] / "attnres_gradient" / "train.py"
spec = importlib.util.spec_from_file_location("k3_attnres_round05_train", path)
if spec is None or spec.loader is None:
raise RuntimeError(f"cannot import Round 05 runner from {path}")
module = importlib.util.module_from_spec(spec)
sys.modules[spec.name] = module
spec.loader.exec_module(module)
return module
parent = load_parent_module()
ACTIVE_VARIANT = "learned_reference"
LAST_MODEL: ForwardInterventionLanguageModel | None = None
LAST_OPTIMIZER: torch.optim.Optimizer | None = None
def extract_wrapper_argument(name: str) -> str:
try:
index = sys.argv.index(name)
except ValueError as error:
raise ValueError(f"missing required wrapper argument: {name}") from error
if index + 1 >= len(sys.argv):
raise ValueError(f"missing value for wrapper argument: {name}")
value = sys.argv[index + 1]
del sys.argv[index : index + 2]
return value
def argument_value(name: str, default: str | None = None) -> str | None:
try:
index = sys.argv.index(name)
except ValueError:
return default
if index + 1 >= len(sys.argv):
raise ValueError(f"missing value for argument: {name}")
return sys.argv[index + 1]
def parameter_names_for_indices(indices: tuple[int, ...]) -> tuple[str, ...]:
names = []
for index in indices:
names.extend(
(
f"mixers.{index}.query",
f"mixers.{index}.key_norm.weight",
)
)
return tuple(names)
class ForwardInterventionLanguageModel(parent.GradientLanguageModel):
"""Round 05 model with one frozen selector and parameter-free uniform mixers."""
def __init__(self, architecture: str):
super().__init__(architecture)
global LAST_MODEL
if architecture != "block":
raise ValueError("Round 08 only permits the block architecture")
if ACTIVE_VARIANT not in VARIANTS:
raise ValueError(f"unknown Round 08 variant: {ACTIVE_VARIANT}")
self.forward_variant = ACTIVE_VARIANT
self.selected_indices = tuple(VARIANTS[ACTIVE_VARIANT])
self.selected_set = frozenset(self.selected_indices)
self.forward_calls = 0
self.depth_visits = [0] * len(self.mixers)
self.output_visits = 0
self.source_counts: dict[int, set[int]] = {
index: set() for index in range(len(self.mixers))
}
self.uniform_weight_max_abs_error = 0.0
selected_names = parameter_names_for_indices(self.selected_indices)
named_parameters = dict(self.named_parameters())
self.selected_initial_tensors = {
name: named_parameters[name].detach().cpu().clone()
for name in selected_names
}
self.gradient_hook_calls = {
name: 0
for name in named_parameters
if name.startswith("mixers.") or name.startswith("output_mixer.")
}
self._gradient_hooks = []
for name, parameter in named_parameters.items():
if name not in self.gradient_hook_calls:
continue
def count_hook(
gradient: torch.Tensor, *, parameter_name: str = name
) -> torch.Tensor:
self.gradient_hook_calls[parameter_name] += 1
return gradient
self._gradient_hooks.append(parameter.register_hook(count_hook))
LAST_MODEL = self
def mix(
self,
mixer_index: int,
sources: list[torch.Tensor],
capture: bool,
) -> tuple[torch.Tensor, dict[str, Any] | None]:
self.depth_visits[mixer_index] += 1
self.source_counts[mixer_index].add(len(sources))
if mixer_index not in self.selected_set:
return self.mixers[mixer_index](sources, capture)
values = torch.stack(sources, dim=0)
logits = torch.zeros(
values.shape[0],
values.shape[1],
values.shape[2],
dtype=torch.float32,
device=values.device,
)
weights = torch.softmax(logits, dim=0)
expected = torch.tensor(
1.0 / len(sources), dtype=weights.dtype, device=weights.device
)
error = (weights - expected).abs().max().detach().cpu().item()
self.uniform_weight_max_abs_error = max(
self.uniform_weight_max_abs_error, error
)
output = torch.einsum(
"nbt,nbtd->btd", weights, values.float()
).to(values.dtype)
if not capture:
return output, None
entropy = -(weights * torch.log(weights.clamp_min(1e-30))).sum(dim=0)
return output, {
"mean_weights": weights.mean(dim=(1, 2)).detach().cpu().tolist(),
"entropy_mean": entropy.mean().detach().cpu().item(),
"sources": len(sources),
}
def forward(
self, input_ids: torch.Tensor, capture: bool = False
) -> tuple[torch.Tensor, parent.ActivationTrace | None]:
self.forward_calls += 1
embedded = self.embed(input_ids)
trace = parent.ActivationTrace([], [], [], [], []) if capture else None
completed = [embedded]
partial: torch.Tensor | None = None
mixer_index = 0
for block in self.blocks:
for branch_index in range(2):
sources = completed + ([] if partial is None else [partial])
branch_input, weights = self.mix(
mixer_index, sources, capture
)
mixer_index += 1
if branch_index == 0:
branch_output = block.attention(
block.attention_norm(branch_input)
)
else:
branch_output = block.mlp(block.mlp_norm(branch_input))
branch_for_residual = branch_output.float()
partial = (
branch_for_residual
if partial is None
else partial + branch_for_residual
)
if trace is not None:
trace.layer_input_rms.append(parent.rms(branch_input))
trace.branch_output_rms.append(parent.rms(branch_output))
trace.stream_state_rms.append(parent.rms(partial))
trace.depth_weights.append(weights or {})
if branch_index == 1:
partial.retain_grad()
trace.block_outputs.append(partial)
if mixer_index % parent.round04.SUBLAYERS_PER_BLOCK == 0:
completed.append(partial)
partial = None
if partial is not None or len(completed) != parent.BLOCK_GROUPS + 1:
raise RuntimeError("Round 08 Block AttnRes aggregation failed")
if self.output_mixer is None:
raise RuntimeError("Round 08 output mixer missing")
self.output_visits += 1
hidden, output_weights = self.output_mixer(completed, capture)
if trace is not None:
trace.output_weights = output_weights
normalized = self.final_norm(hidden)
logits = F.linear(normalized, self.token_embedding.weight)
return logits, trace
def tensor_exact(left: torch.Tensor, right: torch.Tensor) -> bool:
return (
left.dtype == right.dtype
and tuple(left.shape) == tuple(right.shape)
and torch.equal(left.detach().cpu(), right.detach().cpu())
)
def build_intervention_audit(
model: ForwardInterventionLanguageModel,
optimizer: torch.optim.Optimizer,
study_manifest: dict[str, Any],
) -> dict[str, Any]:
selected = tuple(model.selected_indices)
selected_names = set(parameter_names_for_indices(selected))
mixer_parameters = {
name: parameter
for name, parameter in model.named_parameters()
if name.startswith("mixers.") or name.startswith("output_mixer.")
}
optimizer_parameters = {
parameter
for group in optimizer.param_groups
for parameter in group["params"]
}
selected_parameter_checks = {}
for name in sorted(selected_names):
parameter = mixer_parameters[name]
selected_parameter_checks[name] = {
"gradient_hook_calls": model.gradient_hook_calls[name],
"in_optimizer_param_group": parameter in optimizer_parameters,
"optimizer_state_present": parameter in optimizer.state,
"final_equals_initial": tensor_exact(
parameter, model.selected_initial_tensors[name]
),
}
unselected_parameter_checks = {}
for name, parameter in sorted(mixer_parameters.items()):
if name in selected_names:
continue
unselected_parameter_checks[name] = {
"gradient_hook_calls": model.gradient_hook_calls[name],
"in_optimizer_param_group": parameter in optimizer_parameters,
"optimizer_state_present": parameter in optimizer.state,
}
source_counts = {
str(index): sorted(values)
for index, values in model.source_counts.items()
}
selected_source_gate = {
str(index): (
source_counts[str(index)]
== [study_manifest["selected_source_counts"][str(index)]]
== [EXPECTED_SOURCE_COUNTS[index]]
)
for index in selected
}
visit_gate = (
model.forward_calls > 0
and all(value == model.forward_calls for value in model.depth_visits)
and model.output_visits == model.forward_calls
)
selected_parameter_gate = all(
check["gradient_hook_calls"] == 0
and check["in_optimizer_param_group"]
and not check["optimizer_state_present"]
and check["final_equals_initial"]
for check in selected_parameter_checks.values()
)
unselected_parameter_gate = all(
check["gradient_hook_calls"] > 0
and check["in_optimizer_param_group"]
and check["optimizer_state_present"]
for check in unselected_parameter_checks.values()
)
expected_selected = tuple(
study_manifest["variants"]
.get(model.forward_variant, {"selected_depth_indices": []})[
"selected_depth_indices"
]
)
selector_gate = (
selected == expected_selected
and 64 not in selected
and selected_source_gate == {
str(index): True for index in selected
}
)
threshold = study_manifest["thresholds"][
"uniform_weight_max_abs_error"
]
uniform_gate = model.uniform_weight_max_abs_error <= threshold
passed = (
visit_gate
and selector_gate
and selected_parameter_gate
and unselected_parameter_gate
and uniform_gate
)
return {
"passed": passed,
"variant": model.forward_variant,
"selected_depth_indices": list(selected),
"output_mixer_selected": False,
"forward_calls": model.forward_calls,
"depth_visit_counts": model.depth_visits,
"output_visit_count": model.output_visits,
"visit_gate": visit_gate,
"source_counts_by_depth_index": source_counts,
"selected_source_count_checks": selected_source_gate,
"selector_gate": selector_gate,
"uniform_weight_max_abs_error": model.uniform_weight_max_abs_error,
"uniform_weight_threshold": threshold,
"uniform_weight_gate": uniform_gate,
"selected_parameters": selected_parameter_checks,
"selected_parameter_reachability_gate": selected_parameter_gate,
"unselected_parameters": unselected_parameter_checks,
"unselected_parameter_reachability_gate": unselected_parameter_gate,
"semantics": (
"selected depth mixers use parameter-free constant-zero logits "
"with the parent softmax+einsum arithmetic kernel"
),
}
def rewrite_result(
output_path: Path,
study_manifest_path: Path,
study_manifest: dict[str, Any],
) -> None:
if LAST_MODEL is None or LAST_OPTIMIZER is None:
raise RuntimeError("runner capture state missing")
result = json.loads(output_path.read_text())
parent_self_hash = result.pop("canonical_sha256_without_self")
if result["protocol_id"] != PARENT_PROTOCOL_ID:
raise RuntimeError("parent runner protocol drift")
result["schema_version"] = 2
result["protocol_id"] = PROTOCOL_ID
result["parent_protocol_id"] = PARENT_PROTOCOL_ID
result["variant"] = ACTIVE_VARIANT
result["parent_runner_canonical_sha256"] = parent_self_hash
result["study_manifest"] = {
"path": str(study_manifest_path),
"file_sha256": parent.file_sha256(study_manifest_path),
"status": study_manifest["status"],
}
result["forward_intervention"] = build_intervention_audit(
LAST_MODEL, LAST_OPTIMIZER, study_manifest
)
result["canonical_sha256_without_self"] = parent.canonical_sha256(result)
temporary = output_path.with_suffix(output_path.suffix + ".round08.tmp")
temporary.write_text(
json.dumps(result, ensure_ascii=False, indent=2, sort_keys=True) + "\n"
)
os.replace(temporary, output_path)
if not result["forward_intervention"]["passed"]:
raise RuntimeError(
f"forward intervention audit failed: "
f"{result['forward_intervention']}"
)
def main() -> None:
global ACTIVE_VARIANT, LAST_OPTIMIZER
variant = extract_wrapper_argument("--variant")
study_manifest_path = Path(
extract_wrapper_argument("--study-manifest")
).resolve()
if variant not in VARIANTS:
raise ValueError(f"unknown variant: {variant}")
run_kind = argument_value("--run-kind", "formal")
if run_kind in ("formal", "replay") and variant not in FORMAL_VARIANTS:
raise ValueError("learned_reference is smoke-only")
if argument_value("--architecture") != "block":
raise ValueError("Round 08 requires --architecture block")
if argument_value("--depth") != "32":
raise ValueError("Round 08 requires --depth 32")
if run_kind == "replay" and variant != "uniform_groups_6_7_forward":
raise ValueError("the frozen replay uses the primary joint variant")
study_manifest = json.loads(study_manifest_path.read_text())
if (
study_manifest["protocol_id"] != PROTOCOL_ID
or study_manifest["status"] != "frozen-before-model-output"
):
raise ValueError("study manifest is not the frozen Round 08 contract")
expected = tuple(
study_manifest["variants"]
.get(variant, {"selected_depth_indices": []})[
"selected_depth_indices"
]
)
if expected != VARIANTS[variant]:
raise ValueError("study manifest selector drift")
for index, source_count in EXPECTED_SOURCE_COUNTS.items():
if (
study_manifest["selected_source_counts"].get(str(index))
!= source_count
):
raise ValueError(
f"study manifest source-count drift at depth index {index}"
)
output_value = argument_value("--output")
if output_value is None:
raise ValueError("--output is required")
output_path = Path(output_value).resolve()
ACTIVE_VARIANT = variant
parent.GradientLanguageModel = ForwardInterventionLanguageModel
original_adamw = torch.optim.AdamW
def capture_adamw(*args: Any, **kwargs: Any) -> torch.optim.Optimizer:
global LAST_OPTIMIZER
LAST_OPTIMIZER = original_adamw(*args, **kwargs)
return LAST_OPTIMIZER
torch.optim.AdamW = capture_adamw # type: ignore[assignment]
try:
parent.main()
finally:
torch.optim.AdamW = original_adamw # type: ignore[assignment]
rewrite_result(output_path, study_manifest_path, study_manifest)
if __name__ == "__main__":
main()
+294
View File
@@ -0,0 +1,294 @@
#!/usr/bin/env python3
"""Run pre-result Round 08 identity gates."""
from __future__ import annotations
import argparse
import hashlib
import importlib.util
import json
import os
import sys
from pathlib import Path
from typing import Any
import torch
def load_runner() -> Any:
path = Path(__file__).resolve().parent / "train.py"
spec = importlib.util.spec_from_file_location("k3_attnres_round08_train", path)
if spec is None or spec.loader is None:
raise RuntimeError(f"cannot import Round 08 runner from {path}")
module = importlib.util.module_from_spec(spec)
sys.modules[spec.name] = module
spec.loader.exec_module(module)
return module
runner = load_runner()
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser()
subparsers = parser.add_subparsers(dest="command", required=True)
step_zero = subparsers.add_parser("step-zero")
step_zero.add_argument("--cache-dir", type=Path, required=True)
step_zero.add_argument("--parent-manifest", type=Path, required=True)
step_zero.add_argument("--study-manifest", type=Path, required=True)
step_zero.add_argument("--output", type=Path, required=True)
step_zero.add_argument("--seed", type=int, default=2026073001)
smoke = subparsers.add_parser("smoke-compare")
smoke.add_argument("--parent", type=Path, required=True)
smoke.add_argument("--wrapper", type=Path, required=True)
smoke.add_argument("--output", type=Path, required=True)
return parser.parse_args()
def canonical_sha256(value: Any) -> str:
return hashlib.sha256(
json.dumps(
value, ensure_ascii=False, sort_keys=True, separators=(",", ":")
).encode()
).hexdigest()
def tensor_sha256(value: torch.Tensor) -> str:
return hashlib.sha256(runner.parent.tensor_bytes(value)).hexdigest()
def read_and_verify(path: Path) -> dict[str, Any]:
value = json.loads(path.read_text())
expected = value["canonical_sha256_without_self"]
payload = {
key: item
for key, item in value.items()
if key != "canonical_sha256_without_self"
}
if canonical_sha256(payload) != expected:
raise RuntimeError(f"canonical self-hash failed: {path}")
return value
def smoke_compare(args: argparse.Namespace) -> None:
parent_result = read_and_verify(args.parent)
wrapper_result = read_and_verify(args.wrapper)
fields = (
"architecture",
"depth",
"seed",
"steps",
"batch_size",
"target_bytes_seen",
"manifest",
"model",
"optimizer",
"hashes",
"evaluations",
"diagnostics",
"training_history",
"gradient_gate",
"environment",
)
checks = {}
for field in fields:
parent_value = exact_structure(parent_result[field])
wrapper_value = exact_structure(wrapper_result[field])
if field == "manifest":
parent_value.pop("path", None)
wrapper_value.pop("path", None)
checks[field] = parent_value == wrapper_value
wrapper_identity = {
"protocol": wrapper_result["protocol_id"] == runner.PROTOCOL_ID,
"parent_protocol": (
wrapper_result["parent_protocol_id"]
== runner.PARENT_PROTOCOL_ID
),
"variant": wrapper_result["variant"] == "learned_reference",
"forward_audit": wrapper_result["forward_intervention"]["passed"],
}
passed = all(checks.values()) and all(wrapper_identity.values())
result = {
"schema_version": 1,
"protocol_id": runner.PROTOCOL_ID,
"gate": "empty-selector-parent-equivalence",
"passed": passed,
"field_checks": checks,
"wrapper_identity": wrapper_identity,
"excluded_fields": [
"protocol wrapper fields",
"timing",
"self hash",
"parent runner self hash",
"study manifest",
],
"parent_file": str(args.parent.resolve()),
"wrapper_file": str(args.wrapper.resolve()),
}
result["canonical_sha256_without_self"] = canonical_sha256(result)
args.output.parent.mkdir(parents=True, exist_ok=True)
args.output.write_text(
json.dumps(result, ensure_ascii=False, indent=2, sort_keys=True) + "\n"
)
if not passed:
raise RuntimeError(f"empty-selector parent equivalence failed: {checks}")
def exact_structure(value: Any) -> Any:
return json.loads(
json.dumps(value, ensure_ascii=False, sort_keys=True)
)
def step_zero(args: argparse.Namespace) -> None:
if not torch.cuda.is_available():
raise RuntimeError("CUDA is required by the frozen step-zero gate")
study_manifest = json.loads(args.study_manifest.read_text())
if study_manifest["protocol_id"] != runner.PROTOCOL_ID:
raise RuntimeError("study manifest mismatch")
parent_manifest = json.loads(args.parent_manifest.read_text())
if parent_manifest["protocol_id"] != runner.PARENT_PROTOCOL_ID:
raise RuntimeError("parent manifest mismatch")
parent = runner.parent
parent.configure_round04_globals(32)
device = torch.device("cuda")
corpus = parent.round04.ByteCorpus(
args.cache_dir, parent_manifest, device
)
inputs, targets = corpus.fixed_batch(
corpus.diagnostic_starts, 0, 16
)
variants = ("learned_reference",) + tuple(
study_manifest["variants"].keys()
)
observations: dict[str, Any] = {}
reference_payload: dict[str, Any] | None = None
for variant in variants:
parent.configure_determinism(args.seed)
runner.ACTIVE_VARIANT = variant
model = runner.ForwardInterventionLanguageModel("block").to(device)
initial_public = parent.named_state_hash(
model, include_mixers=False
)
initial_mixer = parent.named_state_hash(
model, include_mixers=True
)
with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
logits, trace = model(inputs, capture=True)
loss = parent.cross_entropy(logits, targets)
if trace is None:
raise RuntimeError("step-zero trace missing")
evaluation = parent.evaluate(model, corpus, 64, 8)
diagnostic = parent.diagnostic(model, corpus, 16)
payload = {
"initial_public_hash": initial_public,
"initial_mixer_hash": initial_mixer,
"logits_sha256": tensor_sha256(logits),
"loss_nats": loss.detach().cpu().item(),
"loss_tensor_sha256": tensor_sha256(loss),
"evaluation": exact_structure(evaluation),
"diagnostic": exact_structure(diagnostic),
}
if reference_payload is None:
reference_payload = payload
exact_checks = {
key: payload[key] == reference_payload[key]
for key in payload
}
selected = tuple(runner.VARIANTS[variant])
capture_checks = {}
for index in selected:
summary = trace.depth_weights[index]
source_count = summary["sources"]
capture_checks[str(index)] = {
"source_count": source_count,
"expected_source_count": study_manifest[
"selected_source_counts"
][str(index)],
"capture_summary_exact_vs_learned": (
payload["diagnostic"]["depth_weights"][index]
== reference_payload["diagnostic"]["depth_weights"][index]
),
"passed": (
source_count
== study_manifest["selected_source_counts"][str(index)]
and payload["diagnostic"]["depth_weights"][index]
== reference_payload["diagnostic"]["depth_weights"][index]
),
}
runtime_uniform_gate = (
model.uniform_weight_max_abs_error
<= study_manifest["thresholds"][
"uniform_weight_max_abs_error"
]
)
observations[variant] = {
"payload": payload,
"exact_vs_learned_reference": exact_checks,
"selected_capture_checks": capture_checks,
"pre_reduction_uniform_weight_max_abs_error": (
model.uniform_weight_max_abs_error
),
"pre_reduction_uniform_weight_gate": runtime_uniform_gate,
"passed": (
all(exact_checks.values())
and all(
item["passed"] for item in capture_checks.values()
)
and runtime_uniform_gate
),
}
del model, logits, loss, trace
torch.cuda.empty_cache()
passed = all(item["passed"] for item in observations.values())
result = {
"schema_version": 1,
"protocol_id": runner.PROTOCOL_ID,
"gate": "step-zero-cross-variant-byte-exact",
"seed": args.seed,
"passed": passed,
"variants": observations,
"parent_manifest_sha256": runner.parent.file_sha256(
args.parent_manifest
),
"study_manifest_sha256": runner.parent.file_sha256(
args.study_manifest
),
"environment": {
"gpu": torch.cuda.get_device_name(0),
"torch": torch.__version__,
"cuda": torch.version.cuda,
"cublas_workspace_config": os.environ.get(
"CUBLAS_WORKSPACE_CONFIG"
),
},
}
result["canonical_sha256_without_self"] = canonical_sha256(result)
args.output.parent.mkdir(parents=True, exist_ok=True)
args.output.write_text(
json.dumps(result, ensure_ascii=False, indent=2, sort_keys=True) + "\n"
)
if not passed:
failed = [
name
for name, value in observations.items()
if not value["passed"]
]
raise RuntimeError(f"step-zero exactness failed: {failed}")
def main() -> None:
args = parse_args()
if args.command == "step-zero":
step_zero(args)
else:
smoke_compare(args)
if __name__ == "__main__":
main()
+2
View File
@@ -23,6 +23,7 @@
"check:data:k3-attnres-gradient": "node scripts/check-k3-attnres-gradient-data.mjs",
"check:data:k3-attnres-spike": "node scripts/check-k3-attnres-spike-data.mjs",
"check:data:k3-attnres-local-path": "node scripts/check-k3-attnres-local-path-data.mjs",
"check:data:k3-attnres-forward": "node scripts/check-k3-attnres-forward-data.mjs",
"check:site": "node scripts/check-site.mjs",
"check:moe-browser": "node scripts/check-moe-browser.mjs",
"check:reasoning-browser": "node scripts/check-reasoning-browser.mjs",
@@ -46,6 +47,7 @@
"check:k3-attnres-gradient-browser": "node scripts/check-k3-attnres-gradient-browser.mjs",
"check:k3-attnres-spike-browser": "node scripts/check-k3-attnres-spike-browser.mjs",
"check:k3-attnres-local-path-browser": "node scripts/check-k3-attnres-local-path-browser.mjs",
"check:k3-attnres-forward-browser": "node scripts/check-k3-attnres-forward-browser.mjs",
"check:k3-browser": "node scripts/check-k3-browser.mjs"
},
"dependencies": {
@@ -0,0 +1,112 @@
# Round 08 AttnRes 训练期前向干预:结果审计
审计日期:2026-07-30
协议:`llm-atlas-k3-attnres-forward-training-v1`
结果后 Grok 会话:`019fb28b-a9e1-7643-8e43-06f5e16a2077`
## 1. 一句话结论
冻结 analyzer 的唯一主状态为:
```text
forward_training_attenuation_established_within_reduced_protocol
```
联合 `groups 6+7` 的训练期 uniform-forward 消融在三个预注册 seed 上同时通过
`spike contrast` 与 `peak / mean` 的 20% attenuation 门;逐 seed 与三 seed 平均
validation BPC 也全部通过质量护栏。结果后独立只读复算得到:
```text
blocking_errors = 0
status_confirmed = true
replay_confirmed = true
```
这只是在固定 depth-32 缩小 Block AttnRes、固定数据与 8,000-step 预算中的训练期
架构消融;**不是真实 Kimi-K3 / 2.8T checkpoint 结果,也不是 Figure 5(c) 未公开
telemetry 的复现。**
## 2. 主门复算
冻结定义:
```text
S = layers 21–25
R = other 27 layers
C = mean(g[S]) / mean(g[R])
P = max(g) / mean(g)
D = (X_reference - X_variant) / X_reference
```
主变体为 `uniform_groups_6_7_forward`,只读复算如下:
| seed | C reference | C variant | C drop | P reference | P variant | P drop | ΔBPC |
|---|---:|---:|---:|---:|---:|---:|---:|
| 2026073001 | 3.046269 | 0.806011 | 73.54% | 3.219067 | 1.435618 | 55.40% | +0.005696 |
| 2026073002 | 3.332848 | 0.758201 | 77.25% | 3.618924 | 1.375016 | 62.00% | +0.009598 |
| 2026073003 | 1.881492 | 0.712882 | 62.11% | 2.288267 | 1.548141 | 32.34% | +0.004441 |
因此 attenuation 为 `6/6`;最小 contrast drop 为 `62.11%`,最小 peak drop 为
`32.34%`,都高于冻结的 `20%` 门槛。BPC 三格都低于 `+0.05`,平均
`+0.006578` 低于 `+0.03`,质量门为 `4/4`。
这里的 contrast 下降不等价于“尖峰层被关闭”:它可能由 spike-window 分子下降、
27 层 reference 分母上升,或两者共同造成;被联合消融覆盖的 layers 26–28 仍属于
这个分母。
## 3. 完整性与复现
- 13 个新 raw 文件完整:12 formal + 1 primary replay;
- 每格 `65,536,000` target bytes,新处理总量 `851,968,000`;
- 三个历史 paired reference 合计 `196,608,000` bytes,未在 Round 08 重跑;
- 13/13 raw canonical self-hash、aggregate 与 reproduction self-hash 自洽;
- architecture / depth / steps / batch 固定为 Block / 32 / 8,000 / 32;
- initial public/mixer state、schedule、validation、diagnostic、input-gate、model、
optimizer 与 scientific environment 对同 seed historical reference 配对 exact;
- primary seed 1 的 formal / replay scientific payload exact:
`b85563ca5cb53e60b39c3801d372376206105b8a089a8633b3e81973a7f0c051`。
四个正式 selector 的 visit、source count、uniform arithmetic 与 reachability 均通过。
被选 mixer 的 query / key norm 留在 optimizer param groups,但 forward 不再调用它们:
gradient hook 为 0、Adam state 不存在、最终 tensor 与初始值 byte-exact。未选 mixer 的
optimizer-state 检查是比科学协议更强的实现审计,不参与主 status。
## 4. 描述性 non-additivity
冻结的 bookkeeping residual 为:
```text
I67 = ln(Xref / X67) - ln(Xref / X6) - ln(Xref / X7)
```
step 8,000 的三 seed 平均为:
- spike contrast:`-0.367038`
- peak / mean:`-0.170448`
它来自三套独立训练,只能描述 joint run 与两个 single runs 的 log-effect 残差;不能
写成因果 interaction、Shapley contribution 或“group 6/7 互相抑制”的机制结论。
## 5. 两阶段独立审阅
结果前 Grok 实现审阅指出 smoke-only empty selector 的 visit census 可空真。正式
4×3 路径全部是非空 selector,因此不影响 raw formal 数值;矩阵结束后已删除 early
return、加入 `forward_calls > 0`,并重跑 step-zero、parent smoke、wrapper smoke
与 equivalence。修补后的 learned wrapper 实际执行 39 次 forward,父/包装器 15 组
科学字段仍全部 exact。
结果后 Grok 在只读 sandbox 中从 13 个 raw 与三个 historical references 独立复算
identity、自哈希、selector、pairing、两项主指标、BPC、replay 与 `I67`。它报告
`0 mismatch`、`blocking_errors=0`,确认 analyzer status 与 claim boundary。
## 6. 最终 claim boundary
可以说:在这一固定缩小协议内,联合 group 6+7 的 train-time uniform-forward
architecture ablation 相对历史同 seed reference 达到预注册 attenuation + BPC 门控。
不能说:
- 已定位真实 Kimi-K3 的训练尖峰;
- 已复现 K3 报告 Figure 5(c);
- 已把 forward、natural backward 与 optimizer update 分离成纯因果效应;
- 已证明下游能力等价、总体统计显著性、可加性或因果 interaction。
@@ -0,0 +1,69 @@
# Round 08 Grok Headless 对抗审阅与处置
审阅日期:2026-07-30
审阅会话:`019fb1ce-7ffa-70e3-b860-c4a31a4c6621`
身份:**外部模型的只读方法学审稿,不是论文证据源**
## 1. 调用边界
Grok CLI 使用 single/headless + plan permission 读取:
- `research/K3_ATTNRES_FORWARD_TRAINING_SCOPING.md`
- `research/K3_ATTNRES_FORWARD_TRAINING_PROTOCOL.md`
- `experiments/k3/attnres_forward/manifest.json`
- Round 04 / 05 父 runners
关闭 web search、禁止 subagents;它没有修改文件,也没有运行训练。
## 2. Blocking findings 与处置
| finding | 风险 | 处置 |
|---|---|---|
| `mean` 与父 `softmax+einsum` 的 FP32 归约顺序不保证 byte-exact | step-0 会假失败 | **采纳**:选中路径改成参数无关 constant-zero logits,并复用同一 `softmax+einsum` kernel |
| pairing 要求了父 JSON 不存在的 initial full hash,并把 GPU 名写成 exact | 科学身份正确却因 metadata 失败 | **采纳**:只 hard-gate 初始化、数据、拓扑、optimizer 与确定性/autocast 合同 |
| replay 只排除三个字段,path 会制造假差异 | exact replay 假失败 | **采纳**:冻结 scientific canonical payload 与 path normalization |
| wrapper 若把新 protocol ID 写入 `window_start` salt,会静默换数据 | 历史逐 seed 配对失效 | **采纳**:父 salt 只由父 manifest/runner 管,新 ID 只进 wrapper output |
| 主判定没有在协议正文再写 final-only | analyzer 可能误读六 checkpoints | **采纳**:主 attenuation 与 BPC 只读 step 8,000 |
| spike layer numbering 未 machine-readable 固定 | 可能整体平移一层 | **采纳**:manifest 新增 `spike_layers_1based`,协议钉死 Python index = layer−1 |
## 3. Non-blocking findings 与处置
全部采纳:
- selected 参数留在 AdamW param groups,但因 graph 不可达而没有 state entry,不能写成
“训练了但没动”;
- group 6 / 7 indices 与 source counts 经独立复算正确;
- reference 统一称 historical paired reference;
- layers 26–28 位于 contrast 分母 `R`,必须拆报 `mean(S)` / `mean(R)`;
- `I67` 明确是三次独立训练之间的 log residual;
- missing / non-finite / structural error 统一 `contract_failed`,不能冒充科学失败;
- wrapper 禁止改变 FP32 residual、AdamW grouping、parent salt 与 empty-selector path。
## 4. 算术复核
```text
8,000 × 32 × 256 = 65,536,000 bytes / cell
12 formal cells = 786,432,000 bytes
+ primary replay = 851,968,000 newly processed bytes
historical refs = 196,608,000 bytes(不重跑)
```
selector:
```text
group 6 = 40..47; source N = 6,7,7,7,7,7,7,7
group 7 = 48..55; source N = 7,8,8,8,8,8,8,8
group 7 MLP = 49,51,53,55
groups 6+7 = 40..55
```
主 attenuation 为 `3 seeds × {contrast, peak} = 6` cells;quality 为三个
per-seed BPC gates 加一个 mean gate,共 4 项。两者合取,且只读 final step。
## 5. 复核结论
Grok 判断研究身份、矩阵、source counts 与主公式骨架可以保留;主要风险来自浮点
arithmetic identity、wrapper salt、over-exact metadata 与层号歧义。以上项目已在任何
正式输出出现前全部修订,manifest 状态随后改为 `frozen-before-model-output`。
审稿意见不会进入实验结果、论文事实或官网证据等级;它只用于结果前强化协议。
@@ -0,0 +1,89 @@
# Round 08 runner / analyzer 实现审阅与处置
审阅日期:2026-07-30
Grok 会话:`019fb1f0-bb24-7202-8123-295edda7518f`
身份:**正式文件完成前的外部模型只读实现审计,不是结果或论文证据**
## 1. 审阅边界
Grok Headless 只读检查:
- 冻结协议与 manifest;
- `train.py` / `verify.py` / `analyze.py` / `run_matrix.py`;
- 父 runner 的 salt、schema 与 DepthMixer 算术路径。
明确禁止读取 Round 08 raw formal results、运行训练、修改文件、web search 与 subagents。
## 2. 对正式 4×3 路径的确认
审阅确认:
- 新 protocol ID 没有进入父 `window_start` salt;
- selected 路径用 constant-zero FP32 logits 和父 `softmax+einsum` kernel;
- 非空 selector 的 query / key norm hook、AdamW state 与 final=initial gate 自洽;
- historical pairing 比较的字段在父 JSON 中真实存在;
- replay payload 正确排除 run kind、timing、path、parent self hash 与 GPU/version metadata,
同时保留模型、optimizer、diagnostics、history、selector 与 reachability;
- analyzer 强制 4×3 identity、step 8,000 主判定、1-based layers 21–25、finite/positive、
`D=(ref-variant)/ref`、BPC `variant-ref` 与 descriptive `I67`;
- 12 formal + 1 replay、65,536,000 bytes/cell 与最大两进程算术正确。
没有发现会改变正在运行的四个非空 formal variants 数值语义的 blocking error。
## 3. Blocking finding:smoke-only empty selector 的空审计
`learned_reference` 直接调用父 `forward`,没有递增 wrapper 的 visit counters。因此:
```text
forward_calls = 0
all depth/output visits = 0
visit_gate = all(0 == 0) = true
selected reachability = all([]) = true
```
这不会影响四个正式变体,它们全部是非空 selector;而 empty selector 另有 20-step
parent-equivalence gate,模型/optimizer/evaluations/diagnostics/history/hash 已 exact。
但 `forward_intervention.passed` 本身不能在修复前被当作 learned census 证据。
处置:
- formal matrix 完成后,删除 `selected_indices` 为空时的父路径 early return;
- 让 empty selector 也走同一 copied forward,其中 `mix` 对 64 个节点逐个调用父
`DepthMixer.forward`;
- 加入 `forward_calls > 0`;
- 重跑 step-0、parent smoke、wrapper smoke 与 parent equivalence;
- 只有 copied path 仍逐字段 exact 才保留。
这项修复只强化 smoke audit,不更改任何正式 variant 的 selector 或 forward。
## 4. Non-blocking findings 与处置
| finding | 处置 |
|---|---|
| `smoke_compare` 整体比较 `manifest.path` | 规范化 path 后比较 scientific manifest fields |
| unselected gate 额外要求 optimizer state | 保留为强实现 gate,但在文档中标成 protocol 之外的额外审计,不用于科学 status |
| GPU/version 差异只有 `metadata_equal`,没有 warnings 数组 | aggregate 增加显式 metadata warnings |
| `run_matrix` 失败后只 terminate、不 wait/kill,invalid file 会阻塞重跑 | 加入 terminate→wait→kill 清理,并在 cell failure 时标明 exact invalid target |
| step-0 CE 只比较 Python float | 追加 scalar tensor SHA-256 |
| `EXPECTED_SOURCE_COUNTS` 未使用 | 用于 runner↔manifest 交叉校验 |
上述修订不读取结果、不改变冻结阈值或主公式。
## 5. 矩阵结束后的处置结果
13 个单元全部退出后才应用上述修订;四个正式非空 selector 的 raw 文件未被重写。
修补后的前置闸门结果:
- step-zero 五个 variant 的 logits、loss tensor、evaluation 与 diagnostic exact;
- smoke-only learned wrapper 实际执行 `39` 次 forward,64 个 depth mixer 与
output mixer 的 visit census 全部非零且 exact;
- parent 与 wrapper 的 architecture、seed、schedule、model、optimizer、hash、
evaluation、diagnostic、history、gradient gate 与 scientific environment 共
15 组字段全部 exact;
- primary smoke 的 selected 参数仍为 0 hook、无 optimizer state、final=initial;
- unselected optimizer-state 条件继续作为额外实现闸门,不进入科学 status。
结果后 Grok 会话 `019fb28b-a9e1-7643-8e43-06f5e16a2077` 在只读 sandbox 中独立
复算 13 个 raw、三个 historical references 与 aggregate,报告
`blocking_errors=0`、`status_confirmed=true`、`replay_confirmed=true`。详细数字与
claim boundary 见 `research/K3_ATTNRES_FORWARD_TRAINING_AUDIT.md`。
@@ -0,0 +1,389 @@
# K3 Attention Residuals 训练期前向干预协议
协议 ID:`llm-atlas-k3-attnres-forward-training-v1`
冻结日期:2026-07-30
协议状态:**结果前预注册 frozen;任何语义变更必须更换 protocol ID**
父协议:`llm-atlas-k3-attnres-gradient-scale-v1`
## 0. 研究身份
这是 Round 07 定向线索之后的训练期架构消融。选中 depth mixer 在每一次 train / eval /
diagnostic forward 都用 source states 的算术平均,完全绕过该 mixer 的
`query + key_norm + softmax` 路径。
允许回答:
1. 固定 groups 6+7 的 uniform forward 训练变体,能否在不触发预注册 BPC 失败护栏时,
material 地降低最终固定 activation-gradient spike?
2. group 6、group 7 与 joint 的训练轨迹呈现什么非加性关系?
3. Round 07 指向的 group 7 MLP-only 路径能否独立产生 material response?
不允许回答:
- 真实 Kimi K3 2.8T checkpoint 的梯度或训练动力学;
- 论文 Figure 5(c) 未公开 telemetry 的复现;
- “forward effect” 与 natural backward/update effect 的分离;
- selected query/key 参数如果继续训练会怎样;
- 三 seed 外的总体显著性、置信区间或 p-value;
- 下游能力保持、通用质量等价或最优 AttnRes 设计;
- 单组 effects 的可加性、Shapley value、方差贡献或因果交互;
- 与 Round 07 value-coefficient intervention 同构的“纯 forward”因果复制;
- K3 `A_log` 的官方修复裁决。
## 1. 冻结训练与数据合同
| 字段 | 固定值 |
|---|---|
| architecture | Block AttnRes |
| Transformer depth | 32 |
| aggregation groups | 8 |
| blocks / group | 4 |
| depth / output mixers | 64 / 1 |
| width / heads / FFN | 192 / 6 / 768 |
| context / vocabulary | 256 / byte-256 |
| seeds | 2026073001 / 2026073002 / 2026073003 |
| steps / batch | 8,000 / 32 |
| target bytes / new formal cell | 65,536,000 |
| optimizer | AdamW |
| peak / min LR | 3e-4 / 3e-5 |
| warmup | 400 |
| weight decay | 0.1 for ndim ≥ 2 |
| betas / epsilon | 0.9, 0.95 / 1e-8 |
| clip | global norm 1.0 |
| forward | CUDA BF16 autocast |
| residual accumulation | explicit FP32 |
| validation | fixed 64 × 256-byte windows |
| diagnostic | fixed 16 × 256-byte windows |
| checkpoints | 0 / 100 / 500 / 2,000 / 4,000 / 8,000 |
| concurrency | at most two independent processes |
训练输入 schedule **逐 step 复用父协议**。`window_start` 使用父
`llm-atlas-k3-attnres-gradient-scale-v1` 的 salt;新 protocol ID 只写入 wrapper output
和 study manifest,绝不能进入 `round04.PROTOCOL_ID` 或训练窗口散列。父 manifest
负责 bytes / windows / schedule,新 manifest 只负责 variants / selector / thresholds。
每个 variant / seed 的初始化、optimizer input、validation 与 diagnostic tensors 必须
exact 相同。
runner 必须继承 Round 05 `GradientLanguageModel` 的 explicit FP32 Block residual
accumulation;不得退回 Round 04 的旧累加路径。
新正式处理量:
```text
4 variants × 3 seeds × 65,536,000 = 786,432,000 target bytes
1 primary replay 65,536,000 target bytes
total newly processed 851,968,000 target bytes
historical learned reference 196,608,000 target bytes(不重跑)
```
每格必须使用全新 Python process。最多并行两个;不能共享 model、optimizer、RNG、
CUDA graph 或 output file。
## 2. 冻结正式矩阵
正式 variants:
| variant | exact selected depth indices | layers / branch |
|---|---|---|
| `uniform_group_6_forward` | 40–47 | 21–24 / both |
| `uniform_group_7_forward` | 48–55 | 25–28 / both |
| `uniform_groups_6_7_forward` | 40–55 | 21–28 / both |
| `uniform_group_7_mlp_forward` | 49, 51, 53, 55 | 25–28 / MLP |
`learned_reference` 只允许 smoke,不进入新正式矩阵。output mixer index 64 永远 learned。
正式运行 12 格;另从初始化 replay:
```text
replay / uniform_groups_6_7_forward / seed 2026073001
```
## 3. 唯一 selector 与 forward 语义
runner 必须只有一个 machine-readable selector:
```text
selected(variant, depth_mixer_index) -> bool
```
不得把四个 variant 分叉成四份 model forward。
对未选中 depth mixer 和 output mixer,逐调用父 `DepthMixer.forward`。对选中 mixer,
为了让初始化负控制复用相同浮点归约顺序,使用参数无关的零 logits,但仍走父
`softmax + einsum` 数值 kernel:
```text
values = stack(sources, dim=0)
logits = zeros([N, batch, tokens], dtype=FP32)
weights = softmax(logits, dim=0)
output = einsum("nbt,nbtd->btd", weights, values.float()).to(values.dtype)
```
capture summary 必须仍使用父 schema:
```text
mean_weights = [1 / N] × N
entropy_mean = ln(N)
sources = N
```
选中路径不得调用 `query`、`key_norm` 或 source-dependent logits,也不得用
stop-gradient trick 让这些参数看似参与。这里保留的 softmax 只把常数零 logits 变成
`1/N`,目的是与父路径保持同一 arithmetic kernel;它没有可训练参数。自然结果是选中
mixer 的 `query` 和 `key_norm.weight`:
- gradient hook call count 必须为 0;
- optimizer state entry 必须不存在;
- final tensor 必须与 initial tensor byte-exact。
所有未选中 depth mixer 和 output mixer 的两个参数都必须有正的 gradient hook call
count;这只证明图可达,不要求它们的梯度非零或最终 tensor 一定变化。hook census
从 model 构造后开始,覆盖所有 training backward 与 diagnostic backward;eval
forward 不计 hook。selected 参数允许保留在原 AdamW param groups,但 state entry 必须
不存在,不能表述为“训练了但没有移动”。
## 4. 结果前实现闸门
### 4.1 empty-selector 父等价
`learned_reference` smoke 必须直接走父 forward,不得走常数 uniform 分支;它与父
Round 05 runner 在相同 seed / 20 steps 下必须:
- initial/final model hashes exact;
- final optimizer hash exact;
- evaluations、diagnostics 与 training history exact;
- gradient gate exact;
- input tensor hashes exact。
允许不同字段只限 protocol wrapper identity、study-manifest wrapper、timing、规范化
后的 manifest path、output path 与 self-hash。AdamW 参数分组必须保持父语义;
empty-selector 不能删除任何 mixer 参数。
### 4.2 step-0 identity negative control
父模型所有 depth mixer query 初始化为 0,因此 learned softmax 在 step 0 是 exact
uniform。选中分支用同 dtype 的 constant-zero logits 和同一个
`softmax + einsum` kernel。四个 variant 与 learned reference 在固定 input 上必须:
- logits byte-exact;
- CE byte-exact;
- activation-gradient spectrum byte-exact;
- validation metrics byte-exact;
- selected 的**归约前 weight tensor** 等于 FP32 `1/N`,max absolute error
`≤ 1e-12`;capture 的 `mean_weights` 因 FP32 大规模 mean 可有约 `1e-8` 的归约舍入,
但必须与父 learned capture summary byte-exact;
- 若任何跨 variant byte-exact 比较失败,hard-fail;不得在结果后改成容差 gate。
这个负控制只约束初始化;训练开始后 forward 必须允许分化。
### 4.3 selector census
每个 forward 的 64 个 depth index 必须各访问一次,output 访问一次且保持 learned。
每个 variant 的 selected set 必须与第 2 节 exact。group 内 source counts 必须满足:
```text
group 6: index 40 has N=6; indices 41..47 have N=7
group 7: index 48 has N=7; indices 49..55 have N=8
```
推导前提是 `completed` 含 embedding,且每 8 个 depth mixer 才把 `partial` 聚合为一个
completed group。
正式 output 保存 exact selected indices、实际 visit census、source counts 与
uniform-weight max error。任何漏访、重访、越界或 output 被选中都失败。
### 4.4 数据、有限性与梯度尺度
- 父 manifest、train/validation/diagnostic bytes 与 schedule hashes exact;
- step 0 / 1 / 7,999 optimizer input gate hashes exact;
- loss、logits、所有主 activation gradients 全部 finite;
- global clip 后每一步都执行 optimizer update,不允许 skip;
- smoke 的 diagnostic loss `×2` 时,每层 activation-gradient RMS 比值在
`2 ± 1e-5`,normalized spectrum max delta `≤ 1e-6`。
## 5. 历史 reference 配对合同
reference 固定为:
```text
experiments/k3/attnres_gradient/results/raw/
formal-depth-32-block-seed-{seed}.json
```
analyzer 的 pairing hard gates:
- `protocol_id = llm-atlas-k3-attnres-gradient-scale-v1`;
- formal / block / depth 32 / 8,000 steps / batch 32;
- seed exact;
- initial public 与 mixer hashes 跟对应新 variant exact;
- formal schedule、validation tensor、diagnostic tensor 和三个 input gate hashes exact;
- model topology、optimizer hyperparameters、CUBLAS workspace、deterministic flags 与
autocast 语义 exact。
以下字段明确**不参与 pairing equality**:
- 所有 final hashes、evaluations、diagnostics、training history 与 gradient gate;
- timing、run kind、self-hash、output path 与 manifest path 字符串;
- GPU 名称、driver / CUDA / torch version 的 minor 差异。
环境完整记录;若数值栈变化,aggregate 给出 metadata warning,但只要上述确定性与
autocast 合同相同就不将其误判为 pairing failure。父 JSON 没有 initial full-state 字段,
不得假定它存在;public ∪ mixer 的完整性只用结构/元素 census 自洽。
reference 是历史配对基线,不得写成同期随机对照。若任何合同不等,整轮 aggregate
失败,而不是降级为“近似比较”。
## 6. 固定主对象与公式
每个 diagnostic checkpoint 从:
```text
activation_grad_rms_by_block = [g1, ..., g32]
layer_id ∈ {1,...,32}
g[layer_id] = activation_grad_rms_by_block[layer_id - 1]
S = {21,22,23,24,25} # 1-based
R = {1,...,32} \ S
```
计算:
```text
C = mean(g[S]) / mean(g[R]) # spike contrast
P = max(g) / mean(g) # peak normalized
```
所有 `g`、`C`、`P` 必须 finite 且严格大于 `1e-30`。groups 6+7 还改写落在 `R`
中的 layers 26–28,所以 analyzer 同时报告 `mean(g[S])` 与 `mean(g[R])`,但不把它们
加入主 status。
对同 seed reference `X_ref` 与 variant `X_v`:
```text
D_X(v) = (X_ref - X_v) / X_ref
```
`D>0` 表示 attenuation,`D<0` 表示 amplification。不得取绝对值,不得更换分母。
## 7. 预注册判定
### 7.1 主判定
`uniform_groups_6_7_forward` 的主 attenuation gate **只读取 step=8,000**:
```text
D_C >= 0.20 AND D_P >= 0.20
for all 3 seeds
```
质量 gate 也只读取 `evaluations[step=8000].bits_per_byte`:
```text
delta_bpc(seed) = final_bpc_variant - final_bpc_reference
delta_bpc(seed) <= 0.05 for all 3 seeds
mean(delta_bpc) <= 0.03
```
只有 attenuation 6/6 与 quality 4/4 同时通过,正式 status 才是:
```text
forward_training_attenuation_established_within_reduced_protocol
```
否则按失败位置使用:
```text
attenuation_not_established
quality_guard_failed
attenuation_and_quality_failed
```
不能用次级变体补救主判定。任何 step 8,000 缺失/重复、数组长度错误、hash 不配对、
selector / reachability 失败、`g/C/P/BPC/D/log` 缺失或非 finite 都是
`contract_failed` 并让 analyzer non-zero exit;不能把结构失败包装成上面的科学状态。
### 7.2 次级 material response
group 6、group 7、group 7 MLP-only 各自使用同一 `20% / 3-seed / 2-metric` attenuation
threshold 和同一 quality guard,分别报告:
```text
secondary_material_response / secondary_response_not_established
```
它们不改变主 status,也不升级成 localization。
### 7.3 BPC 护栏的解释
`+0.05 per seed / +0.03 mean` 是预注册的 catastrophic-degradation screen:
- 失败说明不能把 spike 下降当成健康训练的证据;
- 通过不说明能力、校准或下游任务等价;
- BPC 改善也不说明总体架构更优。
## 8. 非加性交互与轨迹
对 `X ∈ {C,P}`、每个 seed、每个 checkpoint:
```text
E6 = ln(X_ref / X_group6)
E7 = ln(X_ref / X_group7)
E67 = ln(X_ref / X_groups6+7)
I67 = E67 - E6 - E7
```
保存 `E6/E7/E67/I67` 原值、对应 `D_C/D_P` 和三 seed mean/range。没有通过阈值、
p-value 或 CI。它是三套独立训练在相同 checkpoint 的跨-run log residual;
`I67` 不能写成可加贡献、独立作用、Shapley value 或因果 interaction estimate。
同时全量保存:
- 六 checkpoints 的 32-layer raw / normalized spectra;
- peak layer、top-five layers;
- validation BPC 与 train-loss trajectory;
- selected/unselected mixer weight summaries;
- selected-parameter reachability audit;
- per-cell timing 与显存(不进入数值结论)。
## 9. replay 与 analyzer 合同
primary seed-1 replay 使用 analyzer 定义的 scientific canonical payload。先删除:
```text
run_kind
timing
canonical_sha256_without_self
manifest.path
study_manifest.path
```
再比较以下固定字段 exact:protocol / variant / architecture / depth / seed / steps / batch /
target bytes、manifest scientific hashes、model、optimizer、initial/final hashes、
evaluations、diagnostics、training history、selector 与 reachability audits,以及
environment 中 deterministic / autocast scientific subset。GPU/版本 metadata 保留在
两份文件中单独展示,不进入 canonical equality。
所有主指标、阈值、status、interaction map 与 compact website artifact 只能由单一
`experiments/k3/attnres_forward/analyze.py` 生成。网站不能在 TypeScript 中重新计算
另一套结论。
analyzer 在任何结构、hash、selector、reachability、finite、reference pairing、
replay 或 threshold contract 失败时必须 non-zero exit,不得输出部分通过结论。
## 10. 报告语言红线
允许:
- “在这个固定缩小模型与训练协议内,局部 uniform-forward 变体……”
- “selected mixer 参数在此架构消融中结构性不可达……”
- “joint log effect 呈现正/负 interaction residual……”
禁止:
- “证明 K3 的训练尖峰来自 group 6/7”
- “只改变 forward,所以这是纯 forward 因果效应”
- “BPC gate 通过,所以能力不受影响”
- “interaction residual 是两个 group 的真实贡献”
- “contrast 下降证明尖峰层本身下降”(未同时检查 `S` / `R` 分拆)
- “step-0 exact 说明训练期始终与 learned forward 恒等”
- “复现了 K3 Figure 5(c)”
- “已经验证官方 2.8T checkpoint”
@@ -0,0 +1,195 @@
# K3 Attention Residuals 训练期前向干预:Round 08 前置定位
研究日期:2026-07-30
阶段身份:**Round 07 后的定向 scoping,不是 Round 08 结果**
上游协议:`llm-atlas-k3-attnres-local-path-v1`
## 1. 为什么还需要一次训练期实验
Round 06 / 07 都保持 learned forward 完全不变,只在固定 diagnostic 的 backward 中
替换 source value coefficients。它们回答的是:
- 全部 65 个 mixer 的 uniform value backward 能不能压低固定尖峰;
- group 6 / 7 的 16 个 depth mixer 在 learned 背景上是否足以复现全局下降;
- 从 all-uniform 背景恢复这些 mixer 是否能反向恢复尖峰。
Round 07 的正式结论是:
```text
groups 6+7 sufficiency:6 / 6 seed×metric cells 通过
groups 6+7 restoration:3 / 6 cells 通过
formal status:one_sided_evidence_localization_not_established
```
这已经足以排除“局部 mask 完全没有反应”,但还不能回答:
> 如果训练的每一次 forward 都真的把这段 depth routing 改成算术平均,模型会怎样适应?
Round 08 把 intervention 放进 optimizer path。它不再追求 forward-identical,而是让
选中 mixer 的输出在训练、验证与诊断中始终为所有 source states 的等权平均。
## 2. 这不是“只改变 forward”
选中 mixer 的 learned 路径原本是:
```text
keys = RMSNorm(sources)
logits = query · keys
weights = softmax(logits over source-depth)
output = Σ weights_i × source_i
```
Round 08 的选中路径是:
```text
output = (1 / N) × Σ source_i
```
因此 intervention 同时改变:
1. forward 的 branch input;
2. 由新 forward 自然产生的 source gradients;
3. 下游 activation、loss 与所有后续 optimizer updates;
4. 选中 mixer 的参数可达性:`query` 与 `key_norm.weight` 不参与图,不得到梯度。
这是一项**训练期架构消融**,不是“只改变 forward、不改变 backward”的可分离因果实验。
结果不能被翻译成 query/key 路径的纯因果效应。
## 3. 为什么只选四个新变体
固定 depth-32 / Block AttnRes 拓扑:
| scope | layers | 0-based depth mixer indices | count |
|---|---:|---:|---:|
| group 6 | 21–24 | 40–47 | 8 |
| group 7 | 25–28 | 48–55 | 8 |
| groups 6+7 | 21–28 | 40–55 | 16 |
| group 7 MLP | 25–28 | 49, 51, 53, 55 | 4 |
四个新训练变体固定为:
1. `uniform_group_6_forward`
2. `uniform_group_7_forward`
3. `uniform_groups_6_7_forward`(主变体)
4. `uniform_group_7_mlp_forward`
选择依据不是 Round 08 结果:
- joint 6+7 是 Round 07 的固定主 scope;
- 单 group 6 / 7 用来构成交互图;
- group 7 MLP-only 是 Round 07 的 branch-level 次级线索;
- output mixer 保持 learned,避免把 local depth intervention 扩成全局 readout 改写。
不加入 attention-only、output-only、all-depth 或 all-65,是为了把本轮正式算力集中在
“局部训练适应”而不是重新展开 14-mode backward census。
## 4. 冻结参考,而不是重新挑基线
paired learned reference 复用 Round 05 已公开的三个:
```text
formal / depth-32 / block / seeds 2026073001..3
```
它们与新变体使用相同:
- 初始化过程;
- 8,000-step deterministic byte schedule;
- 64 validation windows;
- 16 diagnostic windows;
- optimizer、LR schedule、BF16 forward 和 FP32 residual accumulation;
- 六个 diagnostic checkpoints。
runner 的 `learned_reference` smoke mode 必须逐字段复现父 runner 数值。正式 analyzer
还必须验证新旧 cell 的 initial public/mixer hashes、输入 gate tensor hashes、
validation tensor hash、diagnostic tensor hash 与 deterministic / autocast 数值合同。
GPU 名称、驱动 minor 与 timing 只记录,不作为配对 hard gate。
复用参考的意义是避免为已冻结、可复现的三格再消耗 196,608,000 target bytes。报告必须
把它写成**历史配对参考**,不能写成同期随机对照。
## 5. 主问题与质量护栏
主对象沿用 Round 05–07,并且主判定只读取最终 step 8,000:
```text
step = 8,000
position = post-MLP Transformer-block output
reduction = activation-gradient element RMS
S = layers 21–25
R = other 27 layers
C = mean(g[S]) / mean(g[R])
P = max(g) / mean(g)
```
对变体 `v` 和同 seed learned reference:
```text
D_X(v) = (X_ref - X_v) / X_ref
```
主变体 `uniform_groups_6_7_forward` 只有在:
- `C` 与 `P` 的 `D_X ≥ 0.20`;
- 三个 seed 的六个 cell 全部通过;
- 每个 seed 的 final validation BPC 相对 reference 不劣化超过 `+0.05`;
- 三 seed mean BPC delta 不超过 `+0.03`;
时才得到 “forward-training attenuation established within reduced protocol”。
质量阈值在结果前固定。它不是“能力保持”的通用标准,只用于阻止把明显训练失败造成的
梯度塌缩计成正结果。即使通过,也不能推出下游能力等价。
## 6. 非加性局部交互图
对每个 seed、checkpoint 和指标:
```text
E6 = ln(X_ref / X_group6)
E7 = ln(X_ref / X_group7)
E67 = ln(X_ref / X_groups6+7)
I67 = E67 - E6 - E7
```
- `I67 > 0`:joint log attenuation 大于两个 single effects 的和;
- `I67 < 0`:joint log attenuation 小于两个 single effects 的和;
- `I67 = 0`:只是在这个定义下恰好 log-additive。
`I67` 没有预注册显著性阈值,不是 Shapley value、方差分解、独立性检验或因果交互估计。
三个 effect 来自三套独立训练,它只是跨 run 的 log-attenuation residual。它的用途是把
训练轨迹中的补偿/放大关系画清楚,而不是制造一个新的“通过/失败”结论。
groups 6+7 覆盖 layers 21–28,而固定尖峰窗只到 layer 25;layers 26–28 落在 `R`。
所以 `C` 的变化可能同时来自 `S` 下降与 `R` 上升。正式结果必须把两者拆开报告,不能把
contrast 下降单独翻译成“尖峰层被关闭”。
## 7. 真实 K3 checkpoint 的同期边界
截至本轮预检,官方 Kimi-K3 Hugging Face main 仍停在 revision
`9f62e4e9fffbd0a83ddd60e1c209d828994b3569`,remote code 仍按 96 heads 初始化
`A_log`,公开 checkpoint header 仍为 `[128]`。社区 PR #144 / #150 仍是两个未合并、
语义不同的候选修复;没有官方裁决。
所以本轮不下载约 1.56 TB 权重,不声称对真实 K3 forward 做了验证。缩小实验只继承
Block AttnRes 的拓扑动机,不是 K3 checkpoint 的数值替身。
一手状态页:
- [Kimi-K3 official main](https://huggingface.co/moonshotai/Kimi-K3/tree/main)
- [main `modeling_kimi_linear.py`](https://huggingface.co/moonshotai/Kimi-K3/blob/main/modeling_kimi_linear.py)
- [community PR #144](https://huggingface.co/moonshotai/Kimi-K3/discussions/144)
- [community PR #150](https://huggingface.co/moonshotai/Kimi-K3/discussions/150)
## 8. 本轮可证伪交付
Round 08 将在查看正式结果前完成:
1. 冻结协议与 machine-readable manifest;
2. 实现一个 selector,而不是四份分叉 forward;
3. 通过 empty-selector 父等价、step-0 uniform identity、selector census、参数不可达性、
loss-scale 与输入 hash 闸门;
4. 运行 4 variants × 3 seeds × 8,000 steps;
5. 从初始化 replay 主变体 seed 2026073001;
6. 由单一 analyzer 生成主判定、质量闸门、轨迹与非加性交互;
7. 独立审阅机器可读结果;
8. 以五视图交互实验接入网站、开源并发布。
+1 -1
View File
@@ -228,7 +228,7 @@ if (numeric(reliability.initial.passAt) <= numeric(reliability.k2.passAt) || num
if (reliability.nonIdempotent.sideRisk === "LOW") failures.push("非幂等写操作风险没有提升");
if (numeric(rl.wait.utilization) >= numeric(rl.full.utilization) || numeric(rl.wait.lostWork) <= numeric(rl.full.lostWork)) failures.push("wait-all 长尾/重算方向异常");
if (!rl.wait.takeaway.includes("wait-all") || rl.keyboardSelected !== "rl" || rl.keyboardVisible !== "rl") failures.push("长程 RL 解释或键盘导航异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("局部 uniform routing 进入完整训练后") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
if (home.paperCount !== "486" || papers.total !== 486 || !papers.hasAgentFilter || papers.agentVisible < 52) failures.push("论文库 Agent 标签或论文总数异常");
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
+1 -1
View File
@@ -226,7 +226,7 @@ if (!update.steps[0].includes("Fixed preference")) failures.push("DPO 更新流
if (!recipe.family.includes("Multi-effort") || !recipe.regime.includes("9 RL experts") || !recipe.constraints.includes("verbosity")) failures.push("K3 配方合同异常");
if (!recipe.path.some((step) => step.includes("3 domains × 3 efforts")) || !recipe.path.some((step) => step.includes("MOPD"))) failures.push("K3 配方路径异常");
if (recipe.keyboardSelected !== "recipe" || recipe.keyboardVisible !== "recipe") failures.push("实验 tab 键盘导航异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("局部 uniform routing 进入完整训练后") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
if (home.paperCount !== "486" || papers.total !== 486 || !papers.hasAlignmentFilter || papers.alignmentVisible < 35) failures.push("论文库后训练标签或论文总数异常");
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
+1 -1
View File
@@ -234,7 +234,7 @@ if (layout.navLinks !== 20 || mobile.mobileLinks !== 20 || home.navLinks !== 20)
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") {
if (home.releaseCards !== 17 || !home.firstRelease.includes("局部 uniform routing 进入完整训练后") || home.firstHref !== "/k3/") {
failures.push("首页 Transformer 新章入口异常");
}
if (home.paperCount !== "486") failures.push(`首页论文总数异常:${home.paperCount}`);
+1 -1
View File
@@ -1318,7 +1318,7 @@ if (completionDepth.tasks.panel !== "tasks" || completionDepth.tasks.mathCards !
if (completionDepth.hidden.panel !== "hidden" || completionDepth.hidden.stages !== 29 || completionDepth.hidden.selected !== "layer_07" || completionDepth.hidden.points !== 29 || completionDepth.hidden.exact === "1,537 / 1,537" || numeric(completionDepth.hidden.relative) <= 0) failures.push("29 阶段隐藏状态曲线或交互异常");
if (completionDepth.router.panel !== "router" || completionDepth.router.layers !== 26 || completionDepth.router.selected !== "layer 24" || completionDepth.router.points !== 26 || !completionDepth.router.ordered.includes("%") || !completionDepth.router.setExact.includes("%") || numeric(completionDepth.router.tv) <= 0 || completionDepth.router.reproCards !== 4 || completionDepth.router.links !== 4) failures.push("26 层 MoE 路由曲线或复跑证据异常");
if (completionDepth.keyboardSelected !== "tasks" || completionDepth.keyboardVisible !== "tasks") failures.push("完成度与全深度实验键盘 tab 导航异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/" || home.paperCount !== "486") failures.push("首页 K3 首发入口或论文数异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("局部 uniform routing 进入完整训练后") || home.firstHref !== "/k3/" || home.paperCount !== "486") failures.push("首页 K3 首发入口或论文数异常");
if (papers.total !== 486 || !papers.hasFilter || papers.visible < 20 || !papers.hasCoder || !papers.hasEngram) failures.push("论文库 DeepSeek 聚光异常");
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4 || mobile.artifactTabs !== 13 || mobile.behaviorTabs !== 4 || mobile.behaviorSources !== 16 || mobile.behaviorEdges !== 10 || mobile.behaviorDeviceCells !== 29 || mobile.completionDepthTabs !== 4 || mobile.completionDepthHiddenStages !== 29 || mobile.completionDepthRouterLayers !== 26 || mobile.artifactHeatCells !== 64 || mobile.corpusCohorts !== 3 || mobile.lengthDeltaCards !== 4 || mobile.templateLayers !== 6 || mobile.templateScopes !== 2 || mobile.templateModes !== 2 || mobile.templateDomainCards !== 4 || mobile.templateDepthCells !== 24 || mobile.historyLayers !== 6 || mobile.historyScopes !== 2 || mobile.historyModes !== 2 || mobile.historyEffects !== 3 || mobile.historyDomainCards !== 4 || mobile.historyDepthCells !== 24 || mobile.distanceLayers !== 6 || mobile.distanceScopes !== 2 || mobile.distanceModes !== 2 || mobile.distanceContrasts !== 2 || mobile.distanceDomainCards !== 4 || mobile.distanceDepthCells !== 24 || mobile.boundaryLayers !== 6 || mobile.boundaryScopes !== 2 || mobile.boundaryModes !== 2 || mobile.boundaryContrasts !== 3 || mobile.boundaryTokenCards !== 4 || mobile.boundaryDomainCards !== 4 || mobile.boundaryDepthCells !== 24 || mobile.roleLayers !== 6 || mobile.roleScopes !== 2 || mobile.roleModes !== 2 || mobile.roleContrasts !== 3 || mobile.roleLevelCards !== 4 || mobile.roleDomainCards !== 4 || mobile.roleDepthCells !== 24 || mobile.specialLayers !== 6 || mobile.specialScopes !== 2 || mobile.specialModes !== 2 || mobile.specialContrasts !== 4 || mobile.specialTokenCards !== 4 || mobile.specialDomainCards !== 4 || mobile.specialDepthCells !== 24 || mobile.roleBlockLayers !== 6 || mobile.roleBlockScopes !== 2 || mobile.roleBlockModes !== 2 || mobile.roleBlockEffects !== 3 || mobile.roleBlockMatrixCards !== 4 || mobile.roleBlockDomainCards !== 4 || mobile.roleBlockDepthCells !== 24) failures.push("移动端导航或实验异常");
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
+1 -1
View File
@@ -277,7 +277,7 @@ if (numeric(system.initial.success) <= numeric(system.initial.model) || numeric(
if (numeric(system.cheap.success) >= numeric(system.initial.success) || numeric(system.cheap.cost) !== 4) failures.push("低预算没有降低成功率 / 成本");
if (numeric(system.locked.unsafe) !== 0 || numeric(system.locked.overrefusal) <= numeric(system.initial.overrefusal)) failures.push("安全壳没有展现危险服从 / 过拒权衡");
if (system.keyboardSelected !== "judge" || system.keyboardVisible !== "judge") failures.push("实验键盘 tab 导航异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("局部 uniform routing 进入完整训练后") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
if (home.paperCount !== "486" || home.topicCount !== "17" || papers.total !== 486 || !papers.hasFilter || papers.visible < 80) failures.push("首页 / 论文库评测索引异常");
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
+1 -1
View File
@@ -264,7 +264,7 @@ if (!fleet.k3.avoided.includes("320K") || fleet.k3.shortSlo !== "PROTECTED") fai
if (!fleet.failed.state.includes("SECONDARY RE-PREFILL") || !fleet.failed.recompute.includes("FAILED PRIMARY")) failures.push("缓存故障没有触发原子失效后的重算");
if (fleet.bursty.shortSlo !== "VIOLATED") failures.push("平均并发阈值没有暴露长请求突发");
if (fleet.keyboardSelected !== "phase" || fleet.keyboardVisible !== "phase") failures.push("实验键盘 tab 导航异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("局部 uniform routing 进入完整训练后") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
if (home.paperCount !== "486" || papers.total !== 486 || !papers.hasFilter || papers.visible !== 46) failures.push("论文库推理服务标签或总数异常");
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
@@ -0,0 +1,257 @@
import { writeFileSync } from "node:fs";
const cdpPort = process.env.CDP_PORT ?? "9231";
const baseUrl = process.env.SITE_URL ?? "http://127.0.0.1:4330";
const pages = await fetch(`http://127.0.0.1:${cdpPort}/json/list`).then((response) => response.json());
const page = pages.find((entry) => entry.type === "page");
if (!page) throw new Error(`CDP ${cdpPort} 没有可用页面`);
const socket = new WebSocket(page.webSocketDebuggerUrl);
await new Promise((resolve, reject) => {
socket.addEventListener("open", resolve, { once: true });
socket.addEventListener("error", reject, { once: true });
});
let nextId = 0;
const pending = new Map();
const exceptions = [];
socket.addEventListener("message", (event) => {
const message = JSON.parse(event.data);
if (message.id && pending.has(message.id)) {
const { resolve, reject } = pending.get(message.id);
pending.delete(message.id);
if (message.error) reject(new Error(message.error.message));
else resolve(message.result);
}
if (message.method === "Runtime.exceptionThrown") {
exceptions.push(message.params.exceptionDetails.exception?.description ?? message.params.exceptionDetails.text);
}
});
const command = (method, params = {}) => new Promise((resolve, reject) => {
const id = ++nextId;
pending.set(id, { resolve, reject });
socket.send(JSON.stringify({ id, method, params }));
});
const pause = (milliseconds) => new Promise((resolve) => setTimeout(resolve, milliseconds));
const evaluate = async (expression) => {
const result = await command("Runtime.evaluate", { expression, returnByValue: true, awaitPromise: true });
if (result.exceptionDetails) throw new Error(result.exceptionDetails.exception?.description ?? result.exceptionDetails.text);
return result.result.value;
};
const navigate = async (path) => {
await command("Page.navigate", { url: `${baseUrl}${path}` });
for (let attempt = 0; attempt < 100; attempt += 1) {
await pause(100);
if (await evaluate("document.readyState === 'complete'")) return;
}
throw new Error(`${path} 加载超时`);
};
const screenshot = async (path) => {
const result = await command("Page.captureScreenshot", { format: "png", captureBeyondViewport: false });
writeFileSync(path, Buffer.from(result.data, "base64"));
};
await command("Page.enable");
await command("Runtime.enable");
await command("Emulation.setDeviceMetricsOverride", {
width: 1440,
height: 1100,
deviceScaleFactor: 1,
mobile: false,
});
await navigate("/k3/");
const desktop = await evaluate(`(() => {
const root = document.querySelector("[data-forward-lab]");
root.scrollIntoView({ block: "start", behavior: "instant" });
window.scrollBy(0, -78);
const text = (selector) => root.querySelector(selector)?.textContent.replace(/\\s+/g, " ").trim();
const panel = () => root.querySelector("[data-forward-panel]:not([hidden])")?.dataset.forwardPanel;
const setSelect = (selector, value) => {
const node = root.querySelector(selector);
node.value = value;
node.dispatchEvent(new Event("change", { bubbles: true }));
};
const initial = {
panel: panel(),
tabs: root.querySelectorAll("[data-forward-tab]").length,
panels: root.querySelectorAll("[data-forward-panel]").length,
ledger: root.querySelectorAll(".forward-ledger article").length,
targetGroups: root.querySelectorAll(".group-map article.target").length,
spikeLayers: root.querySelectorAll(".group-map i.spike").length,
variants: root.querySelectorAll(".variant-grid article").length,
boundary: root.textContent.includes("K3 的训练尖峰已被定位") &&
root.textContent.includes("Figure 5(c)") &&
root.textContent.includes("训练期架构消融"),
};
root.querySelector('[data-forward-tab="trajectory"]').click();
const trajectoryInitial = {
panel: panel(),
state: text("[data-forward-trajectory-state]"),
points: root.querySelectorAll("[data-forward-trajectory-series] circle").length,
lines: root.querySelectorAll("[data-forward-trajectory-series] polyline").length,
readouts: [...root.querySelectorAll("[data-forward-trajectory-readout] article")]
.map((node) => node.textContent.replace(/\\s+/g, " ").trim()),
};
setSelect("[data-forward-trajectory-seed]", "2026073002");
root.querySelector('[data-forward-trajectory-metric="peak_normalized"]').click();
const trajectoryChanged = {
state: text("[data-forward-trajectory-state]"),
readouts: [...root.querySelectorAll("[data-forward-trajectory-readout] article")]
.map((node) => node.textContent.replace(/\\s+/g, " ").trim()),
};
root.querySelector('[data-forward-tab="gate"]').click();
const gate = {
panel: panel(),
status: text(".status-banner"),
rows: root.querySelectorAll(".gate-table tbody tr").length,
passedRows: root.querySelectorAll(".gate-table tbody td.good:last-child").length,
quality: text(".gate-layout aside"),
claims: root.querySelectorAll(".claim-pair article").length,
};
root.querySelector('[data-forward-tab="interaction"]').click();
const interactionInitial = {
panel: panel(),
state: text("[data-forward-interaction-state]"),
cells: root.querySelectorAll("[data-forward-interaction-cells] article").length,
summary: text("[data-forward-interaction-summary]"),
};
root.querySelector('[data-forward-interaction-metric="peak_normalized"]').click();
const interactionChanged = {
state: text("[data-forward-interaction-state]"),
summary: text("[data-forward-interaction-summary]"),
};
root.querySelector('[data-forward-tab="spectrum"]').click();
const spectrumInitial = {
panel: panel(),
state: text("[data-forward-spectrum-state]"),
points: root.querySelectorAll("[data-forward-spectrum-points] circle").length,
contrast: text("[data-forward-spectrum-contrast]"),
peak: text("[data-forward-spectrum-peak]"),
layer: text("[data-forward-spectrum-layer]"),
audits: root.querySelectorAll(".audit-grid article").length,
replay: root.textContent.includes("scientific exact"),
};
setSelect("[data-forward-spectrum-seed]", "2026073002");
setSelect("[data-forward-spectrum-variant]", "uniform_groups_6_7_forward");
const spectrumChanged = {
state: text("[data-forward-spectrum-state]"),
points: root.querySelectorAll("[data-forward-spectrum-points] circle").length,
contrast: text("[data-forward-spectrum-contrast]"),
peak: text("[data-forward-spectrum-peak]"),
layer: text("[data-forward-spectrum-layer]"),
};
const first = root.querySelector('[data-forward-tab="contract"]');
first.focus();
first.dispatchEvent(new KeyboardEvent("keydown", { key: "ArrowRight", bubbles: true }));
const keyboard = {
selected: root.querySelector('[data-forward-tab][aria-selected="true"]').dataset.forwardTab,
panel: panel(),
};
return {
initial, trajectoryInitial, trajectoryChanged, gate,
interactionInitial, interactionChanged, spectrumInitial, spectrumChanged, keyboard,
documentOverflow: document.documentElement.scrollWidth - document.documentElement.clientWidth,
rootOverflow: root.scrollWidth - root.clientWidth,
};
})()`);
await pause(180);
await screenshot("/tmp/llm-atlas-k3-attnres-forward-desktop.png");
await command("Emulation.setDeviceMetricsOverride", {
width: 390,
height: 844,
deviceScaleFactor: 1,
mobile: true,
});
await navigate("/k3/");
const mobile = await evaluate(`(() => {
const root = document.querySelector("[data-forward-lab]");
root.scrollIntoView({ block: "start", behavior: "instant" });
window.scrollBy(0, -64);
root.querySelector('[data-forward-tab="gate"]').click();
const table = root.querySelector(".gate-table-wrap");
const gateTableScrolls = table.scrollWidth > table.clientWidth;
root.querySelector('[data-forward-tab="spectrum"]').click();
return {
tabs: root.querySelectorAll("[data-forward-tab]").length,
visiblePanel: root.querySelector("[data-forward-panel]:not([hidden])")?.dataset.forwardPanel,
points: root.querySelectorAll("[data-forward-spectrum-points] circle").length,
audits: root.querySelectorAll(".audit-grid article").length,
gateTableScrolls,
documentOverflow: document.documentElement.scrollWidth - document.documentElement.clientWidth,
rootOverflow: root.scrollWidth - root.clientWidth,
};
})()`);
await pause(180);
await screenshot("/tmp/llm-atlas-k3-attnres-forward-mobile.png");
const failures = [];
if (desktop.initial.panel !== "contract" || desktop.initial.tabs !== 5 || desktop.initial.panels !== 5 ||
desktop.initial.ledger !== 6 || desktop.initial.targetGroups !== 2 ||
desktop.initial.spikeLayers !== 5 || desktop.initial.variants !== 4) {
failures.push("五视图、账本、group 或 selector map 结构异常");
}
if (!desktop.initial.boundary) failures.push("reduced-model / K3 / Figure 5(c) claim boundary 缺失");
if (desktop.trajectoryInitial.panel !== "trajectory" || desktop.trajectoryInitial.points !== 24 ||
desktop.trajectoryInitial.lines !== 4 || desktop.trajectoryInitial.readouts.length !== 4 ||
!desktop.trajectoryInitial.state.includes("2026073001") ||
!desktop.trajectoryInitial.readouts.some((value) => value.includes("+73.5%"))) {
failures.push("seed 1 contrast 训练轨迹异常");
}
if (!desktop.trajectoryChanged.state.includes("2026073002") ||
!desktop.trajectoryChanged.state.includes("PEAK / MEAN") ||
!desktop.trajectoryChanged.readouts.some((value) => value.includes("+62.0%"))) {
failures.push("trajectory seed / metric 切换异常");
}
if (desktop.gate.panel !== "gate" || desktop.gate.rows !== 6 || desktop.gate.passedRows !== 6 ||
desktop.gate.claims !== 2 || !desktop.gate.status.includes("ATTENUATION ESTABLISHED") ||
!desktop.gate.status.includes("6 / 6") || !desktop.gate.status.includes("4 / 4") ||
!desktop.gate.quality.includes("+0.0066")) {
failures.push("冻结主门或 BPC quality readout 异常");
}
if (desktop.interactionInitial.panel !== "interaction" || desktop.interactionInitial.cells !== 3 ||
!desktop.interactionInitial.state.includes("8,000") ||
!desktop.interactionInitial.summary.includes("-0.367") ||
!desktop.interactionChanged.state.includes("PEAK / MEAN") ||
!desktop.interactionChanged.summary.includes("-0.170")) {
failures.push("I67 描述性 residual 切换异常");
}
if (desktop.spectrumInitial.panel !== "spectrum" || desktop.spectrumInitial.points !== 32 ||
desktop.spectrumInitial.contrast !== "3.046×" || desktop.spectrumInitial.peak !== "3.219×" ||
desktop.spectrumInitial.layer !== "L21" || desktop.spectrumInitial.audits !== 5 ||
!desktop.spectrumInitial.replay) {
failures.push("reference 32 层谱或 replay 审计异常");
}
if (desktop.spectrumChanged.points !== 32 || !desktop.spectrumChanged.state.includes("2026073002") ||
!desktop.spectrumChanged.state.includes("GROUPS 6+7") ||
desktop.spectrumChanged.contrast !== "0.758×" || desktop.spectrumChanged.peak !== "1.375×" ||
desktop.spectrumChanged.layer !== "L9") {
failures.push("joint variant spectrum 切换异常");
}
if (desktop.keyboard.selected !== "trajectory" || desktop.keyboard.panel !== "trajectory") {
failures.push("键盘 tab 切换异常");
}
if (desktop.documentOverflow > 1 || mobile.documentOverflow > 1) failures.push("页面出现文档级横向溢出");
if (desktop.rootOverflow > 1 || mobile.rootOverflow > 1) failures.push("Round 08 实验室出现横向溢出");
if (mobile.tabs !== 5 || mobile.visiblePanel !== "spectrum" || mobile.points !== 32 ||
mobile.audits !== 5 || !mobile.gateTableScrolls) {
failures.push("390px 移动端布局或局部表格滚动异常");
}
if (exceptions.length) failures.push(`运行时异常:${exceptions.join(" | ")}`);
console.log(JSON.stringify({ desktop, mobile, exceptions }, null, 2));
socket.close();
if (failures.length) {
console.error(`FAIL K3 AttnRes forward-training browser\n- ${failures.join("\n- ")}`);
process.exit(1);
}
console.log("PASS K3 AttnRes forward-training browser interactions");
+106
View File
@@ -0,0 +1,106 @@
import { createHash } from "node:crypto";
import { readdirSync, readFileSync } from "node:fs";
const hash = (bytes) => createHash("sha256").update(bytes).digest("hex");
const read = (path) => {
const bytes = readFileSync(new URL(path, import.meta.url));
return { bytes, json: JSON.parse(bytes), sha256: hash(bytes) };
};
const aggregate = read("../src/data/k3-attnres-forward.json");
const compact = read("../src/data/k3-attnres-forward-compact.json");
const reproduction = read("../experiments/k3/attnres_forward/reproduction.json");
const manifest = read("../experiments/k3/attnres_forward/manifest.json");
const rawDirectory = new URL("../experiments/k3/attnres_forward/results/raw/", import.meta.url);
const failures = [];
const expect = (condition, message) => {
if (!condition) failures.push(message);
};
const close = (actual, expected, tolerance = 1e-15) =>
Math.abs(actual - expected) <= tolerance;
expect(aggregate.sha256 === "f8df928adb8a563d851bb3c1abbf40bcada33626d9177821e4341d854980a097", "aggregate physical SHA-256 changed");
expect(compact.sha256 === "664f6d6226df7c0c9aba6314d54a1cb8ae90f6922016dfa7c7a823729606f0c1", "compact physical SHA-256 changed");
expect(reproduction.sha256 === "b029c333596dd1de957efc2b42f7ebfc1d32f241c687f06021c4d08ffa27b1d0", "reproduction physical SHA-256 changed");
expect(manifest.sha256 === "49546ed5baf36bcb30885062b7c671fafe4ccff2e606624cd7dbe23717f9a712", "manifest physical SHA-256 changed");
expect(aggregate.json.canonical_sha256_without_self === "eecf05c623e473ec5eba8afa54d555d2733a0af2fc5488398f1da095206d50c8", "aggregate canonical SHA-256 changed");
expect(compact.json.canonical_sha256_without_self === "c3e672adb6879f99efccf3bd2e0fabaab1a1159225a8ddaf4f5f9ea7a002784b", "compact canonical SHA-256 changed");
expect(reproduction.json.canonical_sha256_without_self === "57346df80c0d76bd1d306d5fa213feed74ac2094237c22ebcb49f16ea16437e0", "reproduction canonical SHA-256 changed");
expect(compact.json.protocol_id === "llm-atlas-k3-attnres-forward-training-v1", "protocol identity mismatch");
expect(compact.json.status === "forward_training_attenuation_established_within_reduced_protocol", "frozen status changed");
expect(compact.json.primary_step === 8000, "primary step changed");
expect(compact.json.spike_layers_1based.join(",") === "21,22,23,24,25", "spike window changed");
expect(compact.json.thresholds.material_relative_drop === 0.2, "attenuation threshold changed");
expect(compact.json.processed_target_bytes.formal_12_cells === 786432000, "formal target bytes changed");
expect(compact.json.processed_target_bytes.primary_replay === 65536000, "replay target bytes changed");
expect(compact.json.processed_target_bytes.total === 851968000, "new target bytes changed");
expect(compact.json.aggregate_sha256 === aggregate.json.canonical_sha256_without_self, "compact→aggregate canonical link mismatch");
expect(reproduction.json.scientific_payload_sha256 === compact.json.replay.scientific_payload_sha256, "reproduction→replay hash link mismatch");
expect(reproduction.json.passed && compact.json.replay.passed, "full replay is not exact");
expect(reproduction.json.post_result_grok_review.blocking_errors === 0, "post-result audit reports a blocking error");
expect(reproduction.json.post_result_grok_review.status_confirmed, "post-result audit did not confirm status");
expect(reproduction.json.post_result_grok_review.replay_confirmed, "post-result audit did not confirm replay");
expect(compact.json.metadata_warnings.length === 0, "historical pairing metadata warning appeared");
const rawNames = readdirSync(rawDirectory).filter((name) => name.endsWith(".json")).sort();
expect(rawNames.length === 13, "raw matrix is not 12 formal + 1 replay");
const expectedRaw = [
...aggregate.json.input_files.formal.map((item) => item),
{ ...aggregate.json.input_files.replay, variant: "uniform_groups_6_7_forward", seed: 2026073001 },
];
for (const item of expectedRaw) {
const name = item.path.split("/").at(-1);
const raw = read(`../experiments/k3/attnres_forward/results/raw/${name}`);
expect(raw.sha256 === item.sha256, `${name} physical SHA-256 mismatch`);
expect(/^[0-9a-f]{64}$/.test(raw.json.canonical_sha256_without_self), `${name} canonical self-hash missing`);
expect(raw.json.forward_intervention.passed, `${name} forward audit failed`);
expect(raw.json.forward_intervention.forward_calls === 8054, `${name} forward census changed`);
expect(raw.json.target_bytes_seen === 65536000, `${name} target bytes changed`);
}
const primary = compact.json.primary;
expect(primary.material_response_passed, "primary material response failed");
expect(primary.passed_cells === 6 && primary.required_cells === 6, "primary attenuation is not 6/6");
expect(primary.cells.every((cell) => cell.relative_drop >= 0.2 && cell.passed), "a primary attenuation cell fell below 20%");
expect(primary.quality.passed && primary.quality.passed_checks === 4, "BPC quality gate is not 4/4");
expect(close(primary.quality.mean_delta_bpc, 0.0065784582165467525), "mean BPC delta changed");
expect(close(Math.min(...primary.cells.filter((cell) => cell.metric === "spike_contrast").map((cell) => cell.relative_drop)), 0.6211080443849011), "minimum contrast drop changed");
expect(close(Math.min(...primary.cells.filter((cell) => cell.metric === "peak_normalized").map((cell) => cell.relative_drop)), 0.3234439150908039), "minimum peak drop changed");
expect(Object.values(compact.json.secondary_status).every((value) => value === "secondary_material_response"), "secondary status changed");
expect(compact.json.trajectories.length === 90, "trajectory cell count changed");
expect(compact.json.final_spectra.length === 15, "final spectrum count changed");
expect(compact.json.final_spectra.every((item) => item.normalized.length === 32), "a final spectrum is not 32 layers");
expect(compact.json.interaction.cells.length === 36, "interaction cell count changed");
expect(compact.json.interaction.summaries.length === 12, "interaction summary count changed");
expect(close(
compact.json.interaction.summaries.find((item) =>
item.step === 8000 && item.metric === "spike_contrast").mean_interaction_residual,
-0.3670379672721313,
), "final contrast I67 changed");
expect(close(
compact.json.interaction.summaries.find((item) =>
item.step === 8000 && item.metric === "peak_normalized").mean_interaction_residual,
-0.17044758609080035,
), "final peak I67 changed");
if (failures.length) {
console.error(`FAIL K3 AttnRes forward-training data\n- ${failures.join("\n- ")}`);
process.exit(1);
}
console.log(JSON.stringify({
protocol: compact.json.protocol_id,
status: compact.json.status,
rawRuns: rawNames.length,
attenuation: `${primary.passed_cells}/${primary.required_cells}`,
quality: `${primary.quality.passed_checks}/${primary.quality.required_checks}`,
replayExact: compact.json.replay.passed,
bytes: compact.json.processed_target_bytes,
hashes: {
aggregate: aggregate.sha256,
compact: compact.sha256,
reproduction: reproduction.sha256,
},
}, null, 2));
console.log("PASS K3 AttnRes forward-training frozen data");
+9 -3
View File
@@ -88,6 +88,10 @@ const overview = await evaluate(`(() => ({
localPathTabs: document.querySelectorAll("[data-local-tab]").length,
localPathPanels: document.querySelectorAll("[data-local-panel]").length,
localPathVerdict: document.querySelector("#attnres-local-path")?.textContent.includes("localization 未建立"),
forwardTabs: document.querySelectorAll("[data-forward-tab]").length,
forwardPanels: document.querySelectorAll("[data-forward-panel]").length,
forwardVerdict: document.querySelector("#attnres-forward")?.textContent.includes("6 / 6 PASS") &&
document.querySelector("#attnres-forward")?.textContent.includes("reduced protocol only"),
nativeVisionCorrected: document.body.textContent.includes("MoonViT‑V2 从头训练") &&
document.body.textContent.includes("同一个 next-token prediction objective"),
staleVisionClaim: document.body.textContent.includes("先固定语言模型训练视觉组件"),
@@ -300,8 +304,9 @@ const mobile = await evaluate(`(() => {
gradientTabs: document.querySelectorAll("[data-gradient-tab]").length,
spikeTabs: document.querySelectorAll("[data-spike-tab]").length,
localPathTabs: document.querySelectorAll("[data-local-tab]").length,
forwardTabs: document.querySelectorAll("[data-forward-tab]").length,
offenders: [...document.querySelectorAll("body *")]
.filter((node) => !node.closest(".paper-chain, .spec-table-wrap, .cache-strip, .architecture-explorer, [data-k3-lab], [data-k3-artifact-lab], [data-attnres-lab], [data-gradient-lab], [data-spike-lab], [data-local-path-lab]"))
.filter((node) => !node.closest(".paper-chain, .spec-table-wrap, .cache-strip, .architecture-explorer, [data-k3-lab], [data-k3-artifact-lab], [data-attnres-lab], [data-gradient-lab], [data-spike-lab], [data-local-path-lab], [data-forward-lab]"))
.filter((node) => node.getBoundingClientRect().right > document.documentElement.clientWidth + 1)
.slice(0, 15)
.map((node) => ({
@@ -332,7 +337,7 @@ console.log(JSON.stringify(report, null, 2));
const numeric = (text) => Number.parseFloat(text.replaceAll(",", "").replace("−", "-"));
const failures = [];
if (!overview.title.includes("因果环节")) failures.push("K3 二轮标题异常");
if (overview.sections !== 36 || overview.tocLinks !== 36) failures.push("35 个编号专题加阅读链的目录结构异常");
if (overview.sections !== 37 || overview.tocLinks !== 37) failures.push("36 个编号专题加阅读链的目录结构异常");
if (overview.ledgers !== 32 || overview.reportMap !== 9) failures.push("32 张问题账或报告地图异常");
if (overview.figureAtlas !== 21 || overview.paperLinks !== 100 || overview.paperGroups < 12) failures.push("图表审计或 100 节点阅读链异常");
if (overview.labTabs !== 8 || overview.labPanels !== 8) failures.push("八联实验结构异常");
@@ -341,6 +346,7 @@ if (overview.attnresTabs !== 5 || overview.attnresPanels !== 5) failures.push("A
if (overview.gradientTabs !== 5 || overview.gradientPanels !== 5) failures.push("AttnRes 梯度定义扩展五视图异常");
if (overview.spikeTabs !== 5 || overview.spikePanels !== 5) failures.push("AttnRes 尖峰路径五视图异常");
if (overview.localPathTabs !== 5 || overview.localPathPanels !== 5 || !overview.localPathVerdict) failures.push("AttnRes 局部路径五视图或冻结判定异常");
if (overview.forwardTabs !== 5 || overview.forwardPanels !== 5 || !overview.forwardVerdict) failures.push("AttnRes 训练期前向五视图或冻结判定异常");
if (!overview.nativeVisionCorrected || overview.staleVisionClaim) failures.push("原生多模态纠错未生效或旧错误残留");
if (overview.documentOverflow > 1 || mobile.documentOverflow > 1) failures.push("桌面或移动端存在文档级横向溢出");
if (labs.memoryInitial.panel !== "memory" || numeric(labs.memoryInitial.additiveError) <= numeric(labs.memoryInitial.deltaError)) failures.push("Delta memory 初始递推异常");
@@ -365,7 +371,7 @@ if (artifacts.parameterChanged.shape !== "[96,128] F32" || !artifacts.parameterC
if (artifacts.reproductionInitial.panel !== "reproduction" || numeric(artifacts.reproductionInitial.speedup) !== 1.85 || numeric(artifacts.reproductionInitial.localMean) < 2.6 || !artifacts.reproductionInitial.exactSuite || numeric(artifacts.reproductionInitial.cv) < 2) failures.push("FlashKDA H20、本机 exact suite 或 router 初始探针异常");
if (numeric(artifacts.reproductionChanged.speedup) !== 3.27 || numeric(artifacts.reproductionChanged.flash) !== 0.7064 || numeric(artifacts.reproductionChanged.localMean) >= numeric(artifacts.reproductionInitial.localMean) || !artifacts.reproductionChanged.localMode.includes("FP32 state") || numeric(artifacts.reproductionChanged.cv) <= numeric(artifacts.reproductionInitial.cv) || numeric(artifacts.reproductionChanged.zero) <= numeric(artifacts.reproductionInitial.zero)) failures.push("GB200 benchmark、本机 varlen/state 或 synthetic router counterexample 未更新");
if (artifacts.keyboardSelected !== "tensors" || artifacts.keyboardVisible !== "tensors") failures.push("开放工件键盘 tab 导航异常");
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 8 || mobile.artifactTabs !== 4 || mobile.artifactLayers !== 93 || mobile.attnresTabs !== 5 || mobile.gradientTabs !== 5 || mobile.spikeTabs !== 5 || mobile.localPathTabs !== 5) failures.push("移动端导航或实验异常");
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 8 || mobile.artifactTabs !== 4 || mobile.artifactLayers !== 93 || mobile.attnresTabs !== 5 || mobile.gradientTabs !== 5 || mobile.spikeTabs !== 5 || mobile.localPathTabs !== 5 || mobile.forwardTabs !== 5) failures.push("移动端导航或实验异常");
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
+1 -1
View File
@@ -246,7 +246,7 @@ if (ocr.unreported.status !== "OUT OF EVIDENCE" || ocr.unreported.accuracy !== "
if (loop.toolsStart.state !== "OPEN" || loop.toolsEnd.state !== "VERIFIED" || loop.toolsEnd.evidence !== "97%" || loop.toolsEnd.tools !== "3") failures.push("vision-in-the-loop 终局异常");
if (loop.cotEnd.state !== "FAILED" || !loop.cotEnd.takeaway.includes("不能凭空增加")) failures.push("文字 CoT 与新观察没有分开");
if (loop.keyboardSelected !== "connector" || loop.keyboardVisible !== "connector") failures.push("实验键盘 tab 导航异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("局部 uniform routing 进入完整训练后") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
if (home.paperCount !== "486" || papers.total !== 486 || !papers.hasFilter || papers.multimodalVisible < 59) failures.push("论文库多模态标签或总数异常");
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
+1 -1
View File
@@ -277,7 +277,7 @@ if (layout.navLinks !== 20 || mobile.mobileLinks !== 20 || home.navLinks !== 20)
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") {
if (home.releaseCards !== 17 || !home.firstRelease.includes("局部 uniform routing 进入完整训练后") || home.firstHref !== "/k3/") {
failures.push("首页 Transformer 新章入口异常");
}
if (home.paperCount !== "486") failures.push(`首页论文总数异常:${home.paperCount}`);
+1 -1
View File
@@ -289,7 +289,7 @@ if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentO
}
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强")) failures.push("首页 K3 首发入口异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("局部 uniform routing 进入完整训练后")) failures.push("首页 K3 首发入口异常");
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
socket.close();
+1 -1
View File
@@ -288,7 +288,7 @@ if (numeric(residual.attnres.states) !== 9 || !residual.attnres.routeExplain.inc
if (!residual.clamp.activation.includes("V4") || !residual.clamp.bound.includes("100")) failures.push("DeepSeek-V4 clamp 展示异常");
if (!residual.situ.activation.includes("KIMI") || !residual.situ.bound.includes("100")) failures.push("K3 SiTU 上界展示异常");
if (residual.keyboardSelected !== "position" || residual.keyboardVisible !== "position") failures.push("实验键盘 tab 导航异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("局部 uniform routing 进入完整训练后") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
if (home.paperCount !== "486" || home.topicCount !== "17" || papers.total !== 486 || !papers.hasFilter || papers.visible < 30) failures.push("首页 / 论文库表示索引异常");
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
+1 -1
View File
@@ -273,7 +273,7 @@ if (layout.navLinks !== 20 || mobile.mobileLinks !== 20 || home.navLinks !== 20)
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") {
if (home.releaseCards !== 17 || !home.firstRelease.includes("局部 uniform routing 进入完整训练后") || home.firstHref !== "/k3/") {
failures.push("首页 Transformer 新章入口异常");
}
if (home.paperCount !== "486") failures.push(`首页论文总数异常:${home.paperCount}`);
+1 -1
View File
@@ -233,7 +233,7 @@ if (layout.articleSections !== 16 || layout.paperLinks !== 37 || layout.labTabs
if (layout.documentOverflow > 0 || mobile.documentOverflow > 0 || home.documentOverflow > 0) failures.push("页面存在横向溢出");
if (layout.navGap < 0) failures.push(`桌面导航碰撞:${layout.navGap}px`);
if (!mobile.menuVisible || mobile.menuOpen !== "true") failures.push("移动端菜单不可用");
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强")) failures.push("首页 K3 首发入口异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("局部 uniform routing 进入完整训练后")) failures.push("首页 K3 首发入口异常");
if (exceptions.length) failures.push(`浏览器脚本异常:${exceptions.join("; ")}`);
socket.close();
+1 -1
View File
@@ -236,7 +236,7 @@ if (block.family.trim() !== "Hybrid MoE" || !block.kv.includes("3 KDA : 1 Gated
if (!block.path.some((step) => step.includes("KDA × 3")) || !block.note.includes("AttnRes")) failures.push("K3 Block 路径异常");
if (block.context.trim() !== "128K" || numeric(block.mha) !== 400 || numeric(block.kda) !== 1) failures.push("KV 成本缩放异常");
if (block.keyboardSelected !== "block" || block.keyboardVisible !== "block") failures.push("实验 tab 键盘导航异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("16 个局部 mixer 单侧证据很强") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("局部 uniform routing 进入完整训练后") || home.firstHref !== "/k3/") failures.push("首页 K3 首发入口异常");
if (home.paperCount !== "486" || papers.total !== 486 || papers.transformerVisible < 30) failures.push("论文库或首页论文数量异常");
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4) failures.push("移动端导航或实验异常");
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
+703
View File
@@ -0,0 +1,703 @@
---
import rawLab from "@/data/k3-attnres-forward-compact.json";
const lab = rawLab as any;
const json = JSON.stringify(lab).replaceAll("<", "\\u003c");
const seeds = [...new Set(lab.final_spectra.map((item: any) => item.seed))] as number[];
const variants = [
"uniform_group_6_forward",
"uniform_group_7_forward",
"uniform_groups_6_7_forward",
"uniform_group_7_mlp_forward",
];
const variantLabels: Record<string, string> = {
learned_reference: "HISTORICAL LEARNED REFERENCE",
uniform_group_6_forward: "GROUP 6 · UNIFORM FORWARD",
uniform_group_7_forward: "GROUP 7 · UNIFORM FORWARD",
uniform_groups_6_7_forward: "GROUPS 6+7 · UNIFORM FORWARD",
uniform_group_7_mlp_forward: "GROUP 7 MLP · UNIFORM FORWARD",
};
const statusLabels: Record<string, string> = {
forward_training_attenuation_established_within_reduced_protocol:
"ATTENUATION ESTABLISHED · REDUCED PROTOCOL",
attenuation_not_established: "ATTENUATION NOT ESTABLISHED",
quality_guard_failed: "QUALITY GUARD FAILED",
attenuation_and_quality_failed: "ATTENUATION + QUALITY FAILED",
};
const statusLabel = statusLabels[lab.status] ?? lab.status;
const primary = lab.primary;
const quality = primary.quality;
const meanDrop = (effect: any, metric: string) => {
const cells = effect.cells.filter((cell: any) => cell.metric === metric);
return cells.reduce((sum: number, cell: any) => sum + cell.relative_drop, 0) / cells.length;
};
const percent = (value: number) => `${value >= 0 ? "+" : "−"}${Math.abs(value * 100).toFixed(1)}%`;
const shortHash = (value: string) => `${value.slice(0, 10)}…${value.slice(-8)}`;
---
<figure class="forward-lab" data-forward-lab>
<figcaption>
<span>ROUND 08 / TRAIN-TIME FORWARD</span>
<div>
<h3>不再只改 diagnostic backward:让局部 uniform routing 真正进入 8,000-step 训练</h3>
<p>4 variants × 3 seeds · 1 historical paired reference · 1 full replay · frozen analyzer</p>
</div>
<em>REDUCED-MODEL ARCHITECTURE ABLATION</em>
</figcaption>
<div class="forward-ledger">
<article><span>NEW BYTES</span><b>851.968M</b><p>12 formal + 1 replay</p></article>
<article><span>PRIMARY CELLS</span><b>{primary.passed_cells} / {primary.required_cells}</b><p>contrast + peak</p></article>
<article><span>QUALITY</span><b>{quality.passed_checks} / {quality.required_checks}</b><p>BPC degradation screen</p></article>
<article><span>REPLAY</span><b>{lab.replay.passed ? "EXACT" : "FAILED"}</b><p>scientific payload</p></article>
<article><span>SELECTED PARAMS</span><b>UNREACHABLE</b><p>0 hooks · 0 optimizer state</p></article>
<article class:list={{ pass: primary.material_response_passed, warn: !primary.material_response_passed }}>
<span>STATUS</span><b>{primary.material_response_passed ? "PASS" : "NOT ESTABLISHED"}</b><p>reduced protocol only</p>
</article>
</div>
<div class="forward-tabs" role="tablist" aria-label="Round 08 训练期前向干预实验视图">
<button type="button" role="tab" data-forward-tab="contract" aria-selected="true">01 / FORWARD CONTRACT</button>
<button type="button" role="tab" data-forward-tab="trajectory" aria-selected="false">02 / TRAINING TRAJECTORY</button>
<button type="button" role="tab" data-forward-tab="gate" aria-selected="false">03 / FINAL GATE</button>
<button type="button" role="tab" data-forward-tab="interaction" aria-selected="false">04 / NON-ADDITIVITY</button>
<button type="button" role="tab" data-forward-tab="spectrum" aria-selected="false">05 / SPECTRUM × AUDIT</button>
</div>
<section class="forward-panel" data-forward-panel="contract">
<div class="panel-lead">
<div><span>I / WHAT ACTUALLY CHANGED</span><h4>这是训练期架构消融,不是“只改变 forward”的纯因果实验</h4></div>
<p>选中 mixer 每一次 train / eval / diagnostic 都改成参数无关的均匀读取;新的表示自然改变 backward 与后续更新。</p>
</div>
<div class="equation-pair">
<article>
<span>LEARNED DEPTH MIXER</span>
<code>keys = RMSNorm(sources)</code>
<code>w = softmax(query · keys)</code>
<b>output = Σ wᵢ · sourceᵢ</b>
</article>
<i>→</i>
<article class="uniform">
<span>SELECTED UNIFORM MIXER</span>
<code>logits = zeros(N, B, T)</code>
<code>w = softmax(logits) = 1 / N</code>
<b>output = Σ sourceᵢ / N</b>
</article>
</div>
<div class="group-map">
{[1,2,3,4,5,6,7,8].map((group) => (
<article class:list={{ target: group === 6 || group === 7 }}>
<span>GROUP {group}</span>
<b>L{(group - 1) * 4 + 1}–{group * 4}</b>
<div>
{[0,1,2,3].map((offset) => {
const layer = (group - 1) * 4 + offset + 1;
return <i class:list={{ spike: layer >= 21 && layer <= 25 }}>{layer}</i>;
})}
</div>
<small>{group === 6 ? "#40–47" : group === 7 ? "#48–55" : "LEARNED"}</small>
</article>
))}
</div>
<div class="variant-grid">
{variants.map((variant) => {
const effect = lab.effects[variant];
return (
<article class:list={{ primary: variant === "uniform_groups_6_7_forward" }}>
<span>{variant === "uniform_groups_6_7_forward" ? "PRIMARY" : "SECONDARY"}</span>
<b>{variantLabels[variant]}</b>
<dl>
<div><dt>CONTRAST</dt><dd>{percent(meanDrop(effect, "spike_contrast"))}</dd></div>
<div><dt>PEAK</dt><dd>{percent(meanDrop(effect, "peak_normalized"))}</dd></div>
<div><dt>STATUS</dt><dd>{effect.material_response_passed ? "PASS" : "NOT EST."}</dd></div>
</dl>
</article>
);
})}
</div>
<div class="reachability-flow">
<article><span>SELECTED QUERY / KEY NORM</span><b>留在 AdamW param groups</b><p>保持 optimizer 结构与 reference 一致。</p></article>
<i>×</i>
<article><span>COMPUTATION GRAPH</span><b>没有被 forward 调用</b><p>不是 stop-gradient;参数结构性不可达。</p></article>
<i>→</i>
<article><span>FINAL AUDIT</span><b>0 hook · 0 state · byte-exact init</b><p>不能写成“训练了但没有移动”。</p></article>
</div>
</section>
<section class="forward-panel" data-forward-panel="trajectory" hidden>
<div class="panel-lead">
<div><span>II / SIX FROZEN CHECKPOINTS</span><h4>不是只看终点:局部前向干预从什么时候开始分化?</h4></div>
<p>纵轴是相对同 seed historical learned reference 的下降;正值表示 attenuation,负值表示 amplification。</p>
</div>
<div class="forward-controls">
<label>SEED
<select data-forward-trajectory-seed>
{seeds.map((seed) => <option value={String(seed)}>{seed}</option>)}
</select>
</label>
<div>
<button type="button" data-forward-trajectory-metric="spike_contrast" aria-pressed="true">SPIKE CONTRAST</button>
<button type="button" data-forward-trajectory-metric="peak_normalized" aria-pressed="false">PEAK / MEAN</button>
</div>
<span data-forward-trajectory-state></span>
</div>
<div class="trajectory-chart">
<header><b>RELATIVE DROP VS PAIRED REFERENCE</b><span>checkpoints are equally spaced; labels preserve actual steps</span></header>
<svg viewBox="0 0 960 390" role="img" aria-label="四个训练变体的尖峰指标轨迹" data-forward-trajectory-chart>
<g data-forward-trajectory-grid></g>
<g data-forward-trajectory-series></g>
</svg>
<div class="series-key">
<span><i class="g6"></i>GROUP 6</span>
<span><i class="g7"></i>GROUP 7</span>
<span><i class="joint"></i>GROUPS 6+7</span>
<span><i class="mlp"></i>GROUP 7 MLP</span>
</div>
</div>
<div class="trajectory-readout" data-forward-trajectory-readout></div>
<div class="boundary-note">
<b>轨迹不参与主闸门</b>
<p>主 status 只读取 step 8,000;中间 checkpoint 用来观察适应过程,不能挑一个最好看的时点替代终点。</p>
</div>
</section>
<section class="forward-panel" data-forward-panel="gate" hidden>
<div class="panel-lead">
<div><span>III / PREREGISTERED FINAL VERDICT</span><h4>20% attenuation 与 BPC degradation screen 必须同时通过</h4></div>
<p>任何结构、hash、selector、finite 或 replay 错误都会让 analyzer 直接失败,不进入下表的科学状态。</p>
</div>
<div class:list={{ "status-banner": true, pass: primary.material_response_passed, warn: !primary.material_response_passed }}>
<span>FROZEN ANALYZER STATUS</span>
<b>{statusLabel}</b>
<p>ATTENUATION {primary.passed_cells} / {primary.required_cells} · QUALITY {quality.passed_checks} / {quality.required_checks}</p>
</div>
<div class="gate-layout">
<div class="gate-table-wrap">
<table class="gate-table">
<thead><tr><th>SEED</th><th>METRIC</th><th>REFERENCE</th><th>VARIANT</th><th>DROP</th><th>≥20%</th></tr></thead>
<tbody>
{primary.cells.map((cell: any) => (
<tr>
<th>{cell.seed}</th>
<td>{cell.metric === "spike_contrast" ? "CONTRAST" : "PEAK"}</td>
<td>{cell.reference.toFixed(3)}×</td>
<td>{cell.variant.toFixed(3)}×</td>
<td class:list={{ good: cell.passed, bad: !cell.passed }}>{percent(cell.relative_drop)}</td>
<td class:list={{ good: cell.passed, bad: !cell.passed }}>{cell.passed ? "PASS" : "FAIL"}</td>
</tr>
))}
</tbody>
</table>
</div>
<aside>
<span>QUALITY SCREEN</span>
<b>final validation BPC</b>
{Object.entries(quality.per_seed).map(([seed, item]: [string, any]) => (
<div><em>{seed}</em><strong class:list={{ good: item.passed, bad: !item.passed }}>{item.delta_bpc >= 0 ? "+" : ""}{item.delta_bpc.toFixed(4)}</strong><small>≤ +.050</small></div>
))}
<div class="mean"><em>3-SEED MEAN</em><strong class:list={{ good: quality.mean_passed, bad: !quality.mean_passed }}>{quality.mean_delta_bpc >= 0 ? "+" : ""}{quality.mean_delta_bpc.toFixed(4)}</strong><small>≤ +.030</small></div>
</aside>
</div>
<div class="numerator-denominator">
<article><span>SPIKE WINDOW</span><b>S = layers 21–25</b><p>报告 mean(g[S]),但不单独作为 status。</p></article>
<i>÷</i>
<article><span>REFERENCE WINDOW</span><b>R = other 27 layers</b><p>layers 26–28 也被 joint intervention 改写。</p></article>
<i>=</i>
<article class="warning"><span>CONTRAST</span><b>不是“尖峰层关闭”</b><p>下降可能来自 S 降、R 升,或两者同时发生。</p></article>
</div>
<div class="claim-pair">
<article class="yes"><span>可以说</span><b>{statusLabel}</b><p>只限这个固定缩小模型、数据、预算、seed 与 operational metric。</p></article>
<article class="no"><span>不能说</span><b>K3 的训练尖峰已被定位</b><p>没有真实 2.8T checkpoint forward,也没有复现未公开 Figure 5(c) telemetry。</p></article>
</div>
</section>
<section class="forward-panel" data-forward-panel="interaction" hidden>
<div class="panel-lead">
<div><span>IV / DESCRIPTIVE LOG RESIDUAL</span><h4>joint effect 等不等于两个 single-run effects 相加?</h4></div>
<p>三条 effect 来自三套独立训练;这里画的是跨 run 的 bookkeeping residual,不是因果 interaction 或 Shapley contribution。</p>
</div>
<div class="forward-controls">
<label>STEP
<select data-forward-interaction-step>
{[0,100,500,2000,4000,8000].map((step) => <option value={String(step)}>{step.toLocaleString()}</option>)}
</select>
</label>
<div>
<button type="button" data-forward-interaction-metric="spike_contrast" aria-pressed="true">SPIKE CONTRAST</button>
<button type="button" data-forward-interaction-metric="peak_normalized" aria-pressed="false">PEAK / MEAN</button>
</div>
<span data-forward-interaction-state></span>
</div>
<div class="interaction-equation">
<span>E₆₇</span><i>−</i><span>E₆</span><i>−</i><span>E₇</span><b>= I₆₇</b>
<p>E = ln(Xref / Xvariant);I &gt; 0 表示 joint log attenuation 大于两个 single-run effects 的和。</p>
</div>
<div class="interaction-cells" data-forward-interaction-cells></div>
<div class="interaction-summary" data-forward-interaction-summary></div>
<div class="boundary-note">
<b>没有预注册通过线</b>
<p>I₆₇ 只报告原值、三 seed mean 与 range;正负都不能翻译成 group 6 / 7 的真实贡献。</p>
</div>
</section>
<section class="forward-panel" data-forward-panel="spectrum" hidden>
<div class="panel-lead">
<div><span>V / 32-LAYER SHAPE × REPRODUCTION</span><h4>终点平均数之外:峰值移动、S/R 分拆与完整 replay</h4></div>
<p>每条谱按自身 32-layer mean 归一化;选择 reference 或任一训练变体,不改变 analyzer 的正式 status。</p>
</div>
<div class="forward-controls spectrum-controls">
<label>SEED
<select data-forward-spectrum-seed>
{seeds.map((seed) => <option value={String(seed)}>{seed}</option>)}
</select>
</label>
<label>VARIANT
<select data-forward-spectrum-variant>
{["learned_reference", ...variants].map((variant) => <option value={variant}>{variantLabels[variant]}</option>)}
</select>
</label>
<span data-forward-spectrum-state></span>
</div>
<div class="spectrum-layout">
<div class="spectrum-chart">
<header><b>FINAL NORMALIZED ACTIVATION-GRADIENT RMS</b><span>layer mean = 1</span></header>
<svg viewBox="0 0 940 350" role="img" aria-label="训练期前向变体的 32 层梯度谱" data-forward-spectrum-chart>
<rect class="spike-zone" x="0" y="26" width="0" height="280" data-forward-spectrum-zone></rect>
<g data-forward-spectrum-grid></g>
<polyline points="" data-forward-spectrum-line></polyline>
<g data-forward-spectrum-points></g>
</svg>
</div>
<aside>
<span>SELECTED READOUT</span>
<b data-forward-spectrum-label></b>
<dl>
<div><dt>SPIKE MEAN</dt><dd data-forward-spectrum-smean></dd></div>
<div><dt>R MEAN</dt><dd data-forward-spectrum-rmean></dd></div>
<div><dt>CONTRAST</dt><dd data-forward-spectrum-contrast></dd></div>
<div><dt>PEAK / MEAN</dt><dd data-forward-spectrum-peak></dd></div>
<div><dt>PEAK LAYER</dt><dd data-forward-spectrum-layer></dd></div>
</dl>
</aside>
</div>
<div class="audit-grid">
<article><span>STEP-0 IDENTITY</span><b>5 / 5 exact</b><p>logits、CE、validation、32-layer diagnostic 与 capture summary。</p></article>
<article><span>EMPTY SELECTOR</span><b>parent exact</b><p>20-step model、optimizer、history、evaluations 与 hashes。</p></article>
<article><span>PRIMARY REPLAY</span><b>{lab.replay.passed ? "scientific exact" : "FAILED"}</b><p><code>{shortHash(lab.replay.scientific_payload_sha256)}</code></p></article>
<article><span>PROCESSED</span><b>851,968,000 bytes</b><p>历史 references 的 196,608,000 bytes 单列、不重复计入。</p></article>
<article class="boundary"><span>REAL K3</span><b>not executed</b><p><code>A_log [128]↔[96]</code> 尚无官方裁决。</p></article>
</div>
<div class="hash-strip">
<span>PROTOCOL <code>{lab.protocol_id}</code></span>
<span>AGGREGATE <code>{shortHash(lab.aggregate_sha256)}</code></span>
<span>COMPACT <code>{shortHash(lab.canonical_sha256_without_self)}</code></span>
</div>
</section>
<script is:inline type="application/json" data-forward-payload set:html={json}></script>
</figure>
<script>
const initializeForwardLab = (root: HTMLElement) => {
if (root.dataset.ready === "true") return;
root.dataset.ready = "true";
const payload = root.querySelector<HTMLScriptElement>("[data-forward-payload]");
if (!payload) return;
const data = JSON.parse(payload.textContent || "{}");
const variants = [
"uniform_group_6_forward",
"uniform_group_7_forward",
"uniform_groups_6_7_forward",
"uniform_group_7_mlp_forward",
];
const labels: Record<string, string> = {
learned_reference: "HISTORICAL LEARNED REFERENCE",
uniform_group_6_forward: "GROUP 6",
uniform_group_7_forward: "GROUP 7",
uniform_groups_6_7_forward: "GROUPS 6+7",
uniform_group_7_mlp_forward: "GROUP 7 MLP",
};
const colors: Record<string, string> = {
uniform_group_6_forward: "#c98a58",
uniform_group_7_forward: "#8794aa",
uniform_groups_6_7_forward: "#8aae8d",
uniform_group_7_mlp_forward: "#b487a7",
};
const tabs = [...root.querySelectorAll<HTMLButtonElement>("[data-forward-tab]")];
const panels = [...root.querySelectorAll<HTMLElement>("[data-forward-panel]")];
const activate = (name: string, focus = false) => {
tabs.forEach((tab) => {
const selected = tab.dataset.forwardTab === name;
tab.setAttribute("aria-selected", String(selected));
if (selected && focus) tab.focus();
});
panels.forEach((panel) => panel.hidden = panel.dataset.forwardPanel !== name);
};
tabs.forEach((tab, index) => {
tab.addEventListener("click", () => activate(tab.dataset.forwardTab || "contract"));
tab.addEventListener("keydown", (event) => {
if (!["ArrowLeft", "ArrowRight", "Home", "End"].includes(event.key)) return;
event.preventDefault();
let next = index;
if (event.key === "ArrowRight") next = (index + 1) % tabs.length;
if (event.key === "ArrowLeft") next = (index - 1 + tabs.length) % tabs.length;
if (event.key === "Home") next = 0;
if (event.key === "End") next = tabs.length - 1;
activate(tabs[next].dataset.forwardTab || "contract", true);
});
});
const ns = "http://www.w3.org/2000/svg";
const make = (name: string, attributes: Record<string, string>) => {
const node = document.createElementNS(ns, name);
Object.entries(attributes).forEach(([key, value]) => node.setAttribute(key, value));
return node;
};
const setText = (selector: string, value: string) => {
const node = root.querySelector<HTMLElement>(selector);
if (node) node.textContent = value;
};
let trajectoryMetric = "spike_contrast";
const trajectorySeed = root.querySelector<HTMLSelectElement>("[data-forward-trajectory-seed]");
const renderTrajectory = () => {
if (!trajectorySeed) return;
const seed = Number(trajectorySeed.value);
const rows = data.trajectories.filter((row: any) => row.seed === seed && variants.includes(row.variant));
const steps = [...new Set(rows.map((row: any) => row.step))] as number[];
const values = rows.map((row: any) => row.relative_drop[trajectoryMetric]);
const minimum = Math.min(-.05, ...values);
const maximum = Math.max(.25, ...values);
const left = 75, right = 925, top = 35, bottom = 330;
const x = (index: number) => left + index * (right - left) / (steps.length - 1);
const y = (value: number) => bottom - (value - minimum) / (maximum - minimum) * (bottom - top);
const grid = root.querySelector<SVGGElement>("[data-forward-trajectory-grid]");
const series = root.querySelector<SVGGElement>("[data-forward-trajectory-series]");
if (!grid || !series) return;
grid.innerHTML = ""; series.innerHTML = "";
[minimum, 0, .2, maximum].filter((value, index, all) => all.indexOf(value) === index).forEach((value) => {
grid.append(make("line", { x1: String(left), x2: String(right), y1: String(y(value)), y2: String(y(value)), class: value === 0 ? "zero" : "" }));
const label = make("text", { x: "12", y: String(y(value) + 4) });
label.textContent = `${value >= 0 ? "+" : ""}${(value * 100).toFixed(0)}%`;
grid.append(label);
});
steps.forEach((step, index) => {
const label = make("text", { x: String(x(index)), y: "360", "text-anchor": "middle" });
label.textContent = step.toLocaleString();
grid.append(label);
});
variants.forEach((variant) => {
const selected = rows.filter((row: any) => row.variant === variant).sort((a: any, b: any) => a.step - b.step);
const points = selected.map((row: any, index: number) => `${x(index)},${y(row.relative_drop[trajectoryMetric])}`).join(" ");
series.append(make("polyline", { points, fill: "none", stroke: colors[variant], "stroke-width": variant.includes("groups_6_7") ? "3" : "2" }));
selected.forEach((row: any, index: number) => {
const circle = make("circle", { cx: String(x(index)), cy: String(y(row.relative_drop[trajectoryMetric])), r: "4", fill: colors[variant], "data-variant": variant, "data-step": String(row.step) });
series.append(circle);
});
});
const metricLabel = trajectoryMetric === "spike_contrast" ? "SPIKE CONTRAST" : "PEAK / MEAN";
setText("[data-forward-trajectory-state]", `${seed} · ${metricLabel}`);
const readout = root.querySelector<HTMLElement>("[data-forward-trajectory-readout]");
if (readout) {
readout.innerHTML = variants.map((variant) => {
const final = rows.find((row: any) => row.variant === variant && row.step === 8000);
const value = final.relative_drop[trajectoryMetric];
return `<article><span>${labels[variant]}</span><b>${value >= 0 ? "+" : "−"}${Math.abs(value * 100).toFixed(1)}%</b><p>step 8,000 vs same-seed reference</p></article>`;
}).join("");
}
};
trajectorySeed?.addEventListener("change", renderTrajectory);
root.querySelectorAll<HTMLButtonElement>("[data-forward-trajectory-metric]").forEach((button) => {
button.addEventListener("click", () => {
trajectoryMetric = button.dataset.forwardTrajectoryMetric || "spike_contrast";
root.querySelectorAll<HTMLButtonElement>("[data-forward-trajectory-metric]").forEach((candidate) =>
candidate.setAttribute("aria-pressed", String(candidate === button)));
renderTrajectory();
});
});
renderTrajectory();
let interactionMetric = "spike_contrast";
const interactionStep = root.querySelector<HTMLSelectElement>("[data-forward-interaction-step]");
if (interactionStep) interactionStep.value = "8000";
const renderInteraction = () => {
if (!interactionStep) return;
const step = Number(interactionStep.value);
const cells = data.interaction.cells.filter((cell: any) => cell.step === step && cell.metric === interactionMetric);
const container = root.querySelector<HTMLElement>("[data-forward-interaction-cells]");
if (container) {
container.innerHTML = cells.map((cell: any) => {
const residual = cell.interaction_residual;
return `<article data-sign="${residual >= 0 ? "positive" : "negative"}"><span>SEED ${cell.seed}</span><div><em>E6</em><b>${cell.log_effects.group6.toFixed(3)}</b></div><div><em>E7</em><b>${cell.log_effects.group7.toFixed(3)}</b></div><div><em>E67</em><b>${cell.log_effects.groups6_7.toFixed(3)}</b></div><strong>I67 ${residual >= 0 ? "+" : ""}${residual.toFixed(3)}</strong></article>`;
}).join("");
}
const summary = data.interaction.summaries.find((item: any) => item.step === step && item.metric === interactionMetric);
const summaryNode = root.querySelector<HTMLElement>("[data-forward-interaction-summary]");
if (summaryNode && summary) {
summaryNode.innerHTML = `<span>3-SEED DESCRIPTIVE SUMMARY</span><b>mean I67 ${summary.mean_interaction_residual >= 0 ? "+" : ""}${summary.mean_interaction_residual.toFixed(3)}</b><p>range ${summary.minimum.toFixed(3)} → ${summary.maximum.toFixed(3)}</p>`;
}
setText("[data-forward-interaction-state]", `${step.toLocaleString()} · ${interactionMetric === "spike_contrast" ? "SPIKE CONTRAST" : "PEAK / MEAN"}`);
};
interactionStep?.addEventListener("change", renderInteraction);
root.querySelectorAll<HTMLButtonElement>("[data-forward-interaction-metric]").forEach((button) => {
button.addEventListener("click", () => {
interactionMetric = button.dataset.forwardInteractionMetric || "spike_contrast";
root.querySelectorAll<HTMLButtonElement>("[data-forward-interaction-metric]").forEach((candidate) =>
candidate.setAttribute("aria-pressed", String(candidate === button)));
renderInteraction();
});
});
renderInteraction();
const spectrumSeed = root.querySelector<HTMLSelectElement>("[data-forward-spectrum-seed]");
const spectrumVariant = root.querySelector<HTMLSelectElement>("[data-forward-spectrum-variant]");
const renderSpectrum = () => {
if (!spectrumSeed || !spectrumVariant) return;
const seed = Number(spectrumSeed.value);
const variant = spectrumVariant.value;
const record = data.final_spectra.find((item: any) => item.seed === seed && item.variant === variant);
if (!record) return;
const values = record.normalized;
const left = 48, right = 915, top = 28, bottom = 305;
const maximum = Math.max(2, ...values) * 1.08;
const x = (index: number) => left + index * (right - left) / 31;
const y = (value: number) => bottom - value / maximum * (bottom - top);
const grid = root.querySelector<SVGGElement>("[data-forward-spectrum-grid]");
const points = root.querySelector<SVGGElement>("[data-forward-spectrum-points]");
const line = root.querySelector<SVGPolylineElement>("[data-forward-spectrum-line]");
const zone = root.querySelector<SVGRectElement>("[data-forward-spectrum-zone]");
if (!grid || !points || !line || !zone) return;
grid.innerHTML = ""; points.innerHTML = "";
[0, 1, Math.ceil(maximum)].forEach((value) => {
grid.append(make("line", { x1: String(left), x2: String(right), y1: String(y(value)), y2: String(y(value)) }));
const label = make("text", { x: "8", y: String(y(value) + 4) });
label.textContent = `${value}×`; grid.append(label);
});
[1,5,9,13,17,21,25,29,32].forEach((layer) => {
const label = make("text", { x: String(x(layer - 1)), y: "330", "text-anchor": "middle" });
label.textContent = String(layer); grid.append(label);
});
zone.setAttribute("x", String(x(20) - 8));
zone.setAttribute("width", String(x(24) - x(20) + 16));
line.setAttribute("points", values.map((value: number, index: number) => `${x(index)},${y(value)}`).join(" "));
values.forEach((value: number, index: number) => {
points.append(make("circle", { cx: String(x(index)), cy: String(y(value)), r: index + 1 === record.peak_layer_1based ? "4.5" : "2.6", "data-layer": String(index + 1) }));
});
setText("[data-forward-spectrum-state]", `${seed} · ${labels[variant]}`);
setText("[data-forward-spectrum-label]", labels[variant]);
setText("[data-forward-spectrum-smean]", record.spike_mean.toExponential(3));
setText("[data-forward-spectrum-rmean]", record.reference_mean.toExponential(3));
setText("[data-forward-spectrum-contrast]", `${record.spike_contrast.toFixed(3)}×`);
setText("[data-forward-spectrum-peak]", `${record.peak_normalized.toFixed(3)}×`);
setText("[data-forward-spectrum-layer]", `L${record.peak_layer_1based}`);
};
spectrumSeed?.addEventListener("change", renderSpectrum);
spectrumVariant?.addEventListener("change", renderSpectrum);
renderSpectrum();
};
document.querySelectorAll<HTMLElement>("[data-forward-lab]").forEach(initializeForwardLab);
document.addEventListener("astro:page-load", () => {
document.querySelectorAll<HTMLElement>("[data-forward-lab]").forEach(initializeForwardLab);
});
</script>
<style>
.forward-lab {
--ink: #d9d7ce;
--muted: #8e918d;
--line: rgba(217,215,206,.14);
--panel: rgba(14,17,16,.78);
--green: #8aae8d;
--copper: #c98a58;
--blue: #8794aa;
--pink: #b487a7;
--red: #c77768;
margin: 32px 0;
border: 1px solid var(--line);
background: #0d0f0e;
color: var(--ink);
}
.forward-lab h3, .forward-lab h4 { color: var(--ink); }
.forward-lab figcaption { display: grid; grid-template-columns: 150px 1fr auto; gap: 24px; align-items: start; padding: 25px; border-bottom: 1px solid var(--line); }
.forward-lab figcaption > span, .forward-lab figcaption > em { color: var(--copper); font: normal .52rem var(--mono); letter-spacing: .12em; }
.forward-lab figcaption > em { color: var(--muted); text-align: right; }
.forward-lab figcaption h3 { margin: 0 0 8px; font-size: 1.02rem; line-height: 1.35; }
.forward-lab figcaption p { margin: 0; color: var(--muted); font: .56rem/1.5 var(--mono); }
.forward-ledger { display: grid; grid-template-columns: repeat(6, 1fr); border-bottom: 1px solid var(--line); }
.forward-ledger article { min-height: 96px; padding: 15px; border-right: 1px solid var(--line); }
.forward-ledger article:last-child { border-right: 0; }
.forward-ledger span, .variant-grid span, .audit-grid span { color: var(--muted); font: .48rem var(--mono); }
.forward-ledger b { display: block; margin: 11px 0 6px; font: 800 .66rem var(--mono); }
.forward-ledger p { margin: 0; color: var(--muted); font-size: .51rem; }
.forward-ledger .pass b { color: var(--green); }
.forward-ledger .warn b { color: var(--red); }
.forward-tabs { display: grid; grid-template-columns: repeat(5, 1fr); border-bottom: 1px solid var(--line); }
.forward-tabs button { min-height: 48px; border: 0; border-right: 1px solid var(--line); background: transparent; color: var(--muted); font: .51rem var(--mono); cursor: pointer; }
.forward-tabs button:last-child { border-right: 0; }
.forward-tabs button[aria-selected="true"] { background: rgba(201,138,88,.09); color: var(--copper); box-shadow: inset 0 -2px var(--copper); }
.forward-tabs button:focus-visible { outline: 2px solid var(--green); outline-offset: -3px; }
.forward-panel { padding: 26px; }
.panel-lead { display: grid; grid-template-columns: 1fr minmax(260px, 42%); gap: 32px; margin-bottom: 22px; }
.panel-lead span { color: var(--copper); font: .5rem var(--mono); }
.panel-lead h4 { margin: 8px 0 0; font-size: .9rem; }
.panel-lead p { margin: 0; color: var(--muted); font-size: .64rem; line-height: 1.65; }
.equation-pair { display: grid; grid-template-columns: 1fr auto 1fr; gap: 14px; align-items: center; }
.equation-pair article { padding: 18px; border: 1px solid var(--line); background: var(--panel); }
.equation-pair article.uniform { border-color: rgba(138,174,141,.45); }
.equation-pair span { color: var(--muted); font: .5rem var(--mono); }
.equation-pair code, .equation-pair b { display: block; margin-top: 10px; font: .61rem var(--mono); }
.equation-pair code { color: var(--muted); }
.equation-pair b { color: var(--green); }
.equation-pair > i, .reachability-flow > i, .numerator-denominator > i { color: var(--copper); font-style: normal; }
.group-map { display: grid; grid-template-columns: repeat(8, 1fr); margin-top: 18px; border: 1px solid var(--line); }
.group-map article { min-height: 118px; padding: 12px; border-right: 1px solid var(--line); }
.group-map article:last-child { border-right: 0; }
.group-map article.target { background: rgba(201,138,88,.08); box-shadow: inset 0 3px var(--copper); }
.group-map span, .group-map small { color: var(--muted); font: .45rem var(--mono); }
.group-map b { display: block; margin: 9px 0; font: .64rem var(--mono); }
.group-map div { display: grid; grid-template-columns: repeat(4, 1fr); gap: 3px; }
.group-map i { display: grid; place-items: center; aspect-ratio: 1; border: 1px solid var(--line); color: var(--muted); font: normal .44rem var(--mono); }
.group-map i.spike { border-color: rgba(138,174,141,.6); color: var(--green); }
.group-map small { display: block; margin-top: 8px; }
.variant-grid { display: grid; grid-template-columns: repeat(4, 1fr); gap: 12px; margin-top: 18px; }
.variant-grid article { padding: 15px; border: 1px solid var(--line); }
.variant-grid article.primary { border-color: rgba(138,174,141,.5); }
.variant-grid b { display: block; min-height: 34px; margin: 8px 0 12px; font: .6rem/1.4 var(--mono); }
.variant-grid dl { margin: 0; }
.variant-grid dl div { display: flex; justify-content: space-between; padding: 7px 0; border-top: 1px solid var(--line); }
.variant-grid dt, .variant-grid dd { margin: 0; font: .48rem var(--mono); }
.variant-grid dt { color: var(--muted); }
.variant-grid dd { color: var(--green); }
.reachability-flow, .numerator-denominator { display: grid; grid-template-columns: 1fr auto 1fr auto 1fr; gap: 12px; align-items: center; margin-top: 18px; }
.reachability-flow article, .numerator-denominator article { min-height: 108px; padding: 15px; border: 1px solid var(--line); }
.reachability-flow span, .numerator-denominator span { color: var(--muted); font: .47rem var(--mono); }
.reachability-flow b, .numerator-denominator b { display: block; margin: 9px 0; font: .64rem var(--mono); }
.reachability-flow p, .numerator-denominator p { margin: 0; color: var(--muted); font-size: .56rem; line-height: 1.5; }
.numerator-denominator .warning { border-color: rgba(199,119,104,.45); }
.forward-controls { display: flex; align-items: stretch; margin-bottom: 16px; border: 1px solid var(--line); }
.forward-controls label { display: flex; align-items: center; gap: 9px; padding: 10px 13px; border-right: 1px solid var(--line); color: var(--muted); font: .48rem var(--mono); }
.forward-controls select { max-width: 260px; border: 1px solid var(--line); background: #0d0f0e; color: var(--ink); font: .53rem var(--mono); }
.forward-controls > div { display: flex; }
.forward-controls button { border: 0; border-right: 1px solid var(--line); background: transparent; color: var(--muted); font: .48rem var(--mono); }
.forward-controls button[aria-pressed="true"] { color: var(--green); background: rgba(138,174,141,.08); }
.forward-controls > span { display: grid; place-items: center; margin-left: auto; padding: 0 14px; color: var(--green); font: .49rem var(--mono); }
.trajectory-chart, .spectrum-chart { border: 1px solid var(--line); background: var(--panel); }
.trajectory-chart header, .spectrum-chart header { display: flex; justify-content: space-between; padding: 11px 14px; border-bottom: 1px solid var(--line); }
.trajectory-chart header b, .trajectory-chart header span, .spectrum-chart header b, .spectrum-chart header span { font: .49rem var(--mono); }
.trajectory-chart header span, .spectrum-chart header span { color: var(--muted); }
.trajectory-chart svg, .spectrum-chart svg { display: block; width: 100%; height: auto; }
[data-forward-trajectory-grid] :global(line), [data-forward-spectrum-grid] :global(line) { stroke: var(--line); stroke-width: 1; }
[data-forward-trajectory-grid] :global(line.zero) { stroke: rgba(217,215,206,.5); stroke-dasharray: 4 4; }
[data-forward-trajectory-grid] :global(text), [data-forward-spectrum-grid] :global(text) { fill: var(--muted); font: 10px var(--mono); }
.series-key { display: flex; flex-wrap: wrap; gap: 18px; padding: 10px 14px; border-top: 1px solid var(--line); color: var(--muted); font: .48rem var(--mono); }
.series-key span { display: flex; align-items: center; gap: 7px; }
.series-key i { width: 16px; height: 2px; }
.series-key .g6 { background: var(--copper); }
.series-key .g7 { background: var(--blue); }
.series-key .joint { background: var(--green); }
.series-key .mlp { background: var(--pink); }
.trajectory-readout { display: grid; grid-template-columns: repeat(4, 1fr); gap: 12px; margin-top: 14px; }
.trajectory-readout :global(article) { padding: 14px; border: 1px solid var(--line); }
.trajectory-readout :global(span) { color: var(--muted); font: .47rem var(--mono); }
.trajectory-readout :global(b) { display: block; margin: 8px 0; color: var(--green); font: .7rem var(--mono); }
.trajectory-readout :global(p) { margin: 0; color: var(--muted); font-size: .53rem; }
.boundary-note { margin-top: 16px; padding: 15px; border-left: 3px solid var(--copper); background: rgba(201,138,88,.07); }
.boundary-note b { font-size: .64rem; }
.boundary-note p { margin: 6px 0 0; color: var(--muted); font-size: .58rem; line-height: 1.55; }
.status-banner { display: grid; grid-template-columns: 180px 1fr auto; gap: 18px; align-items: center; padding: 16px; border: 1px solid var(--line); }
.status-banner.pass { border-color: rgba(138,174,141,.5); }
.status-banner.warn { border-color: rgba(199,119,104,.5); }
.status-banner span, .status-banner p { color: var(--muted); font: .49rem var(--mono); }
.status-banner b { color: var(--green); font: .72rem var(--mono); }
.status-banner.warn b { color: var(--red); }
.gate-layout { display: grid; grid-template-columns: 1fr 230px; gap: 15px; margin-top: 16px; }
.gate-table-wrap { overflow-x: auto; border: 1px solid var(--line); }
.gate-table { width: 100%; border-collapse: collapse; min-width: 680px; font: .53rem var(--mono); }
.gate-table th, .gate-table td { padding: 12px 10px; border-right: 1px solid var(--line); border-bottom: 1px solid var(--line); text-align: right; }
.gate-table thead { color: var(--muted); }
.good { color: var(--green) !important; }
.bad { color: var(--red) !important; }
.gate-layout aside { padding: 15px; border: 1px solid var(--line); }
.gate-layout aside > span { color: var(--muted); font: .48rem var(--mono); }
.gate-layout aside > b { display: block; margin: 8px 0 14px; font: .63rem var(--mono); }
.gate-layout aside > div { display: grid; grid-template-columns: 1fr auto; gap: 4px 8px; padding: 9px 0; border-top: 1px solid var(--line); }
.gate-layout aside em, .gate-layout aside strong, .gate-layout aside small { font: normal .49rem var(--mono); }
.gate-layout aside small { grid-column: 1 / -1; color: var(--muted); }
.gate-layout aside .mean { margin-top: 5px; }
.claim-pair { display: grid; grid-template-columns: 1fr 1fr; gap: 14px; margin-top: 16px; }
.claim-pair article { padding: 16px; border: 1px solid var(--line); }
.claim-pair span { color: var(--muted); font: .49rem var(--mono); }
.claim-pair b { display: block; margin: 9px 0; font: .65rem var(--mono); }
.claim-pair p { margin: 0; color: var(--muted); font-size: .57rem; line-height: 1.5; }
.claim-pair .yes { border-color: rgba(138,174,141,.45); }
.claim-pair .yes b { color: var(--green); }
.claim-pair .no { border-color: rgba(199,119,104,.45); }
.claim-pair .no b { color: var(--red); }
.interaction-equation { display: flex; flex-wrap: wrap; align-items: center; gap: 12px; padding: 17px; border: 1px solid var(--line); }
.interaction-equation span, .interaction-equation b { padding: 8px 12px; background: var(--panel); font: .68rem var(--mono); }
.interaction-equation i { color: var(--copper); font-style: normal; }
.interaction-equation b { color: var(--green); }
.interaction-equation p { flex-basis: 100%; margin: 0; color: var(--muted); font-size: .56rem; }
.interaction-cells { display: grid; grid-template-columns: repeat(3, 1fr); gap: 14px; margin-top: 16px; }
.interaction-cells :global(article) { padding: 16px; border: 1px solid var(--line); }
.interaction-cells :global(article[data-sign="positive"]) { border-color: rgba(138,174,141,.4); }
.interaction-cells :global(article[data-sign="negative"]) { border-color: rgba(199,119,104,.4); }
.interaction-cells :global(span) { color: var(--muted); font: .48rem var(--mono); }
.interaction-cells :global(div) { display: flex; justify-content: space-between; margin-top: 9px; padding-top: 8px; border-top: 1px solid var(--line); }
.interaction-cells :global(em), .interaction-cells :global(b) { font: normal .52rem var(--mono); }
.interaction-cells :global(em) { color: var(--muted); }
.interaction-cells :global(strong) { display: block; margin-top: 12px; color: var(--green); font: .65rem var(--mono); }
.interaction-summary { margin-top: 14px; padding: 16px; border: 1px solid var(--line); }
.interaction-summary :global(span) { color: var(--muted); font: .48rem var(--mono); }
.interaction-summary :global(b) { display: block; margin: 8px 0; font: .7rem var(--mono); }
.interaction-summary :global(p) { margin: 0; color: var(--muted); font: .52rem var(--mono); }
.spectrum-controls label:nth-child(2) { flex: 1; }
.spectrum-controls label:nth-child(2) select { width: 100%; max-width: none; }
.spectrum-layout { display: grid; grid-template-columns: 1fr 225px; gap: 15px; }
.spike-zone { fill: rgba(201,138,88,.1); }
[data-forward-spectrum-line] { fill: none; stroke: var(--green); stroke-width: 2; }
[data-forward-spectrum-points] :global(circle) { fill: var(--green); stroke: #0d0f0e; stroke-width: 1; }
.spectrum-layout aside { padding: 16px; border: 1px solid var(--line); background: var(--panel); }
.spectrum-layout aside > span { color: var(--muted); font: .48rem var(--mono); }
.spectrum-layout aside > b { display: block; margin: 9px 0 16px; font: .61rem/1.4 var(--mono); }
.spectrum-layout dl { margin: 0; }
.spectrum-layout dl div { display: flex; justify-content: space-between; gap: 8px; padding: 10px 0; border-top: 1px solid var(--line); }
.spectrum-layout dt, .spectrum-layout dd { margin: 0; font: .49rem var(--mono); }
.spectrum-layout dt { color: var(--muted); }
.spectrum-layout dd { color: var(--green); }
.audit-grid { display: grid; grid-template-columns: repeat(5, 1fr); gap: 12px; margin-top: 16px; }
.audit-grid article { padding: 14px; border: 1px solid var(--line); }
.audit-grid b { display: block; margin: 9px 0; font: .63rem var(--mono); }
.audit-grid p { margin: 0; color: var(--muted); font-size: .53rem; line-height: 1.5; }
.audit-grid .boundary { border-color: rgba(199,119,104,.45); }
.hash-strip { display: flex; flex-wrap: wrap; gap: 18px; margin-top: 14px; padding: 12px 14px; border: 1px solid var(--line); color: var(--muted); font: .48rem var(--mono); }
.hash-strip code { color: var(--green); }
@media (max-width: 980px) {
.forward-ledger { grid-template-columns: repeat(3, 1fr); }
.forward-tabs { grid-template-columns: repeat(3, 1fr); }
.group-map { grid-template-columns: repeat(4, 1fr); }
.variant-grid, .trajectory-readout { grid-template-columns: repeat(2, 1fr); }
.gate-layout, .spectrum-layout { grid-template-columns: 1fr; }
.audit-grid { grid-template-columns: repeat(3, 1fr); }
}
@media (max-width: 680px) {
.forward-lab figcaption { grid-template-columns: 1fr; padding: 20px; }
.forward-lab figcaption > span { order: -1; }
.forward-lab figcaption > em { text-align: left; }
.forward-ledger { grid-template-columns: repeat(2, 1fr); }
.forward-tabs { display: flex; overflow-x: auto; }
.forward-tabs button { min-width: 165px; }
.forward-panel { padding: 20px 14px; }
.panel-lead { grid-template-columns: 1fr; gap: 12px; }
.equation-pair, .reachability-flow, .numerator-denominator { grid-template-columns: 1fr; }
.equation-pair > i, .reachability-flow > i, .numerator-denominator > i { text-align: center; transform: rotate(90deg); }
.group-map { grid-template-columns: repeat(2, 1fr); }
.variant-grid, .trajectory-readout, .claim-pair, .interaction-cells, .audit-grid { grid-template-columns: 1fr; }
.forward-controls { flex-wrap: wrap; }
.forward-controls > span { width: 100%; min-height: 34px; border-top: 1px solid var(--line); }
.status-banner { grid-template-columns: 1fr; }
.gate-table { min-width: 680px; }
}
</style>
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+9 -9
View File
@@ -128,20 +128,20 @@ const paths = [
<div class="release-grid">
<a class="release-card k3-release" href="/k3/">
<div>
<p class="eyebrow"><span>NEW / K3 ROUND 07</span> LOCAL MIXER PATH · BIDIRECTIONAL GATE</p>
<h2>16 个局部 mixer 单侧证据很强,但双向定位仍然没有闭合</h2>
<p class="eyebrow"><span>NEW / K3 ROUND 08</span> TRAIN-TIME FORWARD · PREREGISTERED GATE</p>
<h2>局部 uniform routing 进入完整训练后,六个尖峰指标格全部衰减</h2>
<p>
以 14 个冻结 mask 同时检查 sufficiency 与 restoration:groups 6+7 在 learned
背景的 6 / 6 格全部超过 50% global log gap;但从 uniform 背景恢复时,peak 三个
seed 全部未过线。因此只能报告 one-sided evidence,不能宣布尖峰已定位到这 16 个 mixer。
四个前向架构变体各跑三个 8,000-step seed,再完整 replay 主格:groups 6+7 的
contrast / peak 六格降幅为 32.3%–77.3%,BPC 质量门 4 / 4 通过。结论只限缩小
depth-32 Block 协议,不是真实 K3 checkpoint 或 Figure 5(c) 复现。
</p>
</div>
<dl>
<div><dt>MATRIX</dt><dd>14 masks × 3 seeds</dd></div>
<div><dt>SUFFICIENCY</dt><dd>6 / 6 pass</dd></div>
<div><dt>RESTORATION</dt><dd>3 / 6 fail</dd></div>
<div><dt>MATRIX</dt><dd>12 formal + 1 replay</dd></div>
<div><dt>ATTENUATION</dt><dd>6 / 6 pass</dd></div>
<div><dt>QUALITY</dt><dd>4 / 4 pass</dd></div>
</dl>
<span class="release-arrow" aria-hidden="true">进入局部路径图、双向门与 32 层原始谱 →</span>
<span class="release-arrow" aria-hidden="true">进入训练轨迹、主门、non-additivity 与 32 层谱 →</span>
</a>
<a class="release-card deepseek-release" href="/deepseek/">
<div>
+33 -6
View File
@@ -2,6 +2,7 @@
import BaseLayout from "@/layouts/BaseLayout.astro";
import ArchitectureExplorer from "@/components/ArchitectureExplorer.astro";
import K3ArtifactLab from "@/components/K3ArtifactLab.astro";
import K3AttnResForwardLab from "@/components/K3AttnResForwardLab.astro";
import K3AttnResGradientLab from "@/components/K3AttnResGradientLab.astro";
import K3AttnResLocalPathLab from "@/components/K3AttnResLocalPathLab.astro";
import K3AttnResSpikeLab from "@/components/K3AttnResSpikeLab.astro";
@@ -44,7 +45,8 @@ const toc = [
["31", "attnres-gradient", "梯度定义与深度扩展"],
["32", "attnres-spike", "尖峰轨迹与反向路径"],
["33", "attnres-local-path", "局部 mixer 双向干预"],
["34", "audit", "21 张图表审计"],
["34", "attnres-forward", "训练期前向干预"],
["35", "audit", "21 张图表审计"],
["↳", "papers", "100 节点阅读链"],
];
@@ -113,13 +115,13 @@ const paperGroups = [
<BaseLayout
title="Kimi K3 技术报告完整深读:架构、训练、RL、系统与评测"
description="用三十二张问题账、二十一张图表审计、八个机制实验、四个开放工件视图、四轮二十个 AttnRes 独立实验视图与一百个一手阅读节点,逐节读懂 Kimi K3。"
description="用三十二张问题账、二十一张图表审计、八个机制实验、四个开放工件视图、五轮二十五个 AttnRes 独立实验视图与一百个一手阅读节点,逐节读懂 Kimi K3。"
section="k3"
>
<header class="page-hero k3-hero">
<div class="page-hero-inner">
<div>
<p class="eyebrow"><span>ANCHOR REPORT / ROUND 07</span> KIMI K3 · REPORT → ARTIFACTS → INDEPENDENT PROBE</p>
<p class="eyebrow"><span>ANCHOR REPORT / ROUND 08</span> KIMI K3 · REPORT → ARTIFACTS → INDEPENDENT PROBE</p>
<h1>不把报告压成摘要<br />把每个因果环节<br />重新展开</h1>
<p class="lead">
K3 同时扩展序列、深度、宽度、视觉与 Agent 轨迹。真正值得读的不是 2.8T 这个最大数字,
@@ -129,11 +131,11 @@ const paperGroups = [
<dl class="page-facts">
<div><dt>QUESTIONS</dt><dd>32 张问题账</dd></div>
<div><dt>REPORT</dt><dd>16 Figures · 5 Tables</dd></div>
<div><dt>LABS</dt><dd>8 + 4 + 5 + 5 + 5 个交互视图</dd></div>
<div><dt>LABS</dt><dd>8 + 4 + 5 + 5 + 5 + 5 个交互视图</dd></div>
<div><dt>READING</dt><dd>100 个一手 / 官方节点</dd></div>
<div><dt>MODEL</dt><dd>2.78T total / 104.2B active</dd></div>
<div><dt>ARTIFACTS</dt><dd>96 shards · 497,220 tensors</dd></div>
<div><dt>STATUS</dt><dd>K3 七轮 · 局部路径审计</dd></div>
<div><dt>STATUS</dt><dd>K3 八轮 · 训练期消融审计</dd></div>
</dl>
</div>
</header>
@@ -977,8 +979,33 @@ const paperGroups = [
</div>
</section>
<section class="article-section" id="attnres-forward">
<p class="eyebrow"><span>34</span> TRAIN-TIME FORWARD INTERVENTION</p>
<h2>诊断期的局部敏感性进入完整训练后,尖峰指标还会衰减吗?</h2>
<p class="lede">
第八轮不再只改 diagnostic backward,而是让 group 6 / 7 的 parameter-free
uniform mixer 进入每一次 train、eval 与 diagnostic forward。四个变体各跑三个
8,000-step seed,并把同 seed 的 Round 05 learned run 锁为历史配对 reference;
主变体 groups 6+7 的 contrast 与 peak 六格降幅全部超过 20%,同时 BPC 质量门
4 / 4 通过,指定 seed 的完整重训 scientific payload exact。
</p>
<div class="artifact-callout">
<article><span>F / FROZEN</span><b>12 formal + 1 replay</b><p>851,968,000 个新 target bytes;每格 65,536,000。</p></article>
<article><span>X / ATTENUATION</span><b>6 / 6 PASS</b><p>contrast drop 62.1%–77.3%;peak drop 32.3%–62.0%。</p></article>
<article><span>X / QUALITY</span><b>4 / 4 PASS</b><p>三 seed ΔBPC 最大 +0.00960;均值 +0.00658。</p></article>
<article class="warning"><span>B / BOUNDARY</span><b>reduced protocol only</b><p>训练期架构消融;不是 K3 checkpoint 或 Figure 5(c) 复现。</p></article>
</div>
<K3AttnResForwardLab />
<div class="hero-actions">
<a class="button primary" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/research/K3_ATTNRES_FORWARD_TRAINING_AUDIT.md">阅读完整结果审计</a>
<a class="button" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/research/K3_ATTNRES_FORWARD_TRAINING_PROTOCOL.md">核对预注册协议</a>
<a class="button" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/research/K3_ATTNRES_FORWARD_TRAINING_IMPLEMENTATION_REVIEW.md">查看两阶段实现审阅</a>
<a class="button" href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/experiments/k3/attnres_forward">复跑训练、分析与 replay</a>
</div>
</section>
<section class="article-section" id="audit">
<p class="eyebrow"><span>34</span> FIGURE & TABLE AUDIT</p>
<p class="eyebrow"><span>35</span> FIGURE & TABLE AUDIT</p>
<h2>Figure 1–16、Table 1–5:每张图究竟支持什么,不能支持什么</h2>
<div class="figure-atlas">
{k3FigureAtlas.map(([id, report, title, contract]) => (
+3 -3
View File
@@ -50,7 +50,7 @@ const workstreams = [
<div><dt>OVERALL</dt><dd>专题平均 {average}%</dd></div>
<div><dt>READABLE</dt><dd>{published} 个首版可读专题</dd></div>
<div><dt>ACTIVE</dt><dd>{researching} 个研究/写作中</dd></div>
<div><dt>UPDATED</dt><dd>2026-07-30 15:20 CST</dd></div>
<div><dt>UPDATED</dt><dd>2026-07-30 14:42 CST</dd></div>
<div><dt>MODE</dt><dd>持续迭代,不锁死版本</dd></div>
</dl>
</div>
@@ -97,7 +97,7 @@ const workstreams = [
<article><span>✓</span><h3>K3 报告已结构化拆解</h3><p>47 页报告目录、151 条参考来源和架构/后训练/系统主线已经提取。</p></article>
<article><span>✓</span><h3>17 专题知识图</h3><p>从语言模型基础到评测安全,包含先修依赖和三条贯穿案例。</p></article>
<article><span>✓</span><h3>编辑式网站系统</h3><p>响应式导航、章节模板、侧栏、进度、论文链和证据提示组件。</p></article>
<article><span>✓</span><h3>一百零九个原创交互视图</h3><p>K3 三轴图、八联报告实验、四联开放工件实验与四轮二十联 AttnRes 独立实验,DeepSeek 四联公式实验、十三联 Base 工件实验、Chat 行为、completion/full-depth、multi-seed、cross-source 与 task-bootstrap CRN 五轮实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
<article><span>✓</span><h3>一百一十四个原创交互视图</h3><p>K3 三轴图、八联报告实验、四联开放工件实验与五轮二十五联 AttnRes 独立实验,DeepSeek 四联公式实验、十三联 Base 工件实验、Chat 行为、completion/full-depth、multi-seed、cross-source 与 task-bootstrap CRN 五轮实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
<article><span>✓</span><h3>十七篇首版长文</h3><p>K3、语言模型前史、Transformer、表示/位置/残差、DeepSeek、Scaling、数据工程、长上下文、MoE、后训练、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全专题。</p></article>
<article><span>✓</span><h3>语言模型前史深度专题</h3><p>八张独立问题账、33 个正式节点、20 段长文与概率—向量—记忆—对齐四联实验。</p></article>
<article><span>✓</span><h3>Transformer 深度专题</h3><p>十张独立问题账、40 个正式节点、21 段正文与 QKV—Mask—多头位置—Block 成本四联实验。</p></article>
@@ -109,6 +109,7 @@ const workstreams = [
<article><span>✓</span><h3>Kimi K3 五轮梯度定义与深度扩展</h3><p>先确认 Figure 5 没有公开唯一 gradient telemetry 合同,再冻结 16/32 blocks × Baseline/Block × 3 seeds 的 12 个 8,000-step 格。Block 的首尾失衡 6/6 改善但全层 CV 6/6 恶化,两个深度都判为 mixed;指定 32 层格完整重训的模型、优化器与全部冻结字段 exact。</p></article>
<article><span>✓</span><h3>Kimi K3 六轮尖峰轨迹与反向路径</h3><p>严格复用 Round 05 depth-32 Block 的三个正式格:尖峰在 step 500 后形成,六个位置 3/3 seed 可见,四种 reduction 12/12 格稳健。切断 key/softmax 源梯度没有降低尖峰;uniform value-backward 让 contrast 平均下降 70.2%,只判为全局 backward-rule sensitivity。完整 replay 的 16 组冻结字段 exact。</p></article>
<article><span>✓</span><h3>Kimi K3 七轮局部路径双向审计</h3><p>冻结 14 个 same-forward mask,把 groups 6+7 的 16 个 depth mixers 同时放进 sufficiency 与 restoration 两个方向。充分性 6/6 过 50%,恢复性却只有 contrast 3/3 通过、peak 0/3 通过,因此正式状态为 one-sided evidence / localization not established。三个正式格、完整 replay、selector 与 forward identity 全部 exact。</p></article>
<article><span>✓</span><h3>Kimi K3 八轮训练期前向干预</h3><p>四个 forward architecture variants × 三 seed × 8,000 steps,加一格完整 replay;groups 6+7 的 contrast / peak 六格降幅全部超过 20%,BPC 质量门 4/4 通过。13 个 raw、自哈希、historical pairing 与 scientific replay 全部过闸;Grok 结果后复算 blocking error 为 0。结论只限固定缩小协议,不是真实 K3 checkpoint 或 Figure 5(c) 复现。</p></article>
<article><span>✓</span><h3>FlashKDA RTX 5090 执行闸门</h3><p>隔离 CUDA 13.0 / glibc 2.39 编译 sm_120a wheel;6/6 官方参考逐元素相等,并完成 fixed / varlen、三种 state mode 的 1,800 个 CUDA Event samples。</p></article>
<article><span>✓</span><h3>Scaling Laws 深度专题</h3><p>九张账、29 个一手节点、DeepSeek/Kimi 双谱系与曲面—部署—复用—涌现四联实验。</p></article>
<article><span>✓</span><h3>数据工程深度专题</h3><p>十二张账、31 个一手节点、DeepSeek/Kimi 双谱系与流水线—去重—混合—改写四联实验。</p></article>
@@ -137,7 +138,6 @@ const workstreams = [
</div>
<div class="queue-table">
<div class="head"><b>优先级</b><b>专题</b><b>本轮交付</b><b>完成闸门</b></div>
<div><span>P0</span><strong>K3 七轮后续</strong><p>前向训练变体 → 非加性局部交互地图 → 等待 A_log 社区候选的官方裁决后进入真实 checkpoint forward</p><em>局部机制 + 工件边界</em></div>
<div><span>P0</span><strong>DeepSeek 八轮后续</strong><p>干预式 mediation → SM90 FlashMLA / FP8 / pipeline traces → R1-like RL 小模型复现</p><em>运行证据 + 独立复现</em></div>
<div><span>P0</span><strong>Transformer 二轮</strong><p>多头电路逐图 → Pre/Post-LN 真实 traces → Flash/KV 配置与 kernel 对照</p><em>逐图笔记 + 实测边界</em></div>
<div><span>P0</span><strong>表示、位置与残差二轮</strong><p>真实 hidden-state / norm traces → 长上下文位置外推 → mHC / AttnRes 深层稳定性消融</p><em>可复现实验 + 逐图笔记</em></div>