feat: audit K3 open model artifacts
This commit is contained in:
+13
-3
@@ -8,7 +8,7 @@
|
|||||||
|---|---:|---:|---|
|
|---|---:|---:|---|
|
||||||
| 研究框架与规范 | 进行中 | 83% | Scaling Laws 二轮拟合复现与逐图精读 |
|
| 研究框架与规范 | 进行中 | 83% | Scaling Laws 二轮拟合复现与逐图精读 |
|
||||||
| 网站设计系统 | 进行中 | 89% | 打印样式与更多通用可视化组件 |
|
| 网站设计系统 | 进行中 | 89% | 打印样式与更多通用可视化组件 |
|
||||||
| Kimi K3 深读 | 完成二轮 | 88% | 第三轮加入官方权重 traces、独立复现与逐图数值重绘 |
|
| Kimi K3 深读 | 三轮实证进行中 | 92% | 匹配 CUDA 12.9+ 执行 FlashKDA,并接入真实 hidden-state / expert-load traces |
|
||||||
| 语言模型前史 | 完成首版 | 78% | Kneser–Ney、LSTM、Bahdanau 逐图精读与真实小语料复现 |
|
| 语言模型前史 | 完成首版 | 78% | Kneser–Ney、LSTM、Bahdanau 逐图精读与真实小语料复现 |
|
||||||
| Transformer 基础 | 完成首版 | 79% | 多头电路、归一化 traces 与真实 kernel / KV 配置 |
|
| Transformer 基础 | 完成首版 | 79% | 多头电路、归一化 traces 与真实 kernel / KV 配置 |
|
||||||
| 表示、位置与残差高速公路 | 完成首版 | 81% | 真实 hidden-state / norm traces、长上下文位置外推与深层稳定性消融 |
|
| 表示、位置与残差高速公路 | 完成首版 | 81% | 真实 hidden-state / norm traces、长上下文位置外推与深层稳定性消融 |
|
||||||
@@ -41,7 +41,7 @@
|
|||||||
- [x] 完成 486 篇关键论文索引,覆盖 16 个标签专题与 Kimi/DeepSeek 聚光主线。
|
- [x] 完成 486 篇关键论文索引,覆盖 16 个标签专题与 Kimi/DeepSeek 聚光主线。
|
||||||
- [x] 完成可检索、可按专题筛选的论文库页面。
|
- [x] 完成可检索、可按专题筛选的论文库页面。
|
||||||
- [x] 完成 K3、语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全十七篇首版长文。
|
- [x] 完成 K3、语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全十七篇首版长文。
|
||||||
- [x] 完成 K3 三轴架构与八联报告实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 四联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等六十七个原创交互视图。
|
- [x] 完成 K3 三轴架构、八联报告实验与四联开放工件实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 四联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等七十一个原创交互视图。
|
||||||
- [x] 完成长上下文首版:五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。
|
- [x] 完成长上下文首版:五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。
|
||||||
- [x] 核验 FlashAttention、DeepSeek-V2/V3.2/V4、Kimi Linear/K3 等六份论文原文,并建立长上下文研究账本。
|
- [x] 核验 FlashAttention、DeepSeek-V2/V3.2/V4、Kimi Linear/K3 等六份论文原文,并建立长上下文研究账本。
|
||||||
- [x] 核验 Switch、ST-MoE、DeepSeekMoE、Loss-Free、V3、LatentMoE 与 K3 原文,并建立 MoE 研究账本。
|
- [x] 核验 Switch、ST-MoE、DeepSeekMoE、Loss-Free、V3、LatentMoE 与 K3 原文,并建立 MoE 研究账本。
|
||||||
@@ -156,10 +156,17 @@
|
|||||||
- [x] 完成 K3 八联交互实验:Delta Rule、bounded decay、Block AttnRes、LatentMoE payload、SiTU-GLU、Quantile Balancing、MOPD/partial rollout、hybrid prefix cache。
|
- [x] 完成 K3 八联交互实验:Delta Rule、bounded decay、Block AttnRes、LatentMoE payload、SiTU-GLU、Quantile Balancing、MOPD/partial rollout、hybrid prefix cache。
|
||||||
- [x] K3 专属 Chrome 断言通过:32/21/100 内容计数、八个实验计算、键盘 tabs、事实纠错、桌面与 390px 移动端均无异常。
|
- [x] K3 专属 Chrome 断言通过:32/21/100 内容计数、八个实验计算、键盘 tabs、事实纠错、桌面与 390px 移动端均无异常。
|
||||||
- [x] K3 二轮以源提交 `b669615`、不可变镜像 `20260729T040336Z-b669615` 发布;OCI digest `sha256:498b7e43…31cdb3c`,NAS、VPS/Tailscale、NPM、DNS、HTTPS、证书、门户、公开 Forgejo 与十六套生产 Chrome 回归全链路通过;保留 `20260729T031901Z-cd96dab` 回滚。
|
- [x] K3 二轮以源提交 `b669615`、不可变镜像 `20260729T040336Z-b669615` 发布;OCI digest `sha256:498b7e43…31cdb3c`,NAS、VPS/Tailscale、NPM、DNS、HTTPS、证书、门户、公开 Forgejo 与十六套生产 Chrome 回归全链路通过;保留 `20260729T031901Z-cd96dab` 回滚。
|
||||||
|
- [x] 启动 K3 三轮开放工件审计:固定 HF revision `9f62e4e9…b3569` 与 FlashKDA revision `1ce47ea3…ffb0b`,通过 config、index、safetensors headers 与 HTTP Range 避免把 1.56 TB 全量下载写成必要前提。
|
||||||
|
- [x] 闭合 93 层配置与 checkpoint 拓扑:69 KDA / 24 MLA、1 dense / 92 MoE、96 shards、497,220 tensor entries、247,296 expert packed tensors 与同数 scales、187 组 AttnRes projection / norm。
|
||||||
|
- [x] 完成 K3 四联开放工件实验:93 层真实配置条带、routed expert / MLA / MoonViT tensor anatomy、小参数范围审计,以及 FlashKDA 作者 benchmark / 本机编译 / synthetic router 探针边界。
|
||||||
|
- [x] 发现并保留官方工件的 `A_log [128]` vs config / remote code / FlashKDA API expected `[96]` 形状不一致;不宣布 checkpoint 损坏,也不把非标准 channel-wise 假设冒充真实 forward。
|
||||||
|
- [x] FlashKDA 本机编译边界已实测:RTX 5090 架构受支持,但当前 CUDA 12.8 低于官方 12.9+;g++ 13 已推进到 nvcc,随后因 CUDA headers / glibc declarations 冲突停止,kernel 尚未执行。
|
||||||
|
- [x] 第三轮证据快照、可复现探针脚本与正式审计账本已进入开源树;原始权重字节不提交,真实观测 O、推导 D、执行 X、合成 S 与未决 U 分开标记。
|
||||||
|
- [x] K3 新四视图本地真实 Chrome 回归通过:93 层条带、tensor group、参数分布、benchmark/router 切换、键盘 tabs、桌面与 390px 移动端均无异常。
|
||||||
|
|
||||||
## 正在进行
|
## 正在进行
|
||||||
|
|
||||||
- [ ] K3 三轮:使用开放权重与官方实现加入 KDA/AttnRes/MoE 真实 traces、FlashKDA kernel 对照、逐图数值重绘与独立复现。
|
- [ ] K3 三轮下一闸门:在匹配 CUDA 12.9+ 环境执行 FlashKDA correctness / benchmark,获得真实 token hidden states、expert load 与 cache traces,再做逐图数值重绘和独立小模型复现。
|
||||||
- [ ] DeepSeek 三轮:真实专家负载、MLA kernel、FP8 / pipeline 与 R1-like RL traces,外加独立小模型复现。
|
- [ ] DeepSeek 三轮:真实专家负载、MLA kernel、FP8 / pipeline 与 R1-like RL traces,外加独立小模型复现。
|
||||||
- [ ] 表示、位置与残差二轮:真实 hidden-state / norm traces、长上下文位置外推复现与 mHC / AttnRes 深层稳定性消融。
|
- [ ] 表示、位置与残差二轮:真实 hidden-state / norm traces、长上下文位置外推复现与 mHC / AttnRes 深层稳定性消融。
|
||||||
- [ ] 评测安全二轮:真实 cross-harness / pass@k 复跑、Judge 元评测、动态污染与过拒案例。
|
- [ ] 评测安全二轮:真实 cross-harness / pass@k 复跑、Judge 元评测、动态污染与过拒案例。
|
||||||
@@ -278,6 +285,9 @@
|
|||||||
| 2026-07-29 | K3 Figure 1–16 / Table 1–5 全部建立课程视觉契约 | 每张图同时写支持范围与不可外推项;作者报告、论文、推导与 toy model 使用 R/P/D/T 标签 |
|
| 2026-07-29 | K3 Figure 1–16 / Table 1–5 全部建立课程视觉契约 | 每张图同时写支持范围与不可外推项;作者报告、论文、推导与 toy model 使用 R/P/D/T 标签 |
|
||||||
| 2026-07-29 | K3 二轮用八个独立实验闭环 | Delta memory、BF16 decay、AttnRes、LatentMoE、SiTU、QB、MOPD/RL 与 prefix cache 分开操作,不合成伪“架构总分” |
|
| 2026-07-29 | K3 二轮用八个独立实验闭环 | Delta memory、BF16 decay、AttnRes、LatentMoE、SiTU、QB、MOPD/RL 与 prefix cache 分开操作,不合成伪“架构总分” |
|
||||||
| 2026-07-29 | K3 二轮用不可变镜像 `20260729T040336Z-b669615` 发布 | OCI digest `sha256:498b7e43…31cdb3c`;复用 `12010→8080`、NPM host 31 / cert 41、门户 order 180 与公开 Forgejo;十六套生产 Chrome 回归通过,保留 `20260729T031901Z-cd96dab` 回滚 |
|
| 2026-07-29 | K3 二轮用不可变镜像 `20260729T040336Z-b669615` 发布 | OCI digest `sha256:498b7e43…31cdb3c`;复用 `12010→8080`、NPM host 31 / cert 41、门户 order 180 与公开 Forgejo;十六套生产 Chrome 回归通过,保留 `20260729T031901Z-cd96dab` 回滚 |
|
||||||
|
| 2026-07-29 | K3 三轮先审计开放工件,不要求加载 1.56 TB | 固定官方 revisions;用 config、index、96 个 shard headers 与两个小范围权重切片闭合真实 topology、shape 和参数统计 |
|
||||||
|
| 2026-07-29 | 开放工件证据使用 O / D / X / S / U 五种身份 | 观测、推导、本机执行、合成探针与未决矛盾不互相冒充;FlashKDA 作者 benchmark 不写成本机 benchmark |
|
||||||
|
| 2026-07-29 | `A_log [128]` 与 expected `[96]` 保持未决 | 并列报告 checkpoint、config、remote code 与 kernel API;等待官方 loader / 修订解释,不擅自 reshape |
|
||||||
|
|
||||||
## 未决问题
|
## 未决问题
|
||||||
|
|
||||||
|
|||||||
@@ -19,8 +19,13 @@
|
|||||||
|
|
||||||
当前里程碑包含 17 专题学习地图、486 篇关键论文索引、Kimi K3 完整导读,
|
当前里程碑包含 17 专题学习地图、486 篇关键论文索引、Kimi K3 完整导读,
|
||||||
语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 技术谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、工具使用与长程 Agent、原生多模态、训练系统、推理服务、数值优化,以及评测与安全深度专题,
|
语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 技术谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、工具使用与长程 Agent、原生多模态、训练系统、推理服务、数值优化,以及评测与安全深度专题,
|
||||||
以及 67 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、
|
以及 71 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、
|
||||||
8 个交互实验和 100 个一手/官方节点,完整覆盖架构、预训练、后训练、系统、评测、案例与附录。
|
8 个交互实验和 100 个一手/官方节点,完整覆盖架构、预训练、后训练、系统、评测、案例与附录。
|
||||||
|
第三轮已完成首个开放工件里程碑:固定官方模型与 FlashKDA revisions,审计 96 个 checkpoint shards、
|
||||||
|
497,220 个 tensor entries、真实 KDA / MLA / MoE / MoonViT shapes 与小范围参数统计,并用 4 个新视图
|
||||||
|
明确区分官方观测、确定性推导、本机执行、合成探针和未决矛盾。详见
|
||||||
|
[K3_ARTIFACT_AUDIT.md](./research/K3_ARTIFACT_AUDIT.md) 与
|
||||||
|
[checkpoint_probe.py](./experiments/k3/checkpoint_probe.py)。
|
||||||
DeepSeek 二轮专题以 24 张问题账、10 次技术转向、
|
DeepSeek 二轮专题以 24 张问题账、10 次技术转向、
|
||||||
4 个交互实验和 60 个一手/官方节点,串起 Dense、MoE、MLA、V3 协同、R1 与 V4。
|
4 个交互实验和 60 个一手/官方节点,串起 Dense、MoE、MLA、V3 协同、R1 与 V4。
|
||||||
其余专题按进度账本持续扩建。
|
其余专题按进度账本持续扩建。
|
||||||
|
|||||||
+2
-1
@@ -154,7 +154,8 @@ pass^k、校准、动态基准、代码 Verifier、LLM Judge、Arena、Agent 最
|
|||||||
## 三条贯穿式案例
|
## 三条贯穿式案例
|
||||||
|
|
||||||
1. **Kimi K3 解剖**:二轮已完成 32 张问题账、Figure 1–16 / Table 1–5 审计、
|
1. **Kimi K3 解剖**:二轮已完成 32 张问题账、Figure 1–16 / Table 1–5 审计、
|
||||||
8 个交互实验与 100 节点阅读链,把全部专题重新汇入架构—预训练—后训练—系统—评测因果链。
|
8 个报告实验与 100 节点阅读链;三轮首个里程碑进一步固定官方 revisions,审计 96 个 checkpoint shards、
|
||||||
|
497,220 个 tensor entries、真实 KDA / MLA / MoE / MoonViT shapes,并用 4 个工件视图公开复现边界。
|
||||||
2. **DeepSeek 技术谱系**:DeepSeek LLM → DeepSeekMoE → V2/MLA → V3/FP8/MTP/DualPipe → Math/GRPO → R1 → V3.2/DSA → V4 长上下文。
|
2. **DeepSeek 技术谱系**:DeepSeek LLM → DeepSeekMoE → V2/MLA → V3/FP8/MTP/DualPipe → Math/GRPO → R1 → V3.2/DSA → V4 长上下文。
|
||||||
3. **“一个 Token 的旅行”**:从文本分词,经注意力、MoE、GPU 集群、后训练,再到线上推理与工具调用。
|
3. **“一个 Token 的旅行”**:从文本分词,经注意力、MoE、GPU 集群、后训练,再到线上推理与工具调用。
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,388 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Build a small, auditable snapshot from Kimi K3's public model artifacts.
|
||||||
|
|
||||||
|
The script intentionally does not download a checkpoint. It consumes:
|
||||||
|
|
||||||
|
1. the public config and safetensors index;
|
||||||
|
2. safetensors JSON headers fetched with HTTP Range;
|
||||||
|
3. two small byte ranges containing one KDA parameter prefix and one MoE
|
||||||
|
router prefix;
|
||||||
|
4. a local checkout of the official FlashKDA repository.
|
||||||
|
|
||||||
|
Raw model bytes stay local. The generated JSON contains only aggregate
|
||||||
|
statistics, public shapes, revisions, checksums, and a clearly labelled
|
||||||
|
synthetic-input router stress probe.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import hashlib
|
||||||
|
import json
|
||||||
|
import math
|
||||||
|
import platform
|
||||||
|
import re
|
||||||
|
import subprocess
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import torch
|
||||||
|
|
||||||
|
|
||||||
|
def parse_args() -> argparse.Namespace:
|
||||||
|
parser = argparse.ArgumentParser()
|
||||||
|
parser.add_argument("--config", type=Path, required=True)
|
||||||
|
parser.add_argument("--index", type=Path, required=True)
|
||||||
|
parser.add_argument("--hf-model", type=Path, required=True)
|
||||||
|
parser.add_argument("--kda-slice", type=Path, required=True)
|
||||||
|
parser.add_argument("--router-prefix", type=Path, required=True)
|
||||||
|
parser.add_argument("--mla-header", type=Path, required=True)
|
||||||
|
parser.add_argument("--vision-header", type=Path, required=True)
|
||||||
|
parser.add_argument("--flashkda-dir", type=Path, required=True)
|
||||||
|
parser.add_argument("--output", type=Path, required=True)
|
||||||
|
parser.add_argument("--synthetic-tokens", type=int, default=2048)
|
||||||
|
parser.add_argument("--seed", type=int, default=20260729)
|
||||||
|
parser.add_argument("--captured-at", default=None)
|
||||||
|
return parser.parse_args()
|
||||||
|
|
||||||
|
|
||||||
|
def read_json(path: Path) -> dict[str, Any]:
|
||||||
|
return json.loads(path.read_text())
|
||||||
|
|
||||||
|
|
||||||
|
def sha256(path: Path) -> str:
|
||||||
|
digest = hashlib.sha256()
|
||||||
|
with path.open("rb") as handle:
|
||||||
|
for block in iter(lambda: handle.read(1024 * 1024), b""):
|
||||||
|
digest.update(block)
|
||||||
|
return digest.hexdigest()
|
||||||
|
|
||||||
|
|
||||||
|
def summarize(values: torch.Tensor) -> dict[str, Any]:
|
||||||
|
flat = values.detach().float().flatten().cpu()
|
||||||
|
points = torch.tensor([0, 0.01, 0.1, 0.25, 0.5, 0.75, 0.9, 0.99, 1])
|
||||||
|
quantiles = torch.quantile(flat, points).tolist()
|
||||||
|
labels = ["min", "p01", "p10", "p25", "p50", "p75", "p90", "p99", "max"]
|
||||||
|
return {
|
||||||
|
"count": flat.numel(),
|
||||||
|
"mean": flat.mean().item(),
|
||||||
|
"std": flat.std().item(),
|
||||||
|
"quantiles": dict(zip(labels, quantiles, strict=True)),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def tensor_fact(header: dict[str, Any], name: str) -> dict[str, Any]:
|
||||||
|
entry = header[name]
|
||||||
|
return {
|
||||||
|
"name": name,
|
||||||
|
"dtype": entry["dtype"],
|
||||||
|
"shape": entry["shape"],
|
||||||
|
"bytes": entry["data_offsets"][1] - entry["data_offsets"][0],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def parse_benchmark(path: Path, heads: int = 96) -> dict[str, Any]:
|
||||||
|
text = path.read_text()
|
||||||
|
section = text.split(f"### `T=8192`, `H={heads}`, `D=128`", 1)[1].split("###", 1)[0]
|
||||||
|
rows = {}
|
||||||
|
for label, flash, chunk, speedup, gdn, gdn_speedup in re.findall(
|
||||||
|
r"\| ([^|]+?) \| ([0-9.]+) \| ([0-9.]+) \| ([0-9.]+)× \| ([0-9.]+) \| ([0-9.]+)× \|",
|
||||||
|
section,
|
||||||
|
):
|
||||||
|
rows[label.strip()] = {
|
||||||
|
"flash_kda_ms": float(flash),
|
||||||
|
"fla_chunk_kda_ms": float(chunk),
|
||||||
|
"speedup_vs_chunk_kda": float(speedup),
|
||||||
|
"fla_chunk_gdn_ms": float(gdn),
|
||||||
|
"speedup_vs_gdn": float(gdn_speedup),
|
||||||
|
}
|
||||||
|
return {"sequence": 8192, "heads": heads, "dimension": 128, "rows": rows}
|
||||||
|
|
||||||
|
|
||||||
|
def load_metrics(load: torch.Tensor) -> dict[str, Any]:
|
||||||
|
mean = load.mean()
|
||||||
|
ordered = load.sort().values
|
||||||
|
count = load.numel()
|
||||||
|
indices = torch.arange(1, count + 1, device=load.device, dtype=torch.float32)
|
||||||
|
gini = ((2 * indices - count - 1) * ordered).sum() / (count * ordered.sum())
|
||||||
|
quantiles = torch.quantile(
|
||||||
|
load,
|
||||||
|
torch.tensor([0, 0.1, 0.25, 0.5, 0.75, 0.9, 0.99, 1], device=load.device),
|
||||||
|
).tolist()
|
||||||
|
labels = ["min", "p10", "p25", "p50", "p75", "p90", "p99", "max"]
|
||||||
|
return {
|
||||||
|
"mean": mean.item(),
|
||||||
|
"std": load.std().item(),
|
||||||
|
"cv": (load.std() / mean).item(),
|
||||||
|
"gini": gini.item(),
|
||||||
|
"zero_experts": int((load == 0).sum()),
|
||||||
|
"quantiles": dict(zip(labels, quantiles, strict=True)),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
args = parse_args()
|
||||||
|
config = read_json(args.config)
|
||||||
|
text = config["text_config"]
|
||||||
|
index = read_json(args.index)
|
||||||
|
hf_model = read_json(args.hf_model)
|
||||||
|
mla_header = read_json(args.mla_header)
|
||||||
|
vision_header = read_json(args.vision_header)
|
||||||
|
|
||||||
|
names = list(index["weight_map"])
|
||||||
|
shard_files = [
|
||||||
|
item
|
||||||
|
for item in hf_model["siblings"]
|
||||||
|
if re.fullmatch(r"model-\d+-of-\d+\.safetensors", item["rfilename"])
|
||||||
|
]
|
||||||
|
shard_sizes = [item["size"] for item in shard_files]
|
||||||
|
|
||||||
|
kda_raw = args.kda_slice.read_bytes()
|
||||||
|
if len(kda_raw) != 49_664:
|
||||||
|
raise ValueError(f"unexpected KDA slice length: {len(kda_raw)}")
|
||||||
|
a_log = torch.frombuffer(bytearray(kda_raw[:512]), dtype=torch.float32).clone()
|
||||||
|
dt_bias = torch.frombuffer(bytearray(kda_raw[512:]), dtype=torch.float32).clone().view(96, 128)
|
||||||
|
|
||||||
|
router_raw = args.router_prefix.read_bytes()
|
||||||
|
if len(router_raw) != 13_488_640:
|
||||||
|
raise ValueError(f"unexpected router prefix length: {len(router_raw)}")
|
||||||
|
correction_bias = torch.frombuffer(
|
||||||
|
bytearray(router_raw[:3584]), dtype=torch.float32
|
||||||
|
).clone()
|
||||||
|
router_weight = torch.frombuffer(
|
||||||
|
bytearray(router_raw[643_584:13_488_640]), dtype=torch.bfloat16
|
||||||
|
).clone().view(896, 7168)
|
||||||
|
|
||||||
|
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
|
||||||
|
torch.manual_seed(args.seed)
|
||||||
|
if device.type == "cuda":
|
||||||
|
torch.cuda.manual_seed_all(args.seed)
|
||||||
|
router_device = router_weight.to(device)
|
||||||
|
bias_device = correction_bias.to(device)
|
||||||
|
synthetic = torch.randn(
|
||||||
|
args.synthetic_tokens, 7168, device=device, dtype=torch.float32
|
||||||
|
)
|
||||||
|
synthetic *= torch.rsqrt(synthetic.square().mean(-1, keepdim=True) + 1e-6)
|
||||||
|
scores = torch.sigmoid(synthetic.to(torch.bfloat16) @ router_device.T).float()
|
||||||
|
unbiased_ids = scores.topk(16, dim=-1).indices
|
||||||
|
biased_ids = (scores + bias_device).topk(16, dim=-1).indices
|
||||||
|
|
||||||
|
def loads(ids: torch.Tensor) -> torch.Tensor:
|
||||||
|
return torch.bincount(ids.flatten(), minlength=896).float()
|
||||||
|
|
||||||
|
unbiased_load = loads(unbiased_ids)
|
||||||
|
biased_load = loads(biased_ids)
|
||||||
|
overlap = torch.tensor(
|
||||||
|
[
|
||||||
|
len(set(unbiased_ids[row].tolist()) & set(biased_ids[row].tolist()))
|
||||||
|
for row in range(args.synthetic_tokens)
|
||||||
|
],
|
||||||
|
dtype=torch.float32,
|
||||||
|
)
|
||||||
|
router_norms = router_weight.float().norm(dim=1)
|
||||||
|
|
||||||
|
# This is deliberately a hypothesis probe, not a canonical forward pass:
|
||||||
|
# the checkpoint stores A_log[128], while public code/API expect A_log[H=96].
|
||||||
|
channelwise_log_decay = -5.0 * torch.sigmoid(torch.exp(a_log).view(1, 128) * dt_bias)
|
||||||
|
channelwise_retention = torch.exp(channelwise_log_decay)
|
||||||
|
|
||||||
|
per_expert_bytes = 3 * (5_505_024 + 344_064)
|
||||||
|
all_routed_expert_bytes = per_expert_bytes * 896 * 92
|
||||||
|
kda_layers = text["linear_attn_config"]["kda_layers"]
|
||||||
|
mla_layers = text["linear_attn_config"]["full_attn_layers"]
|
||||||
|
|
||||||
|
flash_revision = subprocess.check_output(
|
||||||
|
["git", "-C", str(args.flashkda_dir), "rev-parse", "HEAD"],
|
||||||
|
text=True,
|
||||||
|
).strip()
|
||||||
|
|
||||||
|
captured_at = args.captured_at or datetime.now(timezone.utc).isoformat()
|
||||||
|
result = {
|
||||||
|
"schema_version": 1,
|
||||||
|
"captured_at": captured_at,
|
||||||
|
"evidence_boundary": {
|
||||||
|
"checkpoint_forward_run": False,
|
||||||
|
"raw_weights_committed": False,
|
||||||
|
"router_inputs": "deterministic synthetic RMS-normalized vectors, not token hidden states",
|
||||||
|
"kda_retention_probe": "noncanonical channel-wise interpretation used only to expose the A_log shape ambiguity",
|
||||||
|
},
|
||||||
|
"provenance": {
|
||||||
|
"huggingface_model": "moonshotai/Kimi-K3",
|
||||||
|
"huggingface_revision": hf_model["sha"],
|
||||||
|
"flashkda_revision": flash_revision,
|
||||||
|
"sha256": {
|
||||||
|
"config": sha256(args.config),
|
||||||
|
"index": sha256(args.index),
|
||||||
|
"kda_slice": sha256(args.kda_slice),
|
||||||
|
"router_prefix": sha256(args.router_prefix),
|
||||||
|
},
|
||||||
|
},
|
||||||
|
"checkpoint": {
|
||||||
|
"tensor_data_bytes": index["metadata"]["total_size"],
|
||||||
|
"tensor_data_tb": index["metadata"]["total_size"] / 1e12,
|
||||||
|
"tensor_data_tib": index["metadata"]["total_size"] / 2**40,
|
||||||
|
"shards": len(shard_files),
|
||||||
|
"shard_file_bytes": {
|
||||||
|
"sum": sum(shard_sizes),
|
||||||
|
"min": min(shard_sizes),
|
||||||
|
"max": max(shard_sizes),
|
||||||
|
"mean": sum(shard_sizes) / len(shard_sizes),
|
||||||
|
},
|
||||||
|
"tensor_entries": len(names),
|
||||||
|
"tensor_counts": {
|
||||||
|
"expert_packed": sum(
|
||||||
|
bool(re.search(r"experts\.\d+\.w[123]\.weight_packed$", name))
|
||||||
|
for name in names
|
||||||
|
),
|
||||||
|
"expert_scales": sum(
|
||||||
|
bool(re.search(r"experts\.\d+\.w[123]\.weight_scale$", name))
|
||||||
|
for name in names
|
||||||
|
),
|
||||||
|
"router_weight": sum(
|
||||||
|
name.endswith("block_sparse_moe.gate.weight") for name in names
|
||||||
|
),
|
||||||
|
"router_correction_bias": sum(
|
||||||
|
name.endswith("gate.e_score_correction_bias") for name in names
|
||||||
|
),
|
||||||
|
"attnres_proj": sum(
|
||||||
|
bool(re.search(r"(_res_proj|output_attn_res_proj)\.weight$", name))
|
||||||
|
for name in names
|
||||||
|
),
|
||||||
|
"attnres_norm": sum(
|
||||||
|
bool(re.search(r"(_res_norm|output_attn_res_norm)\.weight$", name))
|
||||||
|
for name in names
|
||||||
|
),
|
||||||
|
"kda_a_log": sum(name.endswith("self_attn.A_log") for name in names),
|
||||||
|
"kda_dt_bias": sum(name.endswith("self_attn.dt_bias") for name in names),
|
||||||
|
"vision": sum(name.startswith("vision_tower.") for name in names),
|
||||||
|
"projector": sum(name.startswith("mm_projector.") for name in names),
|
||||||
|
},
|
||||||
|
"derived_routed_expert_bytes": all_routed_expert_bytes,
|
||||||
|
"derived_routed_expert_share": all_routed_expert_bytes
|
||||||
|
/ index["metadata"]["total_size"],
|
||||||
|
},
|
||||||
|
"configuration": {
|
||||||
|
"layers": text["num_hidden_layers"],
|
||||||
|
"dense_layers": text["first_k_dense_replace"],
|
||||||
|
"hidden": text["hidden_size"],
|
||||||
|
"vocabulary": text["vocab_size"],
|
||||||
|
"context": text["max_position_embeddings"],
|
||||||
|
"kda_layers": kda_layers,
|
||||||
|
"mla_layers": mla_layers,
|
||||||
|
"heads": text["num_attention_heads"],
|
||||||
|
"head_dim": text["linear_attn_config"]["head_dim"],
|
||||||
|
"attnres_block": text["attn_res_block_size"],
|
||||||
|
"experts": text["num_experts"],
|
||||||
|
"active_experts": text["num_experts_per_token"],
|
||||||
|
"shared_experts": text["num_shared_experts"],
|
||||||
|
"latent_width": text["routed_expert_hidden_size"],
|
||||||
|
"expert_intermediate": text["moe_intermediate_size"],
|
||||||
|
"situ_beta": text["activation_situ_beta"],
|
||||||
|
"situ_linear_beta": text["activation_situ_linear_beta"],
|
||||||
|
"mla_nope": text["mla_use_nope"],
|
||||||
|
"mla_output_gate": text["mla_use_output_gate"],
|
||||||
|
},
|
||||||
|
"tensor_examples": {
|
||||||
|
"mla_layer_4": [
|
||||||
|
tensor_fact(
|
||||||
|
mla_header,
|
||||||
|
"language_model.model.layers.3.self_attn.kv_a_proj_with_mqa.weight",
|
||||||
|
),
|
||||||
|
tensor_fact(
|
||||||
|
mla_header,
|
||||||
|
"language_model.model.layers.3.self_attn.kv_b_proj.weight",
|
||||||
|
),
|
||||||
|
tensor_fact(
|
||||||
|
mla_header,
|
||||||
|
"language_model.model.layers.3.self_attn.q_a_proj.weight",
|
||||||
|
),
|
||||||
|
tensor_fact(
|
||||||
|
mla_header,
|
||||||
|
"language_model.model.layers.3.self_attn.q_b_proj.weight",
|
||||||
|
),
|
||||||
|
tensor_fact(
|
||||||
|
mla_header,
|
||||||
|
"language_model.model.layers.3.self_attn.g_proj.weight",
|
||||||
|
),
|
||||||
|
],
|
||||||
|
"vision": [
|
||||||
|
tensor_fact(vision_header, "vision_tower.patch_embed.proj.weight"),
|
||||||
|
tensor_fact(vision_header, "vision_tower.patch_embed.pos_emb.weight"),
|
||||||
|
tensor_fact(vision_header, "vision_tower.encoder.blocks.0.wqkv.weight"),
|
||||||
|
tensor_fact(vision_header, "vision_tower.encoder.blocks.26.wqkv.weight"),
|
||||||
|
tensor_fact(vision_header, "vision_tower.encoder.final_layernorm.weight"),
|
||||||
|
],
|
||||||
|
"routed_expert_0": [
|
||||||
|
{"name": "w1.weight_packed", "dtype": "U8", "shape": [3072, 1792], "bytes": 5_505_024},
|
||||||
|
{"name": "w1.weight_scale", "dtype": "U8", "shape": [3072, 112], "bytes": 344_064},
|
||||||
|
{"name": "w2.weight_packed", "dtype": "U8", "shape": [3584, 1536], "bytes": 5_505_024},
|
||||||
|
{"name": "w2.weight_scale", "dtype": "U8", "shape": [3584, 96], "bytes": 344_064},
|
||||||
|
{"name": "w3.weight_packed", "dtype": "U8", "shape": [3072, 1792], "bytes": 5_505_024},
|
||||||
|
{"name": "w3.weight_scale", "dtype": "U8", "shape": [3072, 112], "bytes": 344_064},
|
||||||
|
],
|
||||||
|
},
|
||||||
|
"parameter_audit": {
|
||||||
|
"a_log_checkpoint_shape": [128],
|
||||||
|
"a_log_public_code_shape": [96],
|
||||||
|
"dt_bias_shape": [96, 128],
|
||||||
|
"beta_projection_shape": [96, 7168],
|
||||||
|
"status": "observed shape inconsistency; runtime meaning unresolved",
|
||||||
|
"a_log": summarize(a_log),
|
||||||
|
"a_rate_exp": summarize(torch.exp(a_log)),
|
||||||
|
"dt_bias": summarize(dt_bias),
|
||||||
|
"channelwise_hypothesis": {
|
||||||
|
"log_decay": summarize(channelwise_log_decay),
|
||||||
|
"one_step_retention": summarize(channelwise_retention),
|
||||||
|
"retention_after_64_steps": summarize(channelwise_retention.pow(64)),
|
||||||
|
},
|
||||||
|
"router_correction_bias": summarize(correction_bias),
|
||||||
|
"router_row_l2": summarize(router_norms),
|
||||||
|
"router_bias_norm_correlation": torch.corrcoef(
|
||||||
|
torch.stack([correction_bias, router_norms])
|
||||||
|
)[0, 1].item(),
|
||||||
|
},
|
||||||
|
"router_stress_probe": {
|
||||||
|
"seed": args.seed,
|
||||||
|
"synthetic_tokens": args.synthetic_tokens,
|
||||||
|
"hidden_rms": synthetic.square().mean().sqrt().item(),
|
||||||
|
"without_correction_bias": load_metrics(unbiased_load),
|
||||||
|
"with_correction_bias": load_metrics(biased_load),
|
||||||
|
"membership_overlap_mean": overlap.mean().item(),
|
||||||
|
"tokens_changed": int((overlap < 16).sum()),
|
||||||
|
"changed_fraction": (overlap < 16).float().mean().item(),
|
||||||
|
"mean_replacements_per_token": (16 - overlap).mean().item(),
|
||||||
|
},
|
||||||
|
"flashkda": {
|
||||||
|
"supported_architectures": ["90a", "100a", "103a", "120a"],
|
||||||
|
"requirements": {"cuda": ">=12.9", "pytorch": ">=2.4", "gpu": "SM90+"},
|
||||||
|
"official_benchmarks": {
|
||||||
|
"h20": parse_benchmark(args.flashkda_dir / "BENCHMARK_H20.md"),
|
||||||
|
"gb200": parse_benchmark(args.flashkda_dir / "BENCHMARK_GB200.md"),
|
||||||
|
},
|
||||||
|
"local_environment": {
|
||||||
|
"python": platform.python_version(),
|
||||||
|
"torch": torch.__version__,
|
||||||
|
"torch_cuda": torch.version.cuda,
|
||||||
|
"gpu": torch.cuda.get_device_name(0) if torch.cuda.is_available() else None,
|
||||||
|
"capability": list(torch.cuda.get_device_capability(0))
|
||||||
|
if torch.cuda.is_available()
|
||||||
|
else None,
|
||||||
|
"libc": list(platform.libc_ver()),
|
||||||
|
},
|
||||||
|
"local_build": {
|
||||||
|
"status": "blocked_before_kernel execution",
|
||||||
|
"attempt_1": "system g++ 15 exceeds CUDA 12.8 host compiler range",
|
||||||
|
"attempt_2": "temporary g++ 13 reaches nvcc, then CUDA 12.8 headers conflict with current glibc math declarations",
|
||||||
|
"interpretation": "GPU architecture is listed by the repository, but the local CUDA 12.8 stack is below the official CUDA 12.9 requirement",
|
||||||
|
},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
args.output.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
args.output.write_text(json.dumps(result, ensure_ascii=False, indent=2) + "\n")
|
||||||
|
print(json.dumps({"output": str(args.output), "bytes": args.output.stat().st_size}))
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,367 @@
|
|||||||
|
# Kimi K3 第三轮:开放工件、权重元数据与可复现实验审计
|
||||||
|
|
||||||
|
> 研究截点:2026-07-29
|
||||||
|
> 官方模型修订:`moonshotai/Kimi-K3@9f62e4e9fffbd0a83ddd60e1c209d828994b3569`
|
||||||
|
> 官方 FlashKDA 修订:`MoonshotAI/FlashKDA@1ce47ea3bb22c84eb9cc665028399cf35e8ffb0b`
|
||||||
|
> 原则:模型卡、配置、远程代码、safetensors header、kernel API、作者 benchmark 和本机实验分开记账。
|
||||||
|
|
||||||
|
## 1. 本轮究竟要推进什么
|
||||||
|
|
||||||
|
第二轮已经把 47 页报告拆成问题、公式、图表和教学实验。第三轮不能只是再写一遍报告,而要回答:
|
||||||
|
|
||||||
|
1. 开放权重仓库实际包含哪些文件、分片和张量?
|
||||||
|
2. 报告的 `69 KDA + 24 MLA`、`1 dense + 92 MoE` 是否能在配置和 tensor names 中闭合?
|
||||||
|
3. Stable LatentMoE 与 MXFP4 在真实 checkpoint 中怎样落成 packed tensors 与 scales?
|
||||||
|
4. `NoPE` 在公开实现中究竟是“删掉位置相关通道”,还是“不施加 rotary transform”?
|
||||||
|
5. AttnRes 在 checkpoint 中留下多少 projection / norm 参数?
|
||||||
|
6. FlashKDA 的公开接口、支持架构、作者 benchmark 和本机可运行边界分别是什么?
|
||||||
|
7. 哪些可称为真实工件 trace,哪些仍只是确定性推导或 synthetic probe?
|
||||||
|
|
||||||
|
本轮不声称完成以下事情:
|
||||||
|
|
||||||
|
- 没有在单张 RTX 5090 上加载 1.56 TB checkpoint;
|
||||||
|
- 没有获得真实 token hidden states、线上 expert load 或生产 cache trace;
|
||||||
|
- 没有把随机向量上的 router 行为写成真实数据分布;
|
||||||
|
- 没有把未成功执行的 FlashKDA kernel 写成“本机 benchmark”;
|
||||||
|
- 没有替官方解释下面发现的 `A_log` 形状不一致。
|
||||||
|
|
||||||
|
## 2. 一手工件与校验
|
||||||
|
|
||||||
|
| 工件 | canonical source | 本地校验 |
|
||||||
|
|---|---|---|
|
||||||
|
| 模型配置 | `https://huggingface.co/moonshotai/Kimi-K3/blob/main/config.json` | `9710e121…379213` |
|
||||||
|
| HF configuration code | `configuration_kimi_k3.py` | `735eb9eb…b416ae` |
|
||||||
|
| 多模态 modeling code | `modeling_kimi_k3.py` | `b9171c96…65ea2` |
|
||||||
|
| 文本 backbone code | `modeling_kimi_linear.py` | `9e3564c7…ff44a` |
|
||||||
|
| tensor index | `model.safetensors.index.json` | `a1c52106…7febd` |
|
||||||
|
| layer-1 KDA 小范围字节 | shard 1 HTTP Range | `dccb10e7…d9b2` |
|
||||||
|
| layer-2 router prefix | shard 2 HTTP Range | `dbe66ff8…23ee` |
|
||||||
|
| FlashKDA | `https://github.com/MoonshotAI/FlashKDA` | git `1ce47ea3…f0b` |
|
||||||
|
|
||||||
|
完整校验值保存在 `src/data/k3-artifact-snapshot.json`。原始权重字节不进入公开仓库。
|
||||||
|
|
||||||
|
### 2.1 为什么只读 Range
|
||||||
|
|
||||||
|
模型 tensor data 为 `1,560,860,324,864` bytes,即:
|
||||||
|
|
||||||
|
- `1.56086 TB`(十进制);
|
||||||
|
- `1.41959 TiB`(二进制);
|
||||||
|
- 96 个 safetensors shards;
|
||||||
|
- 497,220 个 tensor entries。
|
||||||
|
|
||||||
|
当前机器是 32 GB RTX 5090、123 GiB RAM。完整下载和加载既不必要,也不能支持“单卡复现”。
|
||||||
|
safetensors 把 JSON header 放在每个分片开头;header 给出 tensor name、dtype、shape 与 byte offsets。
|
||||||
|
因此可以只读:
|
||||||
|
|
||||||
|
```text
|
||||||
|
8-byte header length
|
||||||
|
→ JSON header
|
||||||
|
→ selected tensor byte ranges
|
||||||
|
```
|
||||||
|
|
||||||
|
这足以审计结构,也能读取少量参数做统计,同时避免把 1.56 TB 下载行为伪装成研究必要条件。
|
||||||
|
|
||||||
|
## 3. checkpoint 拓扑:配置、索引与 header 闭合
|
||||||
|
|
||||||
|
### 3.1 层型
|
||||||
|
|
||||||
|
配置精确给出:
|
||||||
|
|
||||||
|
- 93 层;
|
||||||
|
- KDA layers:69;
|
||||||
|
- full-attention / Gated MLA layers:24;
|
||||||
|
- MLA 位于 1-based `4, 8, …, 92, 93`;
|
||||||
|
- 第一层 dense,后 92 层为 MoE;
|
||||||
|
- AttnRes block size 为 12。
|
||||||
|
|
||||||
|
因此真实层条带是:
|
||||||
|
|
||||||
|
```text
|
||||||
|
L01 KDA + dense FFN
|
||||||
|
L02 KDA + MoE
|
||||||
|
L03 KDA + MoE
|
||||||
|
L04 MLA + MoE
|
||||||
|
...
|
||||||
|
L92 MLA + MoE
|
||||||
|
L93 MLA + MoE
|
||||||
|
```
|
||||||
|
|
||||||
|
最后连续两个 MLA 不是排版错误:`L92` 是最后一个 3:1 block 的 MLA,`L93` 是报告所述额外末层 MLA。
|
||||||
|
|
||||||
|
### 3.2 tensor counts
|
||||||
|
|
||||||
|
| 对象 | index 实际计数 | 为什么是这个数 |
|
||||||
|
|---|---:|---|
|
||||||
|
| KDA `A_log` | 69 | 每个 KDA layer 一份 |
|
||||||
|
| KDA `dt_bias` | 69 | 每个 KDA layer 一份 |
|
||||||
|
| Router weight | 92 | 每个 MoE layer 一份 |
|
||||||
|
| Router correction bias | 92 | 每个 MoE layer 一份 |
|
||||||
|
| Routed expert packed weights | 247,296 | `92 × 896 × 3` |
|
||||||
|
| Routed expert scales | 247,296 | `92 × 896 × 3` |
|
||||||
|
| AttnRes projections | 187 | `93 × 2 + output 1` |
|
||||||
|
| AttnRes norms | 187 | `93 × 2 + output 1` |
|
||||||
|
| Vision tensors | 165 | 27-layer MoonViT-V2 与输入/输出项 |
|
||||||
|
| Projector tensors | 3 | 视觉塔到 7168 text hidden |
|
||||||
|
|
||||||
|
`497,220` 不是“参数数量”,而是 safetensors 中的命名 tensor entry 数。不能和 `2.78T parameters`
|
||||||
|
混为一个口径。
|
||||||
|
|
||||||
|
## 4. 真实 tensor shape:架构不再只靠报告表格
|
||||||
|
|
||||||
|
### 4.1 Dense layer 1 的 KDA
|
||||||
|
|
||||||
|
shard 1 header 给出:
|
||||||
|
|
||||||
|
| tensor | dtype | shape |
|
||||||
|
|---|---|---|
|
||||||
|
| `q/k/v/g_proj.weight` | BF16 | `12288 × 7168` |
|
||||||
|
| `o_proj.weight` | BF16 | `7168 × 12288` |
|
||||||
|
| `q/k/v_conv1d.weight` | F32 | `12288 × 1 × 4` |
|
||||||
|
| `b_proj.weight` | BF16 | `96 × 7168` |
|
||||||
|
| `f_a_proj.weight` | BF16 | `128 × 7168` |
|
||||||
|
| `f_b_proj.weight` | BF16 | `12288 × 128` |
|
||||||
|
| dense FFN gate/up | BF16 | `33792 × 7168` |
|
||||||
|
| dense FFN down | BF16 | `7168 × 33792` |
|
||||||
|
|
||||||
|
`12288 = 96 heads × 128 dims`;ShortConv kernel size 4 也直接出现在 checkpoint shape 中。
|
||||||
|
|
||||||
|
### 4.2 MLA layer 4
|
||||||
|
|
||||||
|
shard 4 header 给出:
|
||||||
|
|
||||||
|
| tensor | shape |
|
||||||
|
|---|---|
|
||||||
|
| `q_a_proj` | `1536 × 7168` |
|
||||||
|
| `q_b_proj` | `18432 × 1536` |
|
||||||
|
| `kv_a_proj_with_mqa` | `576 × 7168` |
|
||||||
|
| `kv_b_proj` | `24576 × 512` |
|
||||||
|
| `g_proj` | `12288 × 7168` |
|
||||||
|
| `o_proj` | `7168 × 12288` |
|
||||||
|
|
||||||
|
关键解释:
|
||||||
|
|
||||||
|
- `576 = 512 KV latent + 64 auxiliary q/k channel`;
|
||||||
|
- `q_b` 输出 `96 × (128 + 64) = 18432`;
|
||||||
|
- 公开 forward 没有对这 64 维施加 RoPE,但仍投影、拼接并进入 attention;
|
||||||
|
- 因此 K3 的 `mla_use_nope=true` 应解释为“不施加显式 rotary transform”,不能改写成“MLA 中完全不存在任何额外 q/k 通道”。
|
||||||
|
|
||||||
|
### 4.3 一个 routed expert 的 MXFP4 表示
|
||||||
|
|
||||||
|
layer 2 / expert 0 的 header:
|
||||||
|
|
||||||
|
| tensor | dtype | packed shape | bytes |
|
||||||
|
|---|---|---:|---:|
|
||||||
|
| `w1.weight_packed` | U8 | `3072 × 1792` | 5,505,024 |
|
||||||
|
| `w1.weight_scale` | U8 | `3072 × 112` | 344,064 |
|
||||||
|
| `w2.weight_packed` | U8 | `3584 × 1536` | 5,505,024 |
|
||||||
|
| `w2.weight_scale` | U8 | `3584 × 96` | 344,064 |
|
||||||
|
| `w3.weight_packed` | U8 | `3072 × 1792` | 5,505,024 |
|
||||||
|
| `w3.weight_scale` | U8 | `3072 × 112` | 344,064 |
|
||||||
|
|
||||||
|
这里能直接读出两个结构:
|
||||||
|
|
||||||
|
1. routed expert 工作在 `3584` latent width,而不是 `7168` full hidden;
|
||||||
|
2. scale group size 为 32:`3584 / 32 = 112`,`3072 / 32 = 96`。
|
||||||
|
|
||||||
|
每个 routed expert 的 packed weights + scales 为 `17,547,264` bytes。按配置与 index 的统一形状推导,
|
||||||
|
92 层全部 routed experts 约占 tensor data 的 `92.67%`。这是 header + 配置的确定性推导,不是运行显存。
|
||||||
|
|
||||||
|
### 4.4 MoonViT-V2
|
||||||
|
|
||||||
|
shard 96 header:
|
||||||
|
|
||||||
|
| tensor | shape |
|
||||||
|
|---|---|
|
||||||
|
| patch projection | `1024 × 3 × 14 × 14` |
|
||||||
|
| learned 2D position table | `64 × 64 × 1024` |
|
||||||
|
| block 0 / 26 QKV | `4608 × 1024` |
|
||||||
|
| block 0 / 26 MLP up | `4096 × 1024` |
|
||||||
|
| block 0 / 26 MLP down | `1024 × 4096` |
|
||||||
|
| final norm | `1024` |
|
||||||
|
|
||||||
|
配置、首层和末层 tensor 同时支持“27 layers、patch 14、vision hidden 1024”;这比只引用模型卡更强。
|
||||||
|
|
||||||
|
## 5. 小参数 Range audit
|
||||||
|
|
||||||
|
### 5.1 KDA layer 1
|
||||||
|
|
||||||
|
只读 shard 1 开头的 49,664 bytes:
|
||||||
|
|
||||||
|
- checkpoint `A_log`: F32 `[128]`;
|
||||||
|
- `dt_bias`: F32 `[12288]`,可按公开配置写成 `[96, 128]`。
|
||||||
|
|
||||||
|
真实参数统计:
|
||||||
|
|
||||||
|
| 参数 | median | p10–p90 | min–max |
|
||||||
|
|---|---:|---:|---:|
|
||||||
|
| `A_log` | −0.1533 | −0.5367–0.0743 | −0.7531–2.4661 |
|
||||||
|
| `exp(A_log)` | 0.8579 | 0.5847–1.0771 | 0.4709–11.7764 |
|
||||||
|
| `dt_bias` | −4.6220 | −6.4912–−2.5690 | −7.8938–0.1792 |
|
||||||
|
|
||||||
|
### 5.2 一个必须公开保留的形状不一致
|
||||||
|
|
||||||
|
四份官方工件目前给出:
|
||||||
|
|
||||||
|
| 工件 | 观测 |
|
||||||
|
|---|---|
|
||||||
|
| `config.json` | `num_heads=96`, `head_dim=128` |
|
||||||
|
| HF remote code | `A_log` 初始化 shape 为 `[num_heads]`,即 `[96]` |
|
||||||
|
| FlashKDA API / C++ check | `A_log` 必须 `[H]` |
|
||||||
|
| checkpoint shard header | `A_log` 实际为 `[128]` |
|
||||||
|
|
||||||
|
同时:
|
||||||
|
|
||||||
|
- `dt_bias` 可闭合为 `[96,128]`;
|
||||||
|
- `b_proj.weight` 为 `[96,7168]`,β 显然按 96 heads;
|
||||||
|
- q/k/v projection 为 `12288 = 96×128`。
|
||||||
|
|
||||||
|
因此,**公开工件中存在可复现的 `A_log [128]` vs expected `[96]` 形状不一致**。
|
||||||
|
|
||||||
|
当前可下的结论只有:
|
||||||
|
|
||||||
|
- 这是 header / code / config / kernel API 四方直接观测,不是 Grok 猜测;
|
||||||
|
- 公开 remote code 按字面构造时会期待 `[96]`;
|
||||||
|
- 需要 Moonshot、实际 vLLM/SGLang loader 或后续权重修订解释转换规则。
|
||||||
|
|
||||||
|
当前不能下的结论:
|
||||||
|
|
||||||
|
- 不能直接宣布 checkpoint 损坏;
|
||||||
|
- 不能擅自把 `[128]` 解释为 per-channel `A_log`;
|
||||||
|
- 不能用某个猜测 reshape 得到的曲线冒充模型真实 retention。
|
||||||
|
|
||||||
|
数据快照保留了一个明确标成 `noncanonical channel-wise hypothesis` 的数值探针,只用于说明:
|
||||||
|
若把 `[128]` 当 channel 参数,能得到怎样的 retention 分布;它不进入正式 forward 结论。
|
||||||
|
|
||||||
|
## 6. Router:真实权重与 synthetic input 必须分层
|
||||||
|
|
||||||
|
读取 layer 2 的:
|
||||||
|
|
||||||
|
- correction bias:896 个 F32;
|
||||||
|
- router weight:`896 × 7168` BF16,约 12.85 MB。
|
||||||
|
|
||||||
|
真实参数统计:
|
||||||
|
|
||||||
|
- correction bias median `0.00493`,min `−0.08442`,max `0.02861`;
|
||||||
|
- router row L2 median `5.0364`,min `2.6019`,max `7.0232`;
|
||||||
|
- bias 与 row norm 的 Pearson correlation `0.4567`。
|
||||||
|
|
||||||
|
为了测试“只拿真实 router weights 是否就能评价 Quantile Balancing”,脚本生成 2,048 个固定 seed、
|
||||||
|
RMS=1 的各向同性随机 hidden vectors,再比较 top-16:
|
||||||
|
|
||||||
|
| 条件 | load CV | Gini | zero-load experts |
|
||||||
|
|---|---:|---:|---:|
|
||||||
|
| 不加 correction bias | 2.085 | 0.831 | 558 |
|
||||||
|
| 加 checkpoint correction bias | 2.529 | 0.879 | 673 |
|
||||||
|
|
||||||
|
两组 top-16 平均只重合 `2.17 / 16`。
|
||||||
|
|
||||||
|
这不是在证明 QB 让真实负载更差,反而证明:
|
||||||
|
|
||||||
|
1. router weights 与 hidden-state distribution 是共同训练的;
|
||||||
|
2. sigmoid top scores 在随机 RMS=1 输入下容易饱和;
|
||||||
|
3. 小 correction bias 会在饱和的近并列区域强烈改写名次;
|
||||||
|
4. 没有真实 hidden traces,就不能用随机向量评价真实负载均衡。
|
||||||
|
|
||||||
|
因此网站把这组结果命名为 **counterexample / synthetic stress probe**,不是“真实 expert load trace”。
|
||||||
|
|
||||||
|
## 7. FlashKDA:作者 benchmark 与本机实验分开
|
||||||
|
|
||||||
|
### 7.1 官方仓库事实
|
||||||
|
|
||||||
|
FlashKDA `1ce47ea3`:
|
||||||
|
|
||||||
|
- CUTLASS kernels;
|
||||||
|
- 支持 `90a / 100a / 103a / 120a`;
|
||||||
|
- README 要求 SM90+、CUDA 12.9+、PyTorch 2.4+;
|
||||||
|
- kernel API 固定 `K=V=128`;
|
||||||
|
- q/k/v/g 为 BF16,`A_log` / `dt_bias` 为 F32;
|
||||||
|
- 支持 fixed length、variable length、initial / final recurrent state。
|
||||||
|
|
||||||
|
官方报告的 `T=8192, H=96, D=128`:
|
||||||
|
|
||||||
|
| device | case | FlashKDA | FLA chunk KDA | 作者报告 speedup |
|
||||||
|
|---|---|---:|---:|---:|
|
||||||
|
| H20 | fixed | 2.6220 ms | 4.8388 ms | 1.85× |
|
||||||
|
| H20 | 8×1024 varlen | 2.0432 ms | 4.6723 ms | 2.29× |
|
||||||
|
| GB200 | fixed | 1.0087 ms | 2.3271 ms | 2.31× |
|
||||||
|
| GB200 | 8×1024 varlen | 0.7064 ms | 2.3105 ms | 3.27× |
|
||||||
|
|
||||||
|
这些是作者仓库 benchmark,不是本站复跑值。
|
||||||
|
|
||||||
|
### 7.2 本机真实构建边界
|
||||||
|
|
||||||
|
本机:
|
||||||
|
|
||||||
|
- RTX 5090,compute capability `12.0`;
|
||||||
|
- PyTorch `2.11.0+cu128`;
|
||||||
|
- PyTorch CUDA `12.8`;
|
||||||
|
- 官方源码明确包含 `sm_120a`,所以不是 GPU architecture 缺失。
|
||||||
|
|
||||||
|
两次可复现构建:
|
||||||
|
|
||||||
|
1. 系统 `g++ 15.2`:PyTorch extension 在编译前拒绝,CUDA 12.8 要求 host compiler `<14`;
|
||||||
|
2. 临时解包 `g++ 13.4`:成功进入 nvcc,但 CUDA 12.8 headers 与当前 glibc math declarations
|
||||||
|
在 `cospi / sinpi / rsqrt` exception specification 处冲突。
|
||||||
|
|
||||||
|
结论:
|
||||||
|
|
||||||
|
- kernel 尚未在本站机器执行;
|
||||||
|
- 失败与 README 的 CUDA 12.9+ 要求一致;
|
||||||
|
- 不能把 `sm_120a` 支持写成本机已经跑通;
|
||||||
|
- 下一次应使用匹配 PyTorch 的 CUDA 12.9+ toolchain 或官方容器后再复跑 correctness + benchmark。
|
||||||
|
|
||||||
|
## 8. 可复现实验入口
|
||||||
|
|
||||||
|
脚本:
|
||||||
|
|
||||||
|
```text
|
||||||
|
experiments/k3/checkpoint_probe.py
|
||||||
|
```
|
||||||
|
|
||||||
|
提交的数据快照:
|
||||||
|
|
||||||
|
```text
|
||||||
|
src/data/k3-artifact-snapshot.json
|
||||||
|
```
|
||||||
|
|
||||||
|
脚本会:
|
||||||
|
|
||||||
|
1. 解析 config、index 与 selected headers;
|
||||||
|
2. 校验小范围字节长度和 SHA-256;
|
||||||
|
3. 统计真实 KDA / router 参数;
|
||||||
|
4. 运行明确标注的 synthetic router counterexample;
|
||||||
|
5. 解析 FlashKDA 官方 H20 / GB200 benchmark;
|
||||||
|
6. 输出本机环境与构建边界;
|
||||||
|
7. 不提交原始权重。
|
||||||
|
|
||||||
|
## 9. 网站实现合同
|
||||||
|
|
||||||
|
第三轮开放工件实验室必须有四个视图:
|
||||||
|
|
||||||
|
1. **Layer map**:93 层真实 config 条带;显示 KDA/MLA、dense/MoE、AttnRes block。
|
||||||
|
2. **Tensor anatomy**:checkpoint / expert / vision / MLA tensor shape 与数量。
|
||||||
|
3. **Parameter audit**:真实 Range statistics,并把 `A_log` mismatch 放在主视区。
|
||||||
|
4. **Reproduction boundary**:官方 benchmark、本站构建失败点、synthetic router counterexample。
|
||||||
|
|
||||||
|
每个视图必须显示证据类型:
|
||||||
|
|
||||||
|
- `O` = official artifact observation;
|
||||||
|
- `D` = deterministic derivation;
|
||||||
|
- `X` = executed local experiment;
|
||||||
|
- `S` = synthetic stress probe;
|
||||||
|
- `U` = unresolved inconsistency。
|
||||||
|
|
||||||
|
## 10. 下一轮证据闸门
|
||||||
|
|
||||||
|
- [x] 官方 config / code / index revision 固定;
|
||||||
|
- [x] 96 个分片与 497,220 tensor entries 审计;
|
||||||
|
- [x] KDA / MLA / MoE / AttnRes / Vision tensor shapes 入账;
|
||||||
|
- [x] selected open-weight ranges 做真实参数统计;
|
||||||
|
- [x] 发现并限定 `A_log` shape inconsistency;
|
||||||
|
- [x] FlashKDA RTX 5090 构建尝试留下可复现边界;
|
||||||
|
- [ ] 使用 CUDA 12.9+ 匹配环境跑 FlashKDA exact correctness;
|
||||||
|
- [ ] 取得真实 hidden-state / router load trace;
|
||||||
|
- [ ] 取得可加载的 reduced checkpoint、官方 trace 或多机资源;
|
||||||
|
- [ ] 对 Figure 3 / 4 / 5 做真实数值重绘;
|
||||||
|
- [ ] 对 AttnRes 读取分布做真实 token / layer trace。
|
||||||
|
|
||||||
@@ -75,6 +75,10 @@ const overview = await evaluate(`(() => ({
|
|||||||
paperGroups: document.querySelectorAll("#papers .paper-group").length,
|
paperGroups: document.querySelectorAll("#papers .paper-group").length,
|
||||||
labTabs: document.querySelectorAll("[data-k3-tab]").length,
|
labTabs: document.querySelectorAll("[data-k3-tab]").length,
|
||||||
labPanels: document.querySelectorAll("[data-k3-panel]").length,
|
labPanels: document.querySelectorAll("[data-k3-panel]").length,
|
||||||
|
artifactTabs: document.querySelectorAll("[data-artifact-tab]").length,
|
||||||
|
artifactPanels: document.querySelectorAll("[data-artifact-panel]").length,
|
||||||
|
artifactLayers: document.querySelectorAll("[data-layer-cell]").length,
|
||||||
|
artifactMismatch: document.querySelector("#artifacts")?.textContent.includes("A_log [128] ≠ expected [96]"),
|
||||||
nativeVisionCorrected: document.body.textContent.includes("MoonViT‑V2 从头训练") &&
|
nativeVisionCorrected: document.body.textContent.includes("MoonViT‑V2 从头训练") &&
|
||||||
document.body.textContent.includes("同一个 next-token prediction objective"),
|
document.body.textContent.includes("同一个 next-token prediction objective"),
|
||||||
staleVisionClaim: document.body.textContent.includes("先固定语言模型训练视觉组件"),
|
staleVisionClaim: document.body.textContent.includes("先固定语言模型训练视觉组件"),
|
||||||
@@ -159,12 +163,100 @@ const labs = await evaluate(`(() => {
|
|||||||
};
|
};
|
||||||
})()`);
|
})()`);
|
||||||
|
|
||||||
|
const artifacts = await evaluate(`(() => {
|
||||||
|
const root = document.querySelector("[data-k3-artifact-lab]");
|
||||||
|
const panel = () => root.querySelector("[data-artifact-panel]:not([hidden])").dataset.artifactPanel;
|
||||||
|
const text = (selector) => root.querySelector(selector).textContent.trim();
|
||||||
|
const input = (selector, value) => {
|
||||||
|
const node = root.querySelector(selector);
|
||||||
|
node.value = value;
|
||||||
|
node.dispatchEvent(new Event("input", { bubbles: true }));
|
||||||
|
};
|
||||||
|
|
||||||
|
const initial = {
|
||||||
|
panel: panel(),
|
||||||
|
layer: text("[data-layer-number]"),
|
||||||
|
attention: text("[data-layer-attention]"),
|
||||||
|
ffn: text("[data-layer-ffn]"),
|
||||||
|
block: text("[data-layer-block]"),
|
||||||
|
};
|
||||||
|
input("[data-layer-slider]", "93");
|
||||||
|
const terminal = {
|
||||||
|
layer: text("[data-layer-number]"),
|
||||||
|
attention: text("[data-layer-attention]"),
|
||||||
|
ffn: text("[data-layer-ffn]"),
|
||||||
|
copy: text("[data-layer-special]"),
|
||||||
|
};
|
||||||
|
|
||||||
|
root.querySelector('[data-artifact-tab="tensors"]').click();
|
||||||
|
const tensors = {
|
||||||
|
panel: panel(),
|
||||||
|
groups: root.querySelectorAll("[data-tensor-tab]").length,
|
||||||
|
expertRows: root.querySelector('[data-tensor-panel="expert"]').querySelectorAll(":scope > div").length,
|
||||||
|
entries: root.textContent.includes("497,220"),
|
||||||
|
share: root.textContent.includes("92.67%"),
|
||||||
|
};
|
||||||
|
root.querySelector('[data-tensor-tab="mla"]').click();
|
||||||
|
const mla = {
|
||||||
|
visible: !root.querySelector('[data-tensor-panel="mla"]').hidden,
|
||||||
|
rows: root.querySelector('[data-tensor-panel="mla"]').querySelectorAll(":scope > div").length,
|
||||||
|
has576: root.querySelector('[data-tensor-panel="mla"]').textContent.includes("576 × 7168"),
|
||||||
|
};
|
||||||
|
|
||||||
|
root.querySelector('[data-artifact-tab="parameters"]').click();
|
||||||
|
const parameterInitial = {
|
||||||
|
panel: panel(),
|
||||||
|
shape: text("[data-parameter-shape]"),
|
||||||
|
conflict: root.textContent.includes("A_log [128]") && root.textContent.includes("A_log [H] = [96]"),
|
||||||
|
};
|
||||||
|
input("[data-parameter-select]", "dt");
|
||||||
|
const parameterChanged = {
|
||||||
|
shape: text("[data-parameter-shape]"),
|
||||||
|
count: text("[data-parameter-count]"),
|
||||||
|
};
|
||||||
|
|
||||||
|
root.querySelector('[data-artifact-tab="reproduction"]').click();
|
||||||
|
const reproductionInitial = {
|
||||||
|
panel: panel(),
|
||||||
|
flash: text("[data-benchmark-flash]"),
|
||||||
|
fla: text("[data-benchmark-fla]"),
|
||||||
|
speedup: text("[data-benchmark-speedup]"),
|
||||||
|
cv: text("[data-router-cv]"),
|
||||||
|
zero: text("[data-router-zero]"),
|
||||||
|
};
|
||||||
|
input("[data-benchmark-device]", "gb200");
|
||||||
|
input("[data-benchmark-case]", "Varlen, \\\`seq_lens\\\`=\\\`1024 x 8\\\`");
|
||||||
|
input("[data-router-mode]", "bias");
|
||||||
|
const reproductionChanged = {
|
||||||
|
flash: text("[data-benchmark-flash]"),
|
||||||
|
speedup: text("[data-benchmark-speedup]"),
|
||||||
|
cv: text("[data-router-cv]"),
|
||||||
|
zero: text("[data-router-zero]"),
|
||||||
|
};
|
||||||
|
|
||||||
|
const first = root.querySelector('[data-artifact-tab="layers"]');
|
||||||
|
first.focus();
|
||||||
|
first.dispatchEvent(new KeyboardEvent("keydown", { key: "ArrowRight", bubbles: true }));
|
||||||
|
return {
|
||||||
|
initial, terminal, tensors, mla, parameterInitial, parameterChanged,
|
||||||
|
reproductionInitial, reproductionChanged,
|
||||||
|
keyboardSelected: root.querySelector('[data-artifact-tab][aria-selected="true"]').dataset.artifactTab,
|
||||||
|
keyboardVisible: panel(),
|
||||||
|
};
|
||||||
|
})()`);
|
||||||
|
|
||||||
await evaluate(`(() => {
|
await evaluate(`(() => {
|
||||||
document.querySelector("[data-k3-lab]").scrollIntoView({ block: "start", behavior: "instant" });
|
document.querySelector("[data-k3-lab]").scrollIntoView({ block: "start", behavior: "instant" });
|
||||||
window.scrollBy(0, -82);
|
window.scrollBy(0, -82);
|
||||||
})()`);
|
})()`);
|
||||||
await pause(180);
|
await pause(180);
|
||||||
await screenshot("/tmp/llm-atlas-k3-lab-desktop.png");
|
await screenshot("/tmp/llm-atlas-k3-lab-desktop.png");
|
||||||
|
await evaluate(`(() => {
|
||||||
|
document.querySelector("[data-k3-artifact-lab]").scrollIntoView({ block: "start", behavior: "instant" });
|
||||||
|
window.scrollBy(0, -82);
|
||||||
|
})()`);
|
||||||
|
await pause(180);
|
||||||
|
await screenshot("/tmp/llm-atlas-k3-artifact-desktop.png");
|
||||||
|
|
||||||
await command("Emulation.setDeviceMetricsOverride", {
|
await command("Emulation.setDeviceMetricsOverride", {
|
||||||
width: 390,
|
width: 390,
|
||||||
@@ -183,8 +275,10 @@ const mobile = await evaluate(`(() => {
|
|||||||
menuVisible: getComputedStyle(toggle).display !== "none",
|
menuVisible: getComputedStyle(toggle).display !== "none",
|
||||||
menuOpen: toggle.getAttribute("aria-expanded"),
|
menuOpen: toggle.getAttribute("aria-expanded"),
|
||||||
tabs: root.querySelectorAll("[data-k3-tab]").length,
|
tabs: root.querySelectorAll("[data-k3-tab]").length,
|
||||||
|
artifactTabs: document.querySelectorAll("[data-artifact-tab]").length,
|
||||||
|
artifactLayers: document.querySelectorAll("[data-layer-cell]").length,
|
||||||
offenders: [...document.querySelectorAll("body *")]
|
offenders: [...document.querySelectorAll("body *")]
|
||||||
.filter((node) => !node.closest(".paper-chain, .spec-table-wrap, .cache-strip, .architecture-explorer, [data-k3-lab]"))
|
.filter((node) => !node.closest(".paper-chain, .spec-table-wrap, .cache-strip, .architecture-explorer, [data-k3-lab], [data-k3-artifact-lab]"))
|
||||||
.filter((node) => node.getBoundingClientRect().right > document.documentElement.clientWidth + 1)
|
.filter((node) => node.getBoundingClientRect().right > document.documentElement.clientWidth + 1)
|
||||||
.slice(0, 15)
|
.slice(0, 15)
|
||||||
.map((node) => ({
|
.map((node) => ({
|
||||||
@@ -201,17 +295,24 @@ await evaluate(`(() => {
|
|||||||
})()`);
|
})()`);
|
||||||
await pause(180);
|
await pause(180);
|
||||||
await screenshot("/tmp/llm-atlas-k3-mobile.png");
|
await screenshot("/tmp/llm-atlas-k3-mobile.png");
|
||||||
|
await evaluate(`(() => {
|
||||||
|
document.querySelector("[data-k3-artifact-lab]").scrollIntoView({ block: "start", behavior: "instant" });
|
||||||
|
window.scrollBy(0, -64);
|
||||||
|
})()`);
|
||||||
|
await pause(180);
|
||||||
|
await screenshot("/tmp/llm-atlas-k3-artifact-mobile.png");
|
||||||
|
|
||||||
const report = { overview, labs, mobile, exceptions };
|
const report = { overview, labs, artifacts, mobile, exceptions };
|
||||||
console.log(JSON.stringify(report, null, 2));
|
console.log(JSON.stringify(report, null, 2));
|
||||||
|
|
||||||
const numeric = (text) => Number.parseFloat(text.replaceAll(",", "").replace("−", "-"));
|
const numeric = (text) => Number.parseFloat(text.replaceAll(",", "").replace("−", "-"));
|
||||||
const failures = [];
|
const failures = [];
|
||||||
if (!overview.title.includes("因果环节")) failures.push("K3 二轮标题异常");
|
if (!overview.title.includes("因果环节")) failures.push("K3 二轮标题异常");
|
||||||
if (overview.sections !== 31 || overview.tocLinks !== 31) failures.push("30 个编号专题加阅读链的目录结构异常");
|
if (overview.sections !== 32 || overview.tocLinks !== 32) failures.push("31 个编号专题加阅读链的目录结构异常");
|
||||||
if (overview.ledgers !== 32 || overview.reportMap !== 9) failures.push("32 张问题账或报告地图异常");
|
if (overview.ledgers !== 32 || overview.reportMap !== 9) failures.push("32 张问题账或报告地图异常");
|
||||||
if (overview.figureAtlas !== 21 || overview.paperLinks !== 100 || overview.paperGroups < 12) failures.push("图表审计或 100 节点阅读链异常");
|
if (overview.figureAtlas !== 21 || overview.paperLinks !== 100 || overview.paperGroups < 12) failures.push("图表审计或 100 节点阅读链异常");
|
||||||
if (overview.labTabs !== 8 || overview.labPanels !== 8) failures.push("八联实验结构异常");
|
if (overview.labTabs !== 8 || overview.labPanels !== 8) failures.push("八联实验结构异常");
|
||||||
|
if (overview.artifactTabs !== 4 || overview.artifactPanels !== 4 || overview.artifactLayers !== 93 || !overview.artifactMismatch) failures.push("开放工件四视图、93 层条带或形状冲突异常");
|
||||||
if (!overview.nativeVisionCorrected || overview.staleVisionClaim) failures.push("原生多模态纠错未生效或旧错误残留");
|
if (!overview.nativeVisionCorrected || overview.staleVisionClaim) failures.push("原生多模态纠错未生效或旧错误残留");
|
||||||
if (overview.documentOverflow > 1 || mobile.documentOverflow > 1) failures.push("桌面或移动端存在文档级横向溢出");
|
if (overview.documentOverflow > 1 || mobile.documentOverflow > 1) failures.push("桌面或移动端存在文档级横向溢出");
|
||||||
if (labs.memoryInitial.panel !== "memory" || numeric(labs.memoryInitial.additiveError) <= numeric(labs.memoryInitial.deltaError)) failures.push("Delta memory 初始递推异常");
|
if (labs.memoryInitial.panel !== "memory" || numeric(labs.memoryInitial.additiveError) <= numeric(labs.memoryInitial.deltaError)) failures.push("Delta memory 初始递推异常");
|
||||||
@@ -227,7 +328,16 @@ if (!labs.rlOver.budget.includes("reward 改为 −1")) failures.push("Reasoning
|
|||||||
if (numeric(labs.cacheInitial.hit) !== 2560 || !labs.cacheInitial.recompute.includes("256")) failures.push("Hybrid prefix cache 默认命中异常");
|
if (numeric(labs.cacheInitial.hit) !== 2560 || !labs.cacheInitial.recompute.includes("256")) failures.push("Hybrid prefix cache 默认命中异常");
|
||||||
if (numeric(labs.cacheSparse.hit) >= numeric(labs.cacheInitial.hit) || numeric(labs.cacheSparse.recompute) <= numeric(labs.cacheInitial.recompute)) failures.push("稀疏 KDA checkpoint 未降低 joint hit");
|
if (numeric(labs.cacheSparse.hit) >= numeric(labs.cacheInitial.hit) || numeric(labs.cacheSparse.recompute) <= numeric(labs.cacheInitial.recompute)) failures.push("稀疏 KDA checkpoint 未降低 joint hit");
|
||||||
if (labs.keyboardSelected !== "decay" || labs.keyboardVisible !== "decay") failures.push("实验键盘 tab 导航异常");
|
if (labs.keyboardSelected !== "decay" || labs.keyboardVisible !== "decay") failures.push("实验键盘 tab 导航异常");
|
||||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 8) failures.push("移动端导航或实验异常");
|
if (artifacts.initial.panel !== "layers" || numeric(artifacts.initial.layer) !== 1 || artifacts.initial.attention !== "KDA" || artifacts.initial.ffn !== "DENSE") failures.push("开放工件初始层视图异常");
|
||||||
|
if (numeric(artifacts.terminal.layer) !== 93 || artifacts.terminal.attention !== "MLA" || artifacts.terminal.ffn !== "MOE" || !artifacts.terminal.copy.includes("L92 / L93")) failures.push("K3 末层真实配置条带异常");
|
||||||
|
if (artifacts.tensors.panel !== "tensors" || artifacts.tensors.groups !== 3 || artifacts.tensors.expertRows !== 6 || !artifacts.tensors.entries || !artifacts.tensors.share) failures.push("checkpoint tensor anatomy 异常");
|
||||||
|
if (!artifacts.mla.visible || artifacts.mla.rows !== 5 || !artifacts.mla.has576) failures.push("MLA header shape 视图异常");
|
||||||
|
if (artifacts.parameterInitial.panel !== "parameters" || artifacts.parameterInitial.shape !== "[128] F32" || !artifacts.parameterInitial.conflict) failures.push("A_log 工件冲突审计异常");
|
||||||
|
if (artifacts.parameterChanged.shape !== "[96,128] F32" || !artifacts.parameterChanged.count.includes("12,288")) failures.push("真实 dt_bias 参数切换异常");
|
||||||
|
if (artifacts.reproductionInitial.panel !== "reproduction" || numeric(artifacts.reproductionInitial.speedup) !== 1.85 || numeric(artifacts.reproductionInitial.cv) < 2) failures.push("FlashKDA H20 或 router 初始探针异常");
|
||||||
|
if (numeric(artifacts.reproductionChanged.speedup) !== 3.27 || numeric(artifacts.reproductionChanged.flash) !== 0.7064 || numeric(artifacts.reproductionChanged.cv) <= numeric(artifacts.reproductionInitial.cv) || numeric(artifacts.reproductionChanged.zero) <= numeric(artifacts.reproductionInitial.zero)) failures.push("GB200 benchmark 或 synthetic router counterexample 未更新");
|
||||||
|
if (artifacts.keyboardSelected !== "tensors" || artifacts.keyboardVisible !== "tensors") failures.push("开放工件键盘 tab 导航异常");
|
||||||
|
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 8 || mobile.artifactTabs !== 4 || mobile.artifactLayers !== 93) failures.push("移动端导航或实验异常");
|
||||||
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
||||||
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,791 @@
|
|||||||
|
---
|
||||||
|
import {
|
||||||
|
k3ArtifactEvidence,
|
||||||
|
k3ArtifactLayers,
|
||||||
|
k3ArtifactSnapshot as snapshot,
|
||||||
|
k3ArtifactViews,
|
||||||
|
} from "@/data/k3Artifacts";
|
||||||
|
|
||||||
|
const checkpoint = snapshot.checkpoint;
|
||||||
|
const audit = snapshot.parameter_audit;
|
||||||
|
const probe = snapshot.router_stress_probe;
|
||||||
|
const flash = snapshot.flashkda;
|
||||||
|
|
||||||
|
const bytes = (value: number) => {
|
||||||
|
if (value >= 2 ** 40) return `${(value / 2 ** 40).toFixed(3)} TiB`;
|
||||||
|
if (value >= 2 ** 30) return `${(value / 2 ** 30).toFixed(2)} GiB`;
|
||||||
|
if (value >= 2 ** 20) return `${(value / 2 ** 20).toFixed(2)} MiB`;
|
||||||
|
return `${value.toLocaleString("en-US")} B`;
|
||||||
|
};
|
||||||
|
|
||||||
|
const parameterRows = [
|
||||||
|
{
|
||||||
|
id: "alog",
|
||||||
|
label: "A_log / checkpoint",
|
||||||
|
shape: "[128] F32",
|
||||||
|
source: audit.a_log,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
id: "dt",
|
||||||
|
label: "dt_bias / checkpoint",
|
||||||
|
shape: "[96,128] F32",
|
||||||
|
source: audit.dt_bias,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
id: "router-bias",
|
||||||
|
label: "router correction bias",
|
||||||
|
shape: "[896] F32",
|
||||||
|
source: audit.router_correction_bias,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
id: "router-norm",
|
||||||
|
label: "router row L2",
|
||||||
|
shape: "[896] derived",
|
||||||
|
source: audit.router_row_l2,
|
||||||
|
},
|
||||||
|
];
|
||||||
|
|
||||||
|
const tensorGroups = [
|
||||||
|
{
|
||||||
|
id: "expert",
|
||||||
|
label: "ROUTED EXPERT 0",
|
||||||
|
boundary: "O / 一个真实 expert 的 packed tensors",
|
||||||
|
rows: snapshot.tensor_examples.routed_expert_0,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
id: "mla",
|
||||||
|
label: "LAYER 4 / GATED MLA",
|
||||||
|
boundary: "O / 真实 shard 4 header",
|
||||||
|
rows: snapshot.tensor_examples.mla_layer_4,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
id: "vision",
|
||||||
|
label: "MOONVIT-V2",
|
||||||
|
boundary: "O / 真实 shard 96 header",
|
||||||
|
rows: snapshot.tensor_examples.vision,
|
||||||
|
},
|
||||||
|
];
|
||||||
|
|
||||||
|
const benchmarkDevices = [
|
||||||
|
["h20", "H20", flash.official_benchmarks.h20],
|
||||||
|
["gb200", "GB200", flash.official_benchmarks.gb200],
|
||||||
|
] as const;
|
||||||
|
---
|
||||||
|
|
||||||
|
<figure class="artifact-lab" data-k3-artifact-lab>
|
||||||
|
<header class="artifact-head">
|
||||||
|
<div>
|
||||||
|
<p>ROUND 03 / OPEN ARTIFACT FORENSICS</p>
|
||||||
|
<h3>不加载 1.56 TB,也能从 config、tensor header 与小范围权重读出真实结构</h3>
|
||||||
|
</div>
|
||||||
|
<p>
|
||||||
|
HF revision <code>{snapshot.provenance.huggingface_revision.slice(0, 9)}</code> ·
|
||||||
|
FlashKDA <code>{snapshot.provenance.flashkda_revision.slice(0, 9)}</code>。
|
||||||
|
原始权重不进仓库;真实观测、推导、执行、合成探针与未决矛盾分别标记。
|
||||||
|
</p>
|
||||||
|
</header>
|
||||||
|
|
||||||
|
<div class="artifact-tabs" role="tablist" aria-label="选择 K3 开放工件实验">
|
||||||
|
{k3ArtifactViews.map(([id, number, title, subtitle], index) => (
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
role="tab"
|
||||||
|
data-artifact-tab={id}
|
||||||
|
aria-selected={index === 0 ? "true" : "false"}
|
||||||
|
tabindex={index === 0 ? "0" : "-1"}
|
||||||
|
>
|
||||||
|
<span>{number}</span><b>{title}</b><small>{subtitle}</small>
|
||||||
|
</button>
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<section class="artifact-panel" data-artifact-panel="layers">
|
||||||
|
<div class="panel-intro">
|
||||||
|
<div><span>O / CONFIG → EXECUTION STRIP</span><h4>拖动一层:同时看 attention、FFN 与 AttnRes block</h4></div>
|
||||||
|
<p>层号按人类阅读使用 1–93;公开 Python 实现内部使用 0-based index。L92 与 L93 连续两个 MLA 是真实配置。</p>
|
||||||
|
</div>
|
||||||
|
<div class="layer-control">
|
||||||
|
<label>
|
||||||
|
<span>SELECT LAYER <output data-layer-label>1 / 93</output></span>
|
||||||
|
<input data-layer-slider type="range" min="1" max="93" value="1" />
|
||||||
|
</label>
|
||||||
|
<div class="layer-legend">
|
||||||
|
<i class="kda"></i><span>KDA · 69</span>
|
||||||
|
<i class="mla"></i><span>MLA · 24</span>
|
||||||
|
<i class="open"></i><span>AttnRes source · 8</span>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<div class="layer-strip" aria-label="K3 真实 93 层配置条带">
|
||||||
|
{k3ArtifactLayers.map((layer) => (
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
data-layer-cell
|
||||||
|
data-layer={layer.layer}
|
||||||
|
data-attention={layer.attention}
|
||||||
|
data-ffn={layer.feedForward}
|
||||||
|
data-block={layer.block}
|
||||||
|
data-open={String(layer.opensResidualBlock)}
|
||||||
|
data-terminal={String(layer.terminalGlobal)}
|
||||||
|
class:list={[
|
||||||
|
layer.attention.toLowerCase(),
|
||||||
|
{ open: layer.opensResidualBlock, terminal: layer.terminalGlobal },
|
||||||
|
]}
|
||||||
|
aria-label={`第 ${layer.layer} 层,${layer.attention},${layer.feedForward},AttnRes block ${layer.block}`}
|
||||||
|
title={`L${layer.layer} · ${layer.attention} · ${layer.feedForward} · B${layer.block}`}
|
||||||
|
>
|
||||||
|
<span>{layer.layer}</span>
|
||||||
|
</button>
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
<div class="layer-readout">
|
||||||
|
<article><span>LAYER</span><b data-layer-number>01</b><p data-layer-special>首层 dense;后续 92 层进入 MoE。</p></article>
|
||||||
|
<article><span>SEQUENCE MIXER</span><b data-layer-attention>KDA</b><p data-layer-attention-copy>96 heads × 128 dims · recurrent state</p></article>
|
||||||
|
<article><span>WIDTH MIXER</span><b data-layer-ffn>DENSE</b><p data-layer-ffn-copy>33792 intermediate · BF16</p></article>
|
||||||
|
<article class="dark"><span>DEPTH SOURCE</span><b data-layer-block>B1 · OPEN</b><p data-layer-block-copy>这一层把 prefix sum 写入 block-level source。</p></article>
|
||||||
|
</div>
|
||||||
|
<div class="boundary"><b>O / exact public config</b><p>条带不代表每层 FLOPs 相同;KDA、MLA、dense 与 896→16 MoE 的状态和执行代价不同。</p></div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="artifact-panel" data-artifact-panel="tensors" hidden>
|
||||||
|
<div class="panel-intro">
|
||||||
|
<div><span>O + D / SAFETENSORS INDEX</span><h4>497,220 个 entries 里,绝大多数为什么来自 experts?</h4></div>
|
||||||
|
<p>entry count、dtype、shape 与 byte offsets 来自公开 index/header;占比是确定性算术,不是运行显存。</p>
|
||||||
|
</div>
|
||||||
|
<div class="artifact-metrics">
|
||||||
|
<article><span>SHARDS</span><b>{checkpoint.shards}</b><p>96 个公开 safetensors</p></article>
|
||||||
|
<article><span>TENSOR DATA</span><b>{checkpoint.tensor_data_tib.toFixed(3)} TiB</b><p>{checkpoint.tensor_data_tb.toFixed(3)} TB decimal</p></article>
|
||||||
|
<article><span>ENTRIES</span><b>{checkpoint.tensor_entries.toLocaleString("en-US")}</b><p>不是 parameter count</p></article>
|
||||||
|
<article class="dark"><span>ROUTED WEIGHT SHARE</span><b>{(checkpoint.derived_routed_expert_share * 100).toFixed(2)}%</b><p>D / packed weights + scales</p></article>
|
||||||
|
</div>
|
||||||
|
<div class="tensor-ledger">
|
||||||
|
<div>
|
||||||
|
<span>247,296</span><b>packed expert tensors</b><p>92 layers × 896 experts × w1/w2/w3</p>
|
||||||
|
</div>
|
||||||
|
<div>
|
||||||
|
<span>247,296</span><b>scale tensors</b><p>每个 packed matrix 独立保存 group scales</p>
|
||||||
|
</div>
|
||||||
|
<div>
|
||||||
|
<span>187 + 187</span><b>AttnRes proj / norm</b><p>每层 attention + MLP 两次读取,再加 output</p>
|
||||||
|
</div>
|
||||||
|
<div>
|
||||||
|
<span>165 + 3</span><b>vision / projector</b><p>MoonViT-V2 与 shared embedding bridge</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<div class="tensor-selector">
|
||||||
|
<span>INSPECT HEADER GROUP</span>
|
||||||
|
{tensorGroups.map((group, index) => (
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
data-tensor-tab={group.id}
|
||||||
|
aria-pressed={index === 0 ? "true" : "false"}
|
||||||
|
>
|
||||||
|
{group.label}
|
||||||
|
</button>
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
{tensorGroups.map((group, index) => (
|
||||||
|
<div class="tensor-table" data-tensor-panel={group.id} hidden={index !== 0}>
|
||||||
|
<header><span>{group.label}</span><b>{group.boundary}</b></header>
|
||||||
|
{group.rows.map((row) => (
|
||||||
|
<div>
|
||||||
|
<code>{row.name.replace("language_model.model.layers.3.self_attn.", "").replace("vision_tower.", "")}</code>
|
||||||
|
<span>{row.dtype}</span>
|
||||||
|
<b>{row.shape.join(" × ")}</b>
|
||||||
|
<small>{bytes(row.bytes)}</small>
|
||||||
|
</div>
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
))}
|
||||||
|
<div class="boundary"><b>O header / D aggregate</b><p>MXFP4 packed shape 不是原始逻辑矩阵 shape;必须结合 latent width、group size 与 loader 格式解释。</p></div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="artifact-panel" data-artifact-panel="parameters" hidden>
|
||||||
|
<div class="panel-intro">
|
||||||
|
<div><span>O + U / SELECTED WEIGHT RANGES</span><h4>几十 KB 真实参数,首先暴露的是一个不能擅自修掉的矛盾</h4></div>
|
||||||
|
<p>只读 shard 1 的 49,664 bytes 与 shard 2 的 13.49 MB prefix;统计可复现,原始参数不再分发。</p>
|
||||||
|
</div>
|
||||||
|
<div class="shape-conflict">
|
||||||
|
<div><span>CONFIG</span><b>96 heads × 128 dims</b><p>q/k/v projection = 12,288</p></div>
|
||||||
|
<i>≠</i>
|
||||||
|
<div class="warn"><span>CHECKPOINT</span><b>A_log [128]</b><p>真实 shard header</p></div>
|
||||||
|
<i>≠</i>
|
||||||
|
<div><span>CODE + KERNEL API</span><b>A_log [H] = [96]</b><p>remote code 与 C++ TORCH_CHECK</p></div>
|
||||||
|
</div>
|
||||||
|
<div class="parameter-control">
|
||||||
|
<label>
|
||||||
|
<span>PARAMETER VIEW</span>
|
||||||
|
<select data-parameter-select>
|
||||||
|
{parameterRows.map((row) => <option value={row.id}>{row.label}</option>)}
|
||||||
|
</select>
|
||||||
|
</label>
|
||||||
|
<p><b data-parameter-shape>[128] F32</b><span data-parameter-count>128 values</span></p>
|
||||||
|
</div>
|
||||||
|
<div class="distribution">
|
||||||
|
{parameterRows.map((row, index) => (
|
||||||
|
<div
|
||||||
|
data-parameter-row={row.id}
|
||||||
|
hidden={index !== 0}
|
||||||
|
data-min={row.source.quantiles.min}
|
||||||
|
data-p10={row.source.quantiles.p10}
|
||||||
|
data-p50={row.source.quantiles.p50}
|
||||||
|
data-p90={row.source.quantiles.p90}
|
||||||
|
data-max={row.source.quantiles.max}
|
||||||
|
data-mean={row.source.mean}
|
||||||
|
data-std={row.source.std}
|
||||||
|
data-count={row.source.count}
|
||||||
|
data-shape={row.shape}
|
||||||
|
>
|
||||||
|
<div class="axis">
|
||||||
|
<i style="left:10%"></i><i style="left:50%"></i><i style="left:90%"></i>
|
||||||
|
</div>
|
||||||
|
<div class="distribution-stats">
|
||||||
|
<article><span>MIN</span><b>{row.source.quantiles.min.toFixed(4)}</b></article>
|
||||||
|
<article><span>P10</span><b>{row.source.quantiles.p10.toFixed(4)}</b></article>
|
||||||
|
<article><span>MEDIAN</span><b>{row.source.quantiles.p50.toFixed(4)}</b></article>
|
||||||
|
<article><span>P90</span><b>{row.source.quantiles.p90.toFixed(4)}</b></article>
|
||||||
|
<article><span>MAX</span><b>{row.source.quantiles.max.toFixed(4)}</b></article>
|
||||||
|
</div>
|
||||||
|
<p>mean {row.source.mean.toFixed(5)} · std {row.source.std.toFixed(5)}</p>
|
||||||
|
</div>
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
<div class="hypothesis-card">
|
||||||
|
<span>U / NONCANONICAL HYPOTHESIS ONLY</span>
|
||||||
|
<h5>若把 checkpoint `[128]` 临时当作 channel-wise 参数</h5>
|
||||||
|
<div>
|
||||||
|
<p><b>{audit.channelwise_hypothesis.one_step_retention.quantiles.p50.toFixed(4)}</b><span>one-step median retention</span></p>
|
||||||
|
<p><b>{audit.channelwise_hypothesis.retention_after_64_steps.quantiles.p50.toExponential(2)}</b><span>64-step median retention</span></p>
|
||||||
|
<p><b>不能定案</b><span>公开 loader / 官方解释仍缺失</span></p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<div class="boundary danger"><b>U / unresolved inconsistency</b><p>本站只报告形状冲突;不宣布 checkpoint 损坏,也不把 channel-wise 猜测冒充真实 K3 forward。</p></div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="artifact-panel" data-artifact-panel="reproduction" hidden>
|
||||||
|
<div class="panel-intro">
|
||||||
|
<div><span>X + S / WHAT ACTUALLY RAN</span><h4>作者 benchmark、本站编译尝试与合成反例,三者不能写成同一种实测</h4></div>
|
||||||
|
<p>RTX 5090 是 sm_120,但本机 PyTorch CUDA 12.8 低于 FlashKDA README 的 12.9+;kernel 尚未执行。</p>
|
||||||
|
</div>
|
||||||
|
<div class="repro-controls">
|
||||||
|
<label><span>OFFICIAL DEVICE</span><select data-benchmark-device>
|
||||||
|
{benchmarkDevices.map(([id, label]) => <option value={id}>{label}</option>)}
|
||||||
|
</select></label>
|
||||||
|
<label><span>OFFICIAL CASE</span><select data-benchmark-case>
|
||||||
|
<option value="Fixed">Fixed T=8192</option>
|
||||||
|
<option value="Varlen, `seq_lens`=`1024 x 8`">8 × 1024 varlen</option>
|
||||||
|
</select></label>
|
||||||
|
<label><span>ROUTER STRESS</span><select data-router-mode>
|
||||||
|
<option value="raw">without correction bias</option>
|
||||||
|
<option value="bias">with checkpoint bias</option>
|
||||||
|
</select></label>
|
||||||
|
</div>
|
||||||
|
<div class="benchmark-readout">
|
||||||
|
<article><span>FLASHKDA</span><b data-benchmark-flash>2.6220 ms</b><p>O / author repository</p></article>
|
||||||
|
<article><span>FLA CHUNK KDA</span><b data-benchmark-fla>4.8388 ms</b><p>O / same author table</p></article>
|
||||||
|
<article class="accent"><span>AUTHOR SPEEDUP</span><b data-benchmark-speedup>1.85×</b><p>不能外推到 RTX 5090</p></article>
|
||||||
|
<article class="dark"><span>LOCAL KERNEL</span><b>NOT RUN</b><p>CUDA 12.8 < official 12.9+</p></article>
|
||||||
|
</div>
|
||||||
|
<div class="local-build">
|
||||||
|
<article><span>ATTEMPT 01</span><b>g++ 15 rejected</b><p>CUDA 12.8 host compiler range要求 <14。</p></article>
|
||||||
|
<i>→</i>
|
||||||
|
<article><span>ATTEMPT 02</span><b>g++ 13 reached nvcc</b><p>随后在 glibc math declarations 处与 CUDA 12.8 headers 冲突。</p></article>
|
||||||
|
<i>→</i>
|
||||||
|
<article class="warn"><span>NEXT GATE</span><b>CUDA 12.9+ matched env</b><p>再跑 exact correctness 与本机 benchmark。</p></article>
|
||||||
|
</div>
|
||||||
|
<div class="router-counterexample">
|
||||||
|
<div>
|
||||||
|
<span>S / REAL WEIGHTS, SYNTHETIC HIDDEN</span>
|
||||||
|
<h5>随机 RMS=1 输入为什么不能评价 Quantile Balancing</h5>
|
||||||
|
<p>2,048 个固定 seed 向量通过真实 `896×7168` router;它们不是模型 token hidden states。</p>
|
||||||
|
</div>
|
||||||
|
<div class="router-stats">
|
||||||
|
<p><span>LOAD CV</span><b data-router-cv>{probe.without_correction_bias.cv.toFixed(3)}</b></p>
|
||||||
|
<p><span>GINI</span><b data-router-gini>{probe.without_correction_bias.gini.toFixed(3)}</b></p>
|
||||||
|
<p><span>ZERO EXPERTS</span><b data-router-zero>{probe.without_correction_bias.zero_experts}</b></p>
|
||||||
|
<p><span>TOP-16 OVERLAP</span><b>{probe.membership_overlap_mean.toFixed(2)} / 16</b></p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<div class="boundary"><b>X/S boundary</b><p>编译失败是本站真实执行结果;router counterexample 只证明 hidden distribution 不可省略,不证明真实 QB 变好或变坏。</p></div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<footer class="evidence-strip">
|
||||||
|
{k3ArtifactEvidence.map(([code, title, copy]) => (
|
||||||
|
<div><span>{code}</span><b>{title}</b><p>{copy}</p></div>
|
||||||
|
))}
|
||||||
|
</footer>
|
||||||
|
</figure>
|
||||||
|
|
||||||
|
<script>
|
||||||
|
const roots = document.querySelectorAll<HTMLElement>("[data-k3-artifact-lab]");
|
||||||
|
|
||||||
|
roots.forEach((root) => {
|
||||||
|
const $ = <T extends HTMLElement = HTMLElement>(selector: string) => root.querySelector<T>(selector)!;
|
||||||
|
const $$ = <T extends HTMLElement = HTMLElement>(selector: string) => [...root.querySelectorAll<T>(selector)];
|
||||||
|
const put = (selector: string, value: string | number) => {
|
||||||
|
const node = $(selector);
|
||||||
|
if (node) node.textContent = String(value);
|
||||||
|
};
|
||||||
|
|
||||||
|
const tabs = $$<HTMLButtonElement>("[data-artifact-tab]");
|
||||||
|
const panels = $$<HTMLElement>("[data-artifact-panel]");
|
||||||
|
const selectTab = (id: string) => {
|
||||||
|
tabs.forEach((tab) => {
|
||||||
|
const active = tab.dataset.artifactTab === id;
|
||||||
|
tab.setAttribute("aria-selected", String(active));
|
||||||
|
tab.tabIndex = active ? 0 : -1;
|
||||||
|
});
|
||||||
|
panels.forEach((panel) => panel.hidden = panel.dataset.artifactPanel !== id);
|
||||||
|
};
|
||||||
|
tabs.forEach((tab, index) => {
|
||||||
|
tab.addEventListener("click", () => selectTab(tab.dataset.artifactTab ?? "layers"));
|
||||||
|
tab.addEventListener("keydown", (event: KeyboardEvent) => {
|
||||||
|
if (!["ArrowLeft", "ArrowRight", "Home", "End"].includes(event.key)) return;
|
||||||
|
event.preventDefault();
|
||||||
|
let next = index;
|
||||||
|
if (event.key === "ArrowLeft") next = (index - 1 + tabs.length) % tabs.length;
|
||||||
|
if (event.key === "ArrowRight") next = (index + 1) % tabs.length;
|
||||||
|
if (event.key === "Home") next = 0;
|
||||||
|
if (event.key === "End") next = tabs.length - 1;
|
||||||
|
selectTab(tabs[next].dataset.artifactTab ?? "layers");
|
||||||
|
tabs[next].focus();
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
const layerSlider = $<HTMLInputElement>("[data-layer-slider]");
|
||||||
|
const cells = $$<HTMLButtonElement>("[data-layer-cell]");
|
||||||
|
const renderLayer = () => {
|
||||||
|
const selected = Number(layerSlider.value);
|
||||||
|
const cell = cells.find((item) => Number(item.dataset.layer) === selected)!;
|
||||||
|
cells.forEach((item) => item.classList.toggle("active", item === cell));
|
||||||
|
const attention = cell.dataset.attention ?? "KDA";
|
||||||
|
const ffn = cell.dataset.ffn ?? "MOE";
|
||||||
|
const block = cell.dataset.block ?? "1";
|
||||||
|
const opens = cell.dataset.open === "true";
|
||||||
|
const terminal = cell.dataset.terminal === "true";
|
||||||
|
put("[data-layer-label]", `${selected} / 93`);
|
||||||
|
put("[data-layer-number]", String(selected).padStart(2, "0"));
|
||||||
|
put("[data-layer-attention]", attention);
|
||||||
|
put("[data-layer-ffn]", ffn);
|
||||||
|
put("[data-layer-block]", `B${block} · ${opens ? "OPEN" : "READ"}`);
|
||||||
|
put("[data-layer-special]", terminal
|
||||||
|
? "额外末层 Gated MLA;因此 L92 / L93 连续全局 attention。"
|
||||||
|
: selected === 1
|
||||||
|
? "首层 dense;后续 92 层进入 MoE。"
|
||||||
|
: `第 ${Math.ceil(selected / 4)} 个 hybrid 节奏位置。`);
|
||||||
|
put("[data-layer-attention-copy]", attention === "KDA"
|
||||||
|
? "96 heads × 128 dims · recurrent state"
|
||||||
|
: "512 latent + 64 auxiliary · No rotary transform");
|
||||||
|
put("[data-layer-ffn-copy]", ffn === "DENSE"
|
||||||
|
? "33792 intermediate · BF16"
|
||||||
|
: "7168 → 3584 latent → 896 choose 16");
|
||||||
|
put("[data-layer-block-copy]", opens
|
||||||
|
? "这一层把 prefix sum 写入新的 block-level source。"
|
||||||
|
: `读取 embedding 与此前 ${opens ? block : Math.min(Number(block), 8)} 个 block source。`);
|
||||||
|
};
|
||||||
|
layerSlider.addEventListener("input", renderLayer);
|
||||||
|
cells.forEach((cell) => cell.addEventListener("click", () => {
|
||||||
|
layerSlider.value = cell.dataset.layer ?? "1";
|
||||||
|
renderLayer();
|
||||||
|
}));
|
||||||
|
renderLayer();
|
||||||
|
|
||||||
|
const tensorTabs = $$<HTMLButtonElement>("[data-tensor-tab]");
|
||||||
|
const tensorPanels = $$<HTMLElement>("[data-tensor-panel]");
|
||||||
|
tensorTabs.forEach((button) => button.addEventListener("click", () => {
|
||||||
|
const id = button.dataset.tensorTab;
|
||||||
|
tensorTabs.forEach((item) => item.setAttribute("aria-pressed", String(item === button)));
|
||||||
|
tensorPanels.forEach((panel) => panel.hidden = panel.dataset.tensorPanel !== id);
|
||||||
|
}));
|
||||||
|
|
||||||
|
const parameterSelect = $<HTMLSelectElement>("[data-parameter-select]");
|
||||||
|
const parameterRows = $$<HTMLElement>("[data-parameter-row]");
|
||||||
|
const renderParameter = () => {
|
||||||
|
const row = parameterRows.find((item) => item.dataset.parameterRow === parameterSelect.value)!;
|
||||||
|
parameterRows.forEach((item) => item.hidden = item !== row);
|
||||||
|
put("[data-parameter-shape]", row.dataset.shape ?? "");
|
||||||
|
put("[data-parameter-count]", `${Number(row.dataset.count).toLocaleString("en-US")} values`);
|
||||||
|
};
|
||||||
|
parameterSelect.addEventListener("input", renderParameter);
|
||||||
|
renderParameter();
|
||||||
|
|
||||||
|
const benchmarks = {
|
||||||
|
h20: {
|
||||||
|
Fixed: { flash: 2.6220, fla: 4.8388, speedup: 1.85 },
|
||||||
|
"Varlen, `seq_lens`=`1024 x 8`": { flash: 2.0432, fla: 4.6723, speedup: 2.29 },
|
||||||
|
},
|
||||||
|
gb200: {
|
||||||
|
Fixed: { flash: 1.0087, fla: 2.3271, speedup: 2.31 },
|
||||||
|
"Varlen, `seq_lens`=`1024 x 8`": { flash: .7064, fla: 2.3105, speedup: 3.27 },
|
||||||
|
},
|
||||||
|
};
|
||||||
|
const device = $<HTMLSelectElement>("[data-benchmark-device]");
|
||||||
|
const benchmarkCase = $<HTMLSelectElement>("[data-benchmark-case]");
|
||||||
|
const renderBenchmark = () => {
|
||||||
|
const row = benchmarks[device.value as keyof typeof benchmarks][benchmarkCase.value as "Fixed" | "Varlen, `seq_lens`=`1024 x 8`"];
|
||||||
|
put("[data-benchmark-flash]", `${row.flash.toFixed(4)} ms`);
|
||||||
|
put("[data-benchmark-fla]", `${row.fla.toFixed(4)} ms`);
|
||||||
|
put("[data-benchmark-speedup]", `${row.speedup.toFixed(2)}×`);
|
||||||
|
};
|
||||||
|
device.addEventListener("input", renderBenchmark);
|
||||||
|
benchmarkCase.addEventListener("input", renderBenchmark);
|
||||||
|
|
||||||
|
const routerMode = $<HTMLSelectElement>("[data-router-mode]");
|
||||||
|
const routerRows = {
|
||||||
|
raw: { cv: 2.0845208168, gini: .8310886025, zero: 558 },
|
||||||
|
bias: { cv: 2.5288832188, gini: .8788146973, zero: 673 },
|
||||||
|
};
|
||||||
|
const renderRouter = () => {
|
||||||
|
const row = routerRows[routerMode.value as keyof typeof routerRows];
|
||||||
|
put("[data-router-cv]", row.cv.toFixed(3));
|
||||||
|
put("[data-router-gini]", row.gini.toFixed(3));
|
||||||
|
put("[data-router-zero]", row.zero);
|
||||||
|
};
|
||||||
|
routerMode.addEventListener("input", renderRouter);
|
||||||
|
});
|
||||||
|
</script>
|
||||||
|
|
||||||
|
<style>
|
||||||
|
.artifact-lab {
|
||||||
|
max-width: 1080px;
|
||||||
|
margin: 42px 0;
|
||||||
|
border: 1px solid var(--ink);
|
||||||
|
background: var(--paper-raised);
|
||||||
|
}
|
||||||
|
.artifact-head,
|
||||||
|
.panel-intro {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: minmax(0, 1.15fr) minmax(280px, .85fr);
|
||||||
|
gap: 30px;
|
||||||
|
padding: 28px;
|
||||||
|
}
|
||||||
|
.artifact-head { color: var(--paper); background: var(--ink); }
|
||||||
|
.artifact-head > div > p,
|
||||||
|
.panel-intro span,
|
||||||
|
.tensor-selector > span {
|
||||||
|
color: var(--copper);
|
||||||
|
font: .6rem/1.2 var(--mono);
|
||||||
|
letter-spacing: .12em;
|
||||||
|
}
|
||||||
|
.artifact-head h3,
|
||||||
|
.panel-intro h4 { margin: 10px 0 0; font-size: 1.23rem; line-height: 1.35; }
|
||||||
|
.artifact-head h3 { color: var(--paper); }
|
||||||
|
.artifact-head > p,
|
||||||
|
.panel-intro > p { margin: 0; font-size: .72rem; line-height: 1.75; }
|
||||||
|
.artifact-head > p { color: rgba(255,255,255,.66); }
|
||||||
|
.artifact-head code { color: var(--paper); font-size: .66rem; }
|
||||||
|
.artifact-tabs {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: repeat(4, 1fr);
|
||||||
|
border-bottom: 1px solid var(--line);
|
||||||
|
}
|
||||||
|
.artifact-tabs button {
|
||||||
|
min-height: 76px;
|
||||||
|
padding: 13px 16px;
|
||||||
|
border: 0;
|
||||||
|
border-right: 1px solid var(--line);
|
||||||
|
color: var(--ink);
|
||||||
|
background: transparent;
|
||||||
|
text-align: left;
|
||||||
|
cursor: pointer;
|
||||||
|
}
|
||||||
|
.artifact-tabs button:last-child { border-right: 0; }
|
||||||
|
.artifact-tabs button[aria-selected="true"] { color: var(--paper); background: var(--copper); }
|
||||||
|
.artifact-tabs span,
|
||||||
|
.artifact-tabs b,
|
||||||
|
.artifact-tabs small { display: block; }
|
||||||
|
.artifact-tabs span { opacity: .65; font: .55rem/1 var(--mono); }
|
||||||
|
.artifact-tabs b { margin-top: 8px; font: 700 .67rem/1 var(--mono); letter-spacing: .05em; }
|
||||||
|
.artifact-tabs small { margin-top: 5px; opacity: .72; font-size: .61rem; }
|
||||||
|
.artifact-panel { padding-bottom: 26px; }
|
||||||
|
.panel-intro { padding: 24px 28px; border-bottom: 1px solid var(--line); }
|
||||||
|
.panel-intro > p { color: var(--muted); }
|
||||||
|
.layer-control,
|
||||||
|
.parameter-control {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: 1.15fr .85fr;
|
||||||
|
gap: 24px;
|
||||||
|
align-items: center;
|
||||||
|
margin: 24px 28px;
|
||||||
|
padding: 16px;
|
||||||
|
border: 1px solid var(--line);
|
||||||
|
background: var(--paper);
|
||||||
|
}
|
||||||
|
.layer-control label,
|
||||||
|
.parameter-control label,
|
||||||
|
.repro-controls label {
|
||||||
|
color: var(--muted);
|
||||||
|
font: .62rem/1.4 var(--mono);
|
||||||
|
}
|
||||||
|
.layer-control label > span { display: flex; justify-content: space-between; }
|
||||||
|
.layer-control output { color: var(--copper); }
|
||||||
|
.layer-control input { width: 100%; margin-top: 12px; accent-color: var(--copper); }
|
||||||
|
.layer-legend { display: grid; grid-template-columns: 12px 1fr 12px 1fr 12px 1fr; gap: 8px; align-items: center; font: .58rem/1.3 var(--mono); }
|
||||||
|
.layer-legend i { width: 12px; height: 12px; }
|
||||||
|
.layer-legend .kda { background: var(--ink); }
|
||||||
|
.layer-legend .mla { background: var(--copper); }
|
||||||
|
.layer-legend .open { border: 2px solid var(--copper); background: transparent; }
|
||||||
|
.layer-strip {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: repeat(31, minmax(7px, 1fr));
|
||||||
|
gap: 3px;
|
||||||
|
margin: 0 28px 24px;
|
||||||
|
}
|
||||||
|
.layer-strip button {
|
||||||
|
position: relative;
|
||||||
|
min-width: 0;
|
||||||
|
height: 32px;
|
||||||
|
padding: 0;
|
||||||
|
border: 1px solid transparent;
|
||||||
|
color: transparent;
|
||||||
|
background: var(--ink);
|
||||||
|
cursor: pointer;
|
||||||
|
}
|
||||||
|
.layer-strip button.mla { background: var(--copper); }
|
||||||
|
.layer-strip button.open::after {
|
||||||
|
content: "";
|
||||||
|
position: absolute;
|
||||||
|
inset: 3px;
|
||||||
|
border: 1px solid var(--paper);
|
||||||
|
}
|
||||||
|
.layer-strip button.terminal { box-shadow: 0 0 0 2px var(--copper); }
|
||||||
|
.layer-strip button.active { outline: 3px solid var(--signal); outline-offset: 1px; z-index: 2; }
|
||||||
|
.layer-strip span { font: .44rem/1 var(--mono); }
|
||||||
|
.layer-readout,
|
||||||
|
.artifact-metrics,
|
||||||
|
.benchmark-readout {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: repeat(4, 1fr);
|
||||||
|
gap: 12px;
|
||||||
|
margin: 0 28px 24px;
|
||||||
|
}
|
||||||
|
.layer-readout article,
|
||||||
|
.artifact-metrics article,
|
||||||
|
.benchmark-readout article {
|
||||||
|
min-width: 0;
|
||||||
|
min-height: 128px;
|
||||||
|
padding: 18px;
|
||||||
|
border: 1px solid var(--line);
|
||||||
|
background: var(--paper);
|
||||||
|
}
|
||||||
|
.layer-readout span,
|
||||||
|
.artifact-metrics span,
|
||||||
|
.benchmark-readout span,
|
||||||
|
.local-build span,
|
||||||
|
.router-counterexample span,
|
||||||
|
.hypothesis-card > span {
|
||||||
|
color: var(--copper);
|
||||||
|
font: .57rem/1.2 var(--mono);
|
||||||
|
letter-spacing: .07em;
|
||||||
|
}
|
||||||
|
.layer-readout b,
|
||||||
|
.artifact-metrics b,
|
||||||
|
.benchmark-readout b { display: block; margin-top: 12px; font: 800 1.08rem/1.1 var(--mono); }
|
||||||
|
.layer-readout p,
|
||||||
|
.artifact-metrics p,
|
||||||
|
.benchmark-readout p,
|
||||||
|
.local-build p { margin: 12px 0 0; color: var(--muted); font-size: .66rem; line-height: 1.55; }
|
||||||
|
.layer-readout article.dark,
|
||||||
|
.artifact-metrics article.dark,
|
||||||
|
.benchmark-readout article.dark { color: var(--paper); background: var(--ink); }
|
||||||
|
.layer-readout article.dark p,
|
||||||
|
.artifact-metrics article.dark p,
|
||||||
|
.benchmark-readout article.dark p { color: rgba(255,255,255,.62); }
|
||||||
|
article.accent { border-color: rgba(173,100,69,.55); background: var(--copper-pale); }
|
||||||
|
.tensor-ledger {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: repeat(4, 1fr);
|
||||||
|
margin: 0 28px 24px;
|
||||||
|
border-top: 1px solid var(--line);
|
||||||
|
border-left: 1px solid var(--line);
|
||||||
|
}
|
||||||
|
.tensor-ledger > div { padding: 16px; border-right: 1px solid var(--line); border-bottom: 1px solid var(--line); }
|
||||||
|
.tensor-ledger span { color: var(--copper); font: 800 .88rem/1 var(--mono); }
|
||||||
|
.tensor-ledger b { display: block; margin-top: 9px; font-size: .73rem; }
|
||||||
|
.tensor-ledger p { margin: 8px 0 0; color: var(--muted); font-size: .62rem; line-height: 1.5; }
|
||||||
|
.tensor-selector {
|
||||||
|
display: flex;
|
||||||
|
gap: 8px;
|
||||||
|
align-items: center;
|
||||||
|
margin: 0 28px 14px;
|
||||||
|
}
|
||||||
|
.tensor-selector > span { margin-right: auto; }
|
||||||
|
.tensor-selector button {
|
||||||
|
padding: 8px 10px;
|
||||||
|
border: 1px solid var(--line);
|
||||||
|
color: var(--ink);
|
||||||
|
background: var(--paper);
|
||||||
|
font: .57rem var(--mono);
|
||||||
|
cursor: pointer;
|
||||||
|
}
|
||||||
|
.tensor-selector button[aria-pressed="true"] { color: var(--paper); border-color: var(--copper); background: var(--copper); }
|
||||||
|
.tensor-table { margin: 0 28px 24px; border: 1px solid var(--line); background: var(--paper); }
|
||||||
|
.tensor-table header,
|
||||||
|
.tensor-table > div {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: minmax(0, 1.8fr) 70px minmax(150px, 1fr) 100px;
|
||||||
|
gap: 14px;
|
||||||
|
align-items: center;
|
||||||
|
padding: 12px 14px;
|
||||||
|
border-bottom: 1px solid var(--line);
|
||||||
|
}
|
||||||
|
.tensor-table > div:last-child { border-bottom: 0; }
|
||||||
|
.tensor-table header { color: var(--paper); background: var(--ink); }
|
||||||
|
.tensor-table header span,
|
||||||
|
.tensor-table header b { font: .57rem/1.3 var(--mono); }
|
||||||
|
.tensor-table header b { grid-column: 2 / -1; }
|
||||||
|
.tensor-table code { overflow-wrap: anywhere; color: var(--ink); font-size: .59rem; }
|
||||||
|
.tensor-table > div span,
|
||||||
|
.tensor-table > div b,
|
||||||
|
.tensor-table > div small { font: .59rem/1.35 var(--mono); }
|
||||||
|
.tensor-table > div span { color: var(--copper); }
|
||||||
|
.tensor-table > div small { color: var(--muted); text-align: right; }
|
||||||
|
.shape-conflict {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: 1fr 24px 1fr 24px 1fr;
|
||||||
|
gap: 10px;
|
||||||
|
align-items: center;
|
||||||
|
margin: 24px 28px;
|
||||||
|
}
|
||||||
|
.shape-conflict > div { min-height: 116px; padding: 18px; border: 1px solid var(--line); background: var(--paper); }
|
||||||
|
.shape-conflict > i,
|
||||||
|
.local-build > i { color: var(--copper); text-align: center; font-style: normal; }
|
||||||
|
.shape-conflict span { color: var(--copper); font: .57rem var(--mono); }
|
||||||
|
.shape-conflict b { display: block; margin-top: 12px; font: 800 .86rem var(--mono); }
|
||||||
|
.shape-conflict p { margin: 10px 0 0; color: var(--muted); font-size: .63rem; }
|
||||||
|
.shape-conflict .warn,
|
||||||
|
.local-build .warn { border-color: var(--signal); background: color-mix(in srgb, var(--signal) 8%, var(--paper)); }
|
||||||
|
.parameter-control label { display: grid; grid-template-columns: 1fr; gap: 9px; }
|
||||||
|
.parameter-control select,
|
||||||
|
.repro-controls select { width: 100%; padding: 9px; border: 1px solid var(--line); color: var(--ink); background: var(--paper-raised); font: .62rem var(--mono); }
|
||||||
|
.parameter-control > p { display: flex; justify-content: space-between; gap: 16px; margin: 0; font: .62rem var(--mono); }
|
||||||
|
.parameter-control > p b { color: var(--copper); }
|
||||||
|
.distribution { margin: 0 28px 24px; padding: 20px; border: 1px solid var(--line); background: var(--paper); }
|
||||||
|
.axis { position: relative; height: 30px; margin: 4px 4px 18px; border-top: 4px solid var(--ink); }
|
||||||
|
.axis::before,
|
||||||
|
.axis::after { content: ""; position: absolute; top: -8px; width: 3px; height: 12px; background: var(--copper); }
|
||||||
|
.axis::before { left: 0; }
|
||||||
|
.axis::after { right: 0; }
|
||||||
|
.axis i { position: absolute; top: -7px; width: 3px; height: 10px; background: var(--signal); }
|
||||||
|
.distribution-stats { display: grid; grid-template-columns: repeat(5, 1fr); gap: 8px; }
|
||||||
|
.distribution-stats article { padding: 12px; border: 1px solid var(--line); }
|
||||||
|
.distribution-stats span,
|
||||||
|
.distribution-stats b { display: block; font: .57rem/1.2 var(--mono); }
|
||||||
|
.distribution-stats span { color: var(--muted); }
|
||||||
|
.distribution-stats b { margin-top: 8px; color: var(--copper); }
|
||||||
|
.distribution [data-parameter-row] > p { margin: 16px 0 0; color: var(--muted); font: .6rem var(--mono); }
|
||||||
|
.hypothesis-card { margin: 0 28px 24px; padding: 20px; border: 1px dashed var(--signal); background: color-mix(in srgb, var(--signal) 6%, var(--paper)); }
|
||||||
|
.hypothesis-card h5,
|
||||||
|
.router-counterexample h5 { margin: 10px 0 16px; font-size: .9rem; }
|
||||||
|
.hypothesis-card > div { display: grid; grid-template-columns: repeat(3, 1fr); gap: 12px; }
|
||||||
|
.hypothesis-card p { margin: 0; padding: 14px; border: 1px solid var(--line); background: var(--paper); }
|
||||||
|
.hypothesis-card p b,
|
||||||
|
.hypothesis-card p span { display: block; }
|
||||||
|
.hypothesis-card p b { font: 800 .9rem var(--mono); }
|
||||||
|
.hypothesis-card p span { margin-top: 8px; color: var(--muted); font-size: .6rem; }
|
||||||
|
.repro-controls {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: repeat(3, 1fr);
|
||||||
|
gap: 12px;
|
||||||
|
margin: 24px 28px;
|
||||||
|
}
|
||||||
|
.repro-controls label { display: grid; gap: 9px; padding: 14px; border: 1px solid var(--line); background: var(--paper); }
|
||||||
|
.local-build {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: 1fr 24px 1fr 24px 1fr;
|
||||||
|
gap: 10px;
|
||||||
|
align-items: center;
|
||||||
|
margin: 0 28px 24px;
|
||||||
|
}
|
||||||
|
.local-build article { min-height: 138px; padding: 18px; border: 1px solid var(--line); background: var(--paper); }
|
||||||
|
.local-build b { display: block; margin-top: 12px; font: 800 .78rem/1.35 var(--mono); }
|
||||||
|
.router-counterexample {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: 1fr 1fr;
|
||||||
|
gap: 22px;
|
||||||
|
margin: 0 28px 24px;
|
||||||
|
padding: 20px;
|
||||||
|
border: 1px solid var(--line);
|
||||||
|
background: var(--paper);
|
||||||
|
}
|
||||||
|
.router-counterexample > div > p { color: var(--muted); font-size: .67rem; line-height: 1.6; }
|
||||||
|
.router-stats { display: grid; grid-template-columns: repeat(2, 1fr); gap: 8px; }
|
||||||
|
.router-stats p { margin: 0; padding: 12px; border: 1px solid var(--line); }
|
||||||
|
.router-stats span,
|
||||||
|
.router-stats b { display: block; }
|
||||||
|
.router-stats b { margin-top: 8px; font: 800 .87rem var(--mono); }
|
||||||
|
.boundary {
|
||||||
|
margin: 0 28px;
|
||||||
|
padding: 14px 16px;
|
||||||
|
border-left: 3px solid var(--copper);
|
||||||
|
background: var(--copper-pale);
|
||||||
|
}
|
||||||
|
.boundary b { font: .61rem var(--mono); }
|
||||||
|
.boundary p { margin: 7px 0 0; color: var(--muted); font-size: .65rem; line-height: 1.55; }
|
||||||
|
.boundary.danger { border-color: var(--signal); background: color-mix(in srgb, var(--signal) 7%, var(--paper)); }
|
||||||
|
.evidence-strip {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: repeat(5, 1fr);
|
||||||
|
margin-top: 26px;
|
||||||
|
border-top: 1px solid var(--line);
|
||||||
|
}
|
||||||
|
.evidence-strip > div { padding: 15px; border-right: 1px solid var(--line); }
|
||||||
|
.evidence-strip > div:last-child { border-right: 0; }
|
||||||
|
.evidence-strip span { color: var(--paper); background: var(--copper); padding: 3px 5px; font: .58rem var(--mono); }
|
||||||
|
.evidence-strip b { display: block; margin-top: 10px; font: .58rem var(--mono); }
|
||||||
|
.evidence-strip p { margin: 8px 0 0; color: var(--muted); font-size: .58rem; line-height: 1.45; }
|
||||||
|
@media (max-width: 760px) {
|
||||||
|
.artifact-head,
|
||||||
|
.panel-intro,
|
||||||
|
.layer-control,
|
||||||
|
.parameter-control,
|
||||||
|
.router-counterexample { grid-template-columns: 1fr; gap: 16px; }
|
||||||
|
.artifact-tabs { grid-template-columns: repeat(2, 1fr); }
|
||||||
|
.artifact-tabs button:nth-child(2) { border-right: 0; }
|
||||||
|
.artifact-tabs button:nth-child(-n+2) { border-bottom: 1px solid var(--line); }
|
||||||
|
.layer-strip { grid-template-columns: repeat(31, minmax(5px, 1fr)); gap: 2px; }
|
||||||
|
.layer-strip button { height: 24px; }
|
||||||
|
.layer-readout,
|
||||||
|
.artifact-metrics,
|
||||||
|
.benchmark-readout,
|
||||||
|
.tensor-ledger { grid-template-columns: repeat(2, 1fr); }
|
||||||
|
.tensor-selector { align-items: stretch; flex-wrap: wrap; }
|
||||||
|
.tensor-selector > span { width: 100%; }
|
||||||
|
.tensor-table { overflow-x: auto; }
|
||||||
|
.tensor-table header,
|
||||||
|
.tensor-table > div { min-width: 690px; }
|
||||||
|
.shape-conflict,
|
||||||
|
.local-build { grid-template-columns: 1fr; }
|
||||||
|
.shape-conflict > i,
|
||||||
|
.local-build > i { transform: rotate(90deg); }
|
||||||
|
.distribution-stats { grid-template-columns: repeat(2, 1fr); }
|
||||||
|
.hypothesis-card > div,
|
||||||
|
.repro-controls { grid-template-columns: 1fr; }
|
||||||
|
.evidence-strip { grid-template-columns: 1fr; }
|
||||||
|
.evidence-strip > div { border-right: 0; border-bottom: 1px solid var(--line); }
|
||||||
|
}
|
||||||
|
@media (max-width: 440px) {
|
||||||
|
.artifact-head,
|
||||||
|
.panel-intro { padding: 21px; }
|
||||||
|
.layer-control,
|
||||||
|
.layer-strip,
|
||||||
|
.layer-readout,
|
||||||
|
.artifact-metrics,
|
||||||
|
.benchmark-readout,
|
||||||
|
.tensor-ledger,
|
||||||
|
.tensor-selector,
|
||||||
|
.tensor-table,
|
||||||
|
.shape-conflict,
|
||||||
|
.parameter-control,
|
||||||
|
.distribution,
|
||||||
|
.hypothesis-card,
|
||||||
|
.repro-controls,
|
||||||
|
.local-build,
|
||||||
|
.router-counterexample,
|
||||||
|
.boundary { margin-left: 16px; margin-right: 16px; }
|
||||||
|
.layer-readout,
|
||||||
|
.artifact-metrics,
|
||||||
|
.benchmark-readout,
|
||||||
|
.tensor-ledger,
|
||||||
|
.router-stats { grid-template-columns: 1fr; }
|
||||||
|
}
|
||||||
|
</style>
|
||||||
@@ -0,0 +1,600 @@
|
|||||||
|
{
|
||||||
|
"schema_version": 1,
|
||||||
|
"captured_at": "2026-07-29T12:21:00+08:00",
|
||||||
|
"evidence_boundary": {
|
||||||
|
"checkpoint_forward_run": false,
|
||||||
|
"raw_weights_committed": false,
|
||||||
|
"router_inputs": "deterministic synthetic RMS-normalized vectors, not token hidden states",
|
||||||
|
"kda_retention_probe": "noncanonical channel-wise interpretation used only to expose the A_log shape ambiguity"
|
||||||
|
},
|
||||||
|
"provenance": {
|
||||||
|
"huggingface_model": "moonshotai/Kimi-K3",
|
||||||
|
"huggingface_revision": "9f62e4e9fffbd0a83ddd60e1c209d828994b3569",
|
||||||
|
"flashkda_revision": "1ce47ea3bb22c84eb9cc665028399cf35e8ffb0b",
|
||||||
|
"sha256": {
|
||||||
|
"config": "9710e121a58d03ac92c8d6da287a19541994319afbbe6d6202af001ffd379213",
|
||||||
|
"index": "a1c5210650ce71d2d3ae9ec5a101ac4afd3cf4b10091be589853437eb967febd",
|
||||||
|
"kda_slice": "dccb10e734a1b6756fe24233ef86c6417b704133a2ec4e56c5c5ddf87274d9b2",
|
||||||
|
"router_prefix": "dbe66ff82e58bbd273514c0a51b74043c519f64df4586166321cf6266e1723ee"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"checkpoint": {
|
||||||
|
"tensor_data_bytes": 1560860324864,
|
||||||
|
"tensor_data_tb": 1.560860324864,
|
||||||
|
"tensor_data_tib": 1.4195941956713796,
|
||||||
|
"shards": 96,
|
||||||
|
"shard_file_bytes": {
|
||||||
|
"sum": 1560936091448,
|
||||||
|
"min": 92289328,
|
||||||
|
"max": 16990916912,
|
||||||
|
"mean": 16259750952.583334
|
||||||
|
},
|
||||||
|
"tensor_entries": 497220,
|
||||||
|
"tensor_counts": {
|
||||||
|
"expert_packed": 247296,
|
||||||
|
"expert_scales": 247296,
|
||||||
|
"router_weight": 92,
|
||||||
|
"router_correction_bias": 92,
|
||||||
|
"attnres_proj": 187,
|
||||||
|
"attnres_norm": 187,
|
||||||
|
"kda_a_log": 69,
|
||||||
|
"kda_dt_bias": 69,
|
||||||
|
"vision": 165,
|
||||||
|
"projector": 3
|
||||||
|
},
|
||||||
|
"derived_routed_expert_bytes": 1446456066048,
|
||||||
|
"derived_routed_expert_share": 0.9267043584915465
|
||||||
|
},
|
||||||
|
"configuration": {
|
||||||
|
"layers": 93,
|
||||||
|
"dense_layers": 1,
|
||||||
|
"hidden": 7168,
|
||||||
|
"vocabulary": 163840,
|
||||||
|
"context": 1048576,
|
||||||
|
"kda_layers": [
|
||||||
|
1,
|
||||||
|
2,
|
||||||
|
3,
|
||||||
|
5,
|
||||||
|
6,
|
||||||
|
7,
|
||||||
|
9,
|
||||||
|
10,
|
||||||
|
11,
|
||||||
|
13,
|
||||||
|
14,
|
||||||
|
15,
|
||||||
|
17,
|
||||||
|
18,
|
||||||
|
19,
|
||||||
|
21,
|
||||||
|
22,
|
||||||
|
23,
|
||||||
|
25,
|
||||||
|
26,
|
||||||
|
27,
|
||||||
|
29,
|
||||||
|
30,
|
||||||
|
31,
|
||||||
|
33,
|
||||||
|
34,
|
||||||
|
35,
|
||||||
|
37,
|
||||||
|
38,
|
||||||
|
39,
|
||||||
|
41,
|
||||||
|
42,
|
||||||
|
43,
|
||||||
|
45,
|
||||||
|
46,
|
||||||
|
47,
|
||||||
|
49,
|
||||||
|
50,
|
||||||
|
51,
|
||||||
|
53,
|
||||||
|
54,
|
||||||
|
55,
|
||||||
|
57,
|
||||||
|
58,
|
||||||
|
59,
|
||||||
|
61,
|
||||||
|
62,
|
||||||
|
63,
|
||||||
|
65,
|
||||||
|
66,
|
||||||
|
67,
|
||||||
|
69,
|
||||||
|
70,
|
||||||
|
71,
|
||||||
|
73,
|
||||||
|
74,
|
||||||
|
75,
|
||||||
|
77,
|
||||||
|
78,
|
||||||
|
79,
|
||||||
|
81,
|
||||||
|
82,
|
||||||
|
83,
|
||||||
|
85,
|
||||||
|
86,
|
||||||
|
87,
|
||||||
|
89,
|
||||||
|
90,
|
||||||
|
91
|
||||||
|
],
|
||||||
|
"mla_layers": [
|
||||||
|
4,
|
||||||
|
8,
|
||||||
|
12,
|
||||||
|
16,
|
||||||
|
20,
|
||||||
|
24,
|
||||||
|
28,
|
||||||
|
32,
|
||||||
|
36,
|
||||||
|
40,
|
||||||
|
44,
|
||||||
|
48,
|
||||||
|
52,
|
||||||
|
56,
|
||||||
|
60,
|
||||||
|
64,
|
||||||
|
68,
|
||||||
|
72,
|
||||||
|
76,
|
||||||
|
80,
|
||||||
|
84,
|
||||||
|
88,
|
||||||
|
92,
|
||||||
|
93
|
||||||
|
],
|
||||||
|
"heads": 96,
|
||||||
|
"head_dim": 128,
|
||||||
|
"attnres_block": 12,
|
||||||
|
"experts": 896,
|
||||||
|
"active_experts": 16,
|
||||||
|
"shared_experts": 2,
|
||||||
|
"latent_width": 3584,
|
||||||
|
"expert_intermediate": 3072,
|
||||||
|
"situ_beta": 4.0,
|
||||||
|
"situ_linear_beta": 25.0,
|
||||||
|
"mla_nope": true,
|
||||||
|
"mla_output_gate": true
|
||||||
|
},
|
||||||
|
"tensor_examples": {
|
||||||
|
"mla_layer_4": [
|
||||||
|
{
|
||||||
|
"name": "language_model.model.layers.3.self_attn.kv_a_proj_with_mqa.weight",
|
||||||
|
"dtype": "BF16",
|
||||||
|
"shape": [
|
||||||
|
576,
|
||||||
|
7168
|
||||||
|
],
|
||||||
|
"bytes": 8257536
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "language_model.model.layers.3.self_attn.kv_b_proj.weight",
|
||||||
|
"dtype": "BF16",
|
||||||
|
"shape": [
|
||||||
|
24576,
|
||||||
|
512
|
||||||
|
],
|
||||||
|
"bytes": 25165824
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "language_model.model.layers.3.self_attn.q_a_proj.weight",
|
||||||
|
"dtype": "BF16",
|
||||||
|
"shape": [
|
||||||
|
1536,
|
||||||
|
7168
|
||||||
|
],
|
||||||
|
"bytes": 22020096
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "language_model.model.layers.3.self_attn.q_b_proj.weight",
|
||||||
|
"dtype": "BF16",
|
||||||
|
"shape": [
|
||||||
|
18432,
|
||||||
|
1536
|
||||||
|
],
|
||||||
|
"bytes": 56623104
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "language_model.model.layers.3.self_attn.g_proj.weight",
|
||||||
|
"dtype": "BF16",
|
||||||
|
"shape": [
|
||||||
|
12288,
|
||||||
|
7168
|
||||||
|
],
|
||||||
|
"bytes": 176160768
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"vision": [
|
||||||
|
{
|
||||||
|
"name": "vision_tower.patch_embed.proj.weight",
|
||||||
|
"dtype": "BF16",
|
||||||
|
"shape": [
|
||||||
|
1024,
|
||||||
|
3,
|
||||||
|
14,
|
||||||
|
14
|
||||||
|
],
|
||||||
|
"bytes": 1204224
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "vision_tower.patch_embed.pos_emb.weight",
|
||||||
|
"dtype": "BF16",
|
||||||
|
"shape": [
|
||||||
|
64,
|
||||||
|
64,
|
||||||
|
1024
|
||||||
|
],
|
||||||
|
"bytes": 8388608
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "vision_tower.encoder.blocks.0.wqkv.weight",
|
||||||
|
"dtype": "BF16",
|
||||||
|
"shape": [
|
||||||
|
4608,
|
||||||
|
1024
|
||||||
|
],
|
||||||
|
"bytes": 9437184
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "vision_tower.encoder.blocks.26.wqkv.weight",
|
||||||
|
"dtype": "BF16",
|
||||||
|
"shape": [
|
||||||
|
4608,
|
||||||
|
1024
|
||||||
|
],
|
||||||
|
"bytes": 9437184
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "vision_tower.encoder.final_layernorm.weight",
|
||||||
|
"dtype": "BF16",
|
||||||
|
"shape": [
|
||||||
|
1024
|
||||||
|
],
|
||||||
|
"bytes": 2048
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"routed_expert_0": [
|
||||||
|
{
|
||||||
|
"name": "w1.weight_packed",
|
||||||
|
"dtype": "U8",
|
||||||
|
"shape": [
|
||||||
|
3072,
|
||||||
|
1792
|
||||||
|
],
|
||||||
|
"bytes": 5505024
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "w1.weight_scale",
|
||||||
|
"dtype": "U8",
|
||||||
|
"shape": [
|
||||||
|
3072,
|
||||||
|
112
|
||||||
|
],
|
||||||
|
"bytes": 344064
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "w2.weight_packed",
|
||||||
|
"dtype": "U8",
|
||||||
|
"shape": [
|
||||||
|
3584,
|
||||||
|
1536
|
||||||
|
],
|
||||||
|
"bytes": 5505024
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "w2.weight_scale",
|
||||||
|
"dtype": "U8",
|
||||||
|
"shape": [
|
||||||
|
3584,
|
||||||
|
96
|
||||||
|
],
|
||||||
|
"bytes": 344064
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "w3.weight_packed",
|
||||||
|
"dtype": "U8",
|
||||||
|
"shape": [
|
||||||
|
3072,
|
||||||
|
1792
|
||||||
|
],
|
||||||
|
"bytes": 5505024
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "w3.weight_scale",
|
||||||
|
"dtype": "U8",
|
||||||
|
"shape": [
|
||||||
|
3072,
|
||||||
|
112
|
||||||
|
],
|
||||||
|
"bytes": 344064
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"parameter_audit": {
|
||||||
|
"a_log_checkpoint_shape": [
|
||||||
|
128
|
||||||
|
],
|
||||||
|
"a_log_public_code_shape": [
|
||||||
|
96
|
||||||
|
],
|
||||||
|
"dt_bias_shape": [
|
||||||
|
96,
|
||||||
|
128
|
||||||
|
],
|
||||||
|
"beta_projection_shape": [
|
||||||
|
96,
|
||||||
|
7168
|
||||||
|
],
|
||||||
|
"status": "observed shape inconsistency; runtime meaning unresolved",
|
||||||
|
"a_log": {
|
||||||
|
"count": 128,
|
||||||
|
"mean": -0.16894778609275818,
|
||||||
|
"std": 0.3667531907558441,
|
||||||
|
"quantiles": {
|
||||||
|
"min": -0.7530547380447388,
|
||||||
|
"p01": -0.7163181900978088,
|
||||||
|
"p10": -0.5367198586463928,
|
||||||
|
"p25": -0.4014696776866913,
|
||||||
|
"p50": -0.1533127874135971,
|
||||||
|
"p75": 0.0,
|
||||||
|
"p90": 0.07427694648504257,
|
||||||
|
"p99": 0.7512762546539307,
|
||||||
|
"max": 2.4661004543304443
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"a_rate_exp": {
|
||||||
|
"count": 128,
|
||||||
|
"mean": 0.9474229216575623,
|
||||||
|
"std": 1.0007015466690063,
|
||||||
|
"quantiles": {
|
||||||
|
"min": 0.47092580795288086,
|
||||||
|
"p01": 0.48882821202278137,
|
||||||
|
"p10": 0.5846699476242065,
|
||||||
|
"p25": 0.6693384051322937,
|
||||||
|
"p50": 0.8578758239746094,
|
||||||
|
"p75": 1.0,
|
||||||
|
"p90": 1.0771057605743408,
|
||||||
|
"p99": 2.1350440979003906,
|
||||||
|
"max": 11.776434898376465
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"dt_bias": {
|
||||||
|
"count": 12288,
|
||||||
|
"mean": -4.57477331161499,
|
||||||
|
"std": 1.423690915107727,
|
||||||
|
"quantiles": {
|
||||||
|
"min": -7.8938117027282715,
|
||||||
|
"p01": -6.962827682495117,
|
||||||
|
"p10": -6.4912495613098145,
|
||||||
|
"p25": -5.797565460205078,
|
||||||
|
"p50": -4.622031211853027,
|
||||||
|
"p75": -3.4063913822174072,
|
||||||
|
"p90": -2.568972587585449,
|
||||||
|
"p99": -1.895374059677124,
|
||||||
|
"max": 0.179231196641922
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"channelwise_hypothesis": {
|
||||||
|
"log_decay": {
|
||||||
|
"count": 12288,
|
||||||
|
"mean": -0.22606194019317627,
|
||||||
|
"std": 0.26349756121635437,
|
||||||
|
"quantiles": {
|
||||||
|
"min": -2.685786724090576,
|
||||||
|
"p01": -1.122031569480896,
|
||||||
|
"p10": -0.6017768383026123,
|
||||||
|
"p25": -0.3287131190299988,
|
||||||
|
"p50": -0.12320166081190109,
|
||||||
|
"p75": -0.035499051213264465,
|
||||||
|
"p90": -0.008835914544761181,
|
||||||
|
"p99": -8.967738722276408e-06,
|
||||||
|
"max": 0.0
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"one_step_retention": {
|
||||||
|
"count": 12288,
|
||||||
|
"mean": 0.821946918964386,
|
||||||
|
"std": 0.17656706273555756,
|
||||||
|
"quantiles": {
|
||||||
|
"min": 0.068167544901371,
|
||||||
|
"p01": 0.32561764121055603,
|
||||||
|
"p10": 0.5478373765945435,
|
||||||
|
"p25": 0.7198495268821716,
|
||||||
|
"p50": 0.8840853571891785,
|
||||||
|
"p75": 0.9651236534118652,
|
||||||
|
"p90": 0.9912030100822449,
|
||||||
|
"p99": 0.9999910593032837,
|
||||||
|
"max": 1.0
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"retention_after_64_steps": {
|
||||||
|
"count": 12288,
|
||||||
|
"mean": 0.13120755553245544,
|
||||||
|
"std": 0.253229558467865,
|
||||||
|
"quantiles": {
|
||||||
|
"min": 0.0,
|
||||||
|
"p01": 6.506081464145866e-32,
|
||||||
|
"p10": 1.8781597374798298e-17,
|
||||||
|
"p25": 7.302480842241721e-10,
|
||||||
|
"p50": 0.00037638185312971473,
|
||||||
|
"p75": 0.10311195999383926,
|
||||||
|
"p90": 0.5680769681930542,
|
||||||
|
"p99": 0.9994263052940369,
|
||||||
|
"max": 1.0
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"router_correction_bias": {
|
||||||
|
"count": 896,
|
||||||
|
"mean": 6.285762310653809e-07,
|
||||||
|
"std": 0.02034943550825119,
|
||||||
|
"quantiles": {
|
||||||
|
"min": -0.08442370593547821,
|
||||||
|
"p01": -0.06806425005197525,
|
||||||
|
"p10": -0.0273849219083786,
|
||||||
|
"p25": -0.009373871609568596,
|
||||||
|
"p50": 0.004928342066705227,
|
||||||
|
"p75": 0.014873139560222626,
|
||||||
|
"p90": 0.020546497777104378,
|
||||||
|
"p99": 0.02667994052171707,
|
||||||
|
"max": 0.028605474159121513
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"router_row_l2": {
|
||||||
|
"count": 896,
|
||||||
|
"mean": 4.949629783630371,
|
||||||
|
"std": 0.895085871219635,
|
||||||
|
"quantiles": {
|
||||||
|
"min": 2.6018569469451904,
|
||||||
|
"p01": 2.9133214950561523,
|
||||||
|
"p10": 3.7345614433288574,
|
||||||
|
"p25": 4.310352325439453,
|
||||||
|
"p50": 5.036390781402588,
|
||||||
|
"p75": 5.6203179359436035,
|
||||||
|
"p90": 6.068215370178223,
|
||||||
|
"p99": 6.65656042098999,
|
||||||
|
"max": 7.023199081420898
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"router_bias_norm_correlation": 0.456669420003891
|
||||||
|
},
|
||||||
|
"router_stress_probe": {
|
||||||
|
"seed": 20260729,
|
||||||
|
"synthetic_tokens": 2048,
|
||||||
|
"hidden_rms": 0.999999463558197,
|
||||||
|
"without_correction_bias": {
|
||||||
|
"mean": 36.57143020629883,
|
||||||
|
"std": 76.2339096069336,
|
||||||
|
"cv": 2.0845208168029785,
|
||||||
|
"gini": 0.8310886025428772,
|
||||||
|
"zero_experts": 558,
|
||||||
|
"quantiles": {
|
||||||
|
"min": 0.0,
|
||||||
|
"p10": 0.0,
|
||||||
|
"p25": 0.0,
|
||||||
|
"p50": 0.0,
|
||||||
|
"p75": 16.0,
|
||||||
|
"p90": 166.0,
|
||||||
|
"p99": 310.0999755859375,
|
||||||
|
"max": 369.0
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"with_correction_bias": {
|
||||||
|
"mean": 36.57143020629883,
|
||||||
|
"std": 92.48487854003906,
|
||||||
|
"cv": 2.528883218765259,
|
||||||
|
"gini": 0.878814697265625,
|
||||||
|
"zero_experts": 673,
|
||||||
|
"quantiles": {
|
||||||
|
"min": 0.0,
|
||||||
|
"p10": 0.0,
|
||||||
|
"p25": 0.0,
|
||||||
|
"p50": 0.0,
|
||||||
|
"p75": 0.0,
|
||||||
|
"p90": 167.0,
|
||||||
|
"p99": 398.0999755859375,
|
||||||
|
"max": 451.0
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"membership_overlap_mean": 2.1708984375,
|
||||||
|
"tokens_changed": 2048,
|
||||||
|
"changed_fraction": 1.0,
|
||||||
|
"mean_replacements_per_token": 13.8291015625
|
||||||
|
},
|
||||||
|
"flashkda": {
|
||||||
|
"supported_architectures": [
|
||||||
|
"90a",
|
||||||
|
"100a",
|
||||||
|
"103a",
|
||||||
|
"120a"
|
||||||
|
],
|
||||||
|
"requirements": {
|
||||||
|
"cuda": ">=12.9",
|
||||||
|
"pytorch": ">=2.4",
|
||||||
|
"gpu": "SM90+"
|
||||||
|
},
|
||||||
|
"official_benchmarks": {
|
||||||
|
"h20": {
|
||||||
|
"sequence": 8192,
|
||||||
|
"heads": 96,
|
||||||
|
"dimension": 128,
|
||||||
|
"rows": {
|
||||||
|
"Fixed": {
|
||||||
|
"flash_kda_ms": 2.622,
|
||||||
|
"fla_chunk_kda_ms": 4.8388,
|
||||||
|
"speedup_vs_chunk_kda": 1.85,
|
||||||
|
"fla_chunk_gdn_ms": 3.1985,
|
||||||
|
"speedup_vs_gdn": 1.22
|
||||||
|
},
|
||||||
|
"Varlen, `seq_lens`=[1300, 547, 2048, 963, 271, 3063]": {
|
||||||
|
"flash_kda_ms": 2.3449,
|
||||||
|
"fla_chunk_kda_ms": 4.8291,
|
||||||
|
"speedup_vs_chunk_kda": 2.06,
|
||||||
|
"fla_chunk_gdn_ms": 3.0541,
|
||||||
|
"speedup_vs_gdn": 1.3
|
||||||
|
},
|
||||||
|
"Varlen, `seq_lens`=`1024 x 8`": {
|
||||||
|
"flash_kda_ms": 2.0432,
|
||||||
|
"fla_chunk_kda_ms": 4.6723,
|
||||||
|
"speedup_vs_chunk_kda": 2.29,
|
||||||
|
"fla_chunk_gdn_ms": 2.9117,
|
||||||
|
"speedup_vs_gdn": 1.43
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"gb200": {
|
||||||
|
"sequence": 8192,
|
||||||
|
"heads": 96,
|
||||||
|
"dimension": 128,
|
||||||
|
"rows": {
|
||||||
|
"Fixed": {
|
||||||
|
"flash_kda_ms": 1.0087,
|
||||||
|
"fla_chunk_kda_ms": 2.3271,
|
||||||
|
"speedup_vs_chunk_kda": 2.31,
|
||||||
|
"fla_chunk_gdn_ms": 1.2792,
|
||||||
|
"speedup_vs_gdn": 1.27
|
||||||
|
},
|
||||||
|
"Varlen, `seq_lens`=[1300, 547, 2048, 963, 271, 3063]": {
|
||||||
|
"flash_kda_ms": 0.8597,
|
||||||
|
"fla_chunk_kda_ms": 2.334,
|
||||||
|
"speedup_vs_chunk_kda": 2.71,
|
||||||
|
"fla_chunk_gdn_ms": 1.2962,
|
||||||
|
"speedup_vs_gdn": 1.51
|
||||||
|
},
|
||||||
|
"Varlen, `seq_lens`=`1024 x 8`": {
|
||||||
|
"flash_kda_ms": 0.7064,
|
||||||
|
"fla_chunk_kda_ms": 2.3105,
|
||||||
|
"speedup_vs_chunk_kda": 3.27,
|
||||||
|
"fla_chunk_gdn_ms": 1.2744,
|
||||||
|
"speedup_vs_gdn": 1.8
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"local_environment": {
|
||||||
|
"python": "3.10.14",
|
||||||
|
"torch": "2.11.0+cu128",
|
||||||
|
"torch_cuda": "12.8",
|
||||||
|
"gpu": "NVIDIA GeForce RTX 5090",
|
||||||
|
"capability": [
|
||||||
|
12,
|
||||||
|
0
|
||||||
|
],
|
||||||
|
"libc": [
|
||||||
|
"glibc",
|
||||||
|
"2.43"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"local_build": {
|
||||||
|
"status": "blocked_before_kernel execution",
|
||||||
|
"attempt_1": "system g++ 15 exceeds CUDA 12.8 host compiler range",
|
||||||
|
"attempt_2": "temporary g++ 13 reaches nvcc, then CUDA 12.8 headers conflict with current glibc math declarations",
|
||||||
|
"interpretation": "GPU architecture is listed by the repository, but the local CUDA 12.8 stack is below the official CUDA 12.9 requirement"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,36 @@
|
|||||||
|
import snapshot from "./k3-artifact-snapshot.json";
|
||||||
|
|
||||||
|
export { snapshot as k3ArtifactSnapshot };
|
||||||
|
|
||||||
|
const mlaLayerSet = new Set(snapshot.configuration.mla_layers);
|
||||||
|
|
||||||
|
export const k3ArtifactLayers = Array.from(
|
||||||
|
{ length: snapshot.configuration.layers },
|
||||||
|
(_, index) => {
|
||||||
|
const layer = index + 1;
|
||||||
|
const attention = mlaLayerSet.has(layer) ? "MLA" : "KDA";
|
||||||
|
return {
|
||||||
|
layer,
|
||||||
|
attention,
|
||||||
|
feedForward: layer === 1 ? "DENSE" : "MOE",
|
||||||
|
block: Math.floor(index / snapshot.configuration.attnres_block) + 1,
|
||||||
|
opensResidualBlock: index % snapshot.configuration.attnres_block === 0,
|
||||||
|
terminalGlobal: layer === snapshot.configuration.layers,
|
||||||
|
};
|
||||||
|
},
|
||||||
|
);
|
||||||
|
|
||||||
|
export const k3ArtifactViews = [
|
||||||
|
["layers", "01", "LAYER MAP", "93 层配置"],
|
||||||
|
["tensors", "02", "TENSOR ANATOMY", "497,220 entries"],
|
||||||
|
["parameters", "03", "PARAMETER AUDIT", "真实权重小切片"],
|
||||||
|
["reproduction", "04", "REPRODUCTION", "作者值与本机边界"],
|
||||||
|
] as const;
|
||||||
|
|
||||||
|
export const k3ArtifactEvidence = [
|
||||||
|
["O", "OFFICIAL ARTIFACT", "官方 config、code、index、header 或参数字节直接观测。"],
|
||||||
|
["D", "DERIVATION", "由公开 shape 与计数做确定性算术。"],
|
||||||
|
["X", "EXECUTED", "本站机器实际执行并留下环境与结果。"],
|
||||||
|
["S", "SYNTHETIC", "真实权重加合成输入;只测试边界,不代表真实 token。"],
|
||||||
|
["U", "UNRESOLVED", "工件之间存在不一致,当前不擅自补解释。"],
|
||||||
|
] as const;
|
||||||
@@ -128,19 +128,19 @@ const paths = [
|
|||||||
<div class="release-grid">
|
<div class="release-grid">
|
||||||
<a class="release-card k3-release" href="/k3/">
|
<a class="release-card k3-release" href="/k3/">
|
||||||
<div>
|
<div>
|
||||||
<p class="eyebrow"><span>NEW / K3 ROUND 02</span> REPORT · FORMULA · SYSTEM · EVIDENCE</p>
|
<p class="eyebrow"><span>NEW / K3 ROUND 03</span> REPORT · CHECKPOINT · KERNEL · BOUNDARY</p>
|
||||||
<h2>47 页不再压成摘要:把 K3 的每个因果环节重新展开</h2>
|
<h2>47 页不再压成摘要:再把 1.56 TB 开放工件接回报告</h2>
|
||||||
<p>
|
<p>
|
||||||
用三十二张问题账逐节读完 KDA、Gated MLA、AttnRes、Stable LatentMoE、原生视觉、
|
在三十二张报告问题账之外,继续审计 96 个 safetensors 分片、497,220 个 tensor entries、
|
||||||
预训练、九专家 MOPD、Agent 环境、FlashKDA / MoonEP、混合 prefix cache、评测与案例边界。
|
真实 KDA / MLA / MoE / MoonViT shape、小范围权重统计、FlashKDA 编译边界与未决形状矛盾。
|
||||||
</p>
|
</p>
|
||||||
</div>
|
</div>
|
||||||
<dl>
|
<dl>
|
||||||
<div><dt>REPORT</dt><dd>16 Figures · 5 Tables</dd></div>
|
<div><dt>REPORT</dt><dd>16 Figures · 5 Tables</dd></div>
|
||||||
<div><dt>NODES</dt><dd>100 个一手 / 官方节点</dd></div>
|
<div><dt>ARTIFACTS</dt><dd>96 shards · 497,220 entries</dd></div>
|
||||||
<div><dt>LAB</dt><dd>Delta · Decay · AttnRes · MoE · QB · RL · Cache</dd></div>
|
<div><dt>LAB</dt><dd>8 报告实验 + 4 工件视图</dd></div>
|
||||||
</dl>
|
</dl>
|
||||||
<span class="release-arrow" aria-hidden="true">从报告目录进入完整因果链 →</span>
|
<span class="release-arrow" aria-hidden="true">从报告目录进入开放工件证据链 →</span>
|
||||||
</a>
|
</a>
|
||||||
<a class="release-card deepseek-release" href="/deepseek/">
|
<a class="release-card deepseek-release" href="/deepseek/">
|
||||||
<div>
|
<div>
|
||||||
|
|||||||
@@ -1,6 +1,7 @@
|
|||||||
---
|
---
|
||||||
import BaseLayout from "@/layouts/BaseLayout.astro";
|
import BaseLayout from "@/layouts/BaseLayout.astro";
|
||||||
import ArchitectureExplorer from "@/components/ArchitectureExplorer.astro";
|
import ArchitectureExplorer from "@/components/ArchitectureExplorer.astro";
|
||||||
|
import K3ArtifactLab from "@/components/K3ArtifactLab.astro";
|
||||||
import K3ReportLab from "@/components/K3ReportLab.astro";
|
import K3ReportLab from "@/components/K3ReportLab.astro";
|
||||||
import { k3FigureAtlas, k3Ledgers, k3PaperChain, k3ReportMap } from "@/data/k3";
|
import { k3FigureAtlas, k3Ledgers, k3PaperChain, k3ReportMap } from "@/data/k3";
|
||||||
|
|
||||||
@@ -34,7 +35,8 @@ const toc = [
|
|||||||
["26", "cases", "案例边界"],
|
["26", "cases", "案例边界"],
|
||||||
["27", "xtml", "XTML 协议"],
|
["27", "xtml", "XTML 协议"],
|
||||||
["28", "lab", "八联交互实验"],
|
["28", "lab", "八联交互实验"],
|
||||||
["29", "audit", "21 张图表审计"],
|
["29", "artifacts", "开放权重工件审计"],
|
||||||
|
["30", "audit", "21 张图表审计"],
|
||||||
["↳", "papers", "100 节点阅读链"],
|
["↳", "papers", "100 节点阅读链"],
|
||||||
];
|
];
|
||||||
|
|
||||||
@@ -103,13 +105,13 @@ const paperGroups = [
|
|||||||
|
|
||||||
<BaseLayout
|
<BaseLayout
|
||||||
title="Kimi K3 技术报告完整深读:架构、训练、RL、系统与评测"
|
title="Kimi K3 技术报告完整深读:架构、训练、RL、系统与评测"
|
||||||
description="用三十二张问题账、二十一张图表审计、八个交互实验与一百个一手阅读节点,逐节读懂 Kimi K3 技术报告。"
|
description="用三十二张问题账、二十一张图表审计、八个机制实验、四个开放工件视图与一百个一手阅读节点,逐节读懂 Kimi K3。"
|
||||||
section="k3"
|
section="k3"
|
||||||
>
|
>
|
||||||
<header class="page-hero k3-hero">
|
<header class="page-hero k3-hero">
|
||||||
<div class="page-hero-inner">
|
<div class="page-hero-inner">
|
||||||
<div>
|
<div>
|
||||||
<p class="eyebrow"><span>ANCHOR REPORT / ROUND 02</span> KIMI K3 · 47 PAGES</p>
|
<p class="eyebrow"><span>ANCHOR REPORT / ROUND 03</span> KIMI K3 · REPORT → OPEN ARTIFACTS</p>
|
||||||
<h1>不把报告压成摘要<br />把每个因果环节<br />重新展开</h1>
|
<h1>不把报告压成摘要<br />把每个因果环节<br />重新展开</h1>
|
||||||
<p class="lead">
|
<p class="lead">
|
||||||
K3 同时扩展序列、深度、宽度、视觉与 Agent 轨迹。真正值得读的不是 2.8T 这个最大数字,
|
K3 同时扩展序列、深度、宽度、视觉与 Agent 轨迹。真正值得读的不是 2.8T 这个最大数字,
|
||||||
@@ -119,10 +121,11 @@ const paperGroups = [
|
|||||||
<dl class="page-facts">
|
<dl class="page-facts">
|
||||||
<div><dt>QUESTIONS</dt><dd>32 张问题账</dd></div>
|
<div><dt>QUESTIONS</dt><dd>32 张问题账</dd></div>
|
||||||
<div><dt>REPORT</dt><dd>16 Figures · 5 Tables</dd></div>
|
<div><dt>REPORT</dt><dd>16 Figures · 5 Tables</dd></div>
|
||||||
<div><dt>LABS</dt><dd>8 个可操作实验</dd></div>
|
<div><dt>LABS</dt><dd>8 个机制实验 + 4 个工件视图</dd></div>
|
||||||
<div><dt>READING</dt><dd>100 个一手 / 官方节点</dd></div>
|
<div><dt>READING</dt><dd>100 个一手 / 官方节点</dd></div>
|
||||||
<div><dt>MODEL</dt><dd>2.78T total / 104.2B active</dd></div>
|
<div><dt>MODEL</dt><dd>2.78T total / 104.2B active</dd></div>
|
||||||
<div><dt>STATUS</dt><dd>K3 二轮深读</dd></div>
|
<div><dt>ARTIFACTS</dt><dd>96 shards · 497,220 tensors</dd></div>
|
||||||
|
<div><dt>STATUS</dt><dd>K3 三轮进行中</dd></div>
|
||||||
</dl>
|
</dl>
|
||||||
</div>
|
</div>
|
||||||
</header>
|
</header>
|
||||||
@@ -137,7 +140,8 @@ const paperGroups = [
|
|||||||
</ol>
|
</ol>
|
||||||
<div class="rail-note">
|
<div class="rail-note">
|
||||||
<b>证据约定</b>
|
<b>证据约定</b>
|
||||||
R = K3 报告;P = 原论文 / 官方实现;D = 确定性推导;T = 教学模型。四者不互相冒充。
|
报告层用 R / P / D / T 区分报告、原始来源、推导与教学模型;
|
||||||
|
工件层用 O / D / X / S / U 区分观测、推导、本站执行、合成探针与未决矛盾。
|
||||||
</div>
|
</div>
|
||||||
</aside>
|
</aside>
|
||||||
|
|
||||||
@@ -844,8 +848,29 @@ const paperGroups = [
|
|||||||
<K3ReportLab />
|
<K3ReportLab />
|
||||||
</section>
|
</section>
|
||||||
|
|
||||||
|
<section class="article-section" id="artifacts">
|
||||||
|
<p class="eyebrow"><span>29</span> OPEN ARTIFACT FORENSICS</p>
|
||||||
|
<h2>从“报告说了什么”走到“公开 checkpoint 实际长什么样”</h2>
|
||||||
|
<p class="lede">
|
||||||
|
第三轮固定到官方 Hugging Face revision,读取 config、remote code、60 MB tensor index、
|
||||||
|
四个 safetensors headers 和两个小范围参数切片。原始权重不进入本站仓库;
|
||||||
|
结构、shape、计数、参数统计与本机编译边界都可以从公开脚本重复生成。
|
||||||
|
</p>
|
||||||
|
<div class="artifact-callout">
|
||||||
|
<article><span>O / OBSERVED</span><b>1.4196 TiB tensor data</b><p>96 shards、497,220 entries;不是运行显存,也不是参数量口径。</p></article>
|
||||||
|
<article><span>D / CLOSED LOOP</span><b>69 KDA · 24 MLA · 92 MoE</b><p>配置、tensor names 与 header shape 三方闭合。</p></article>
|
||||||
|
<article class="warning"><span>U / UNRESOLVED</span><b>A_log [128] ≠ expected [96]</b><p>checkpoint 与公开代码 / kernel API 的形状冲突保留在主视区,不擅自解释。</p></article>
|
||||||
|
</div>
|
||||||
|
<K3ArtifactLab />
|
||||||
|
<div class="hero-actions">
|
||||||
|
<a class="button primary" href="https://huggingface.co/moonshotai/Kimi-K3">打开官方开放权重</a>
|
||||||
|
<a class="button" href="https://github.com/MoonshotAI/FlashKDA">打开 FlashKDA 官方实现</a>
|
||||||
|
<a class="button" href="https://github.com/MoonshotAI/FlashKDA/blob/master/BENCHMARK_GB200.md">核对作者 GB200 benchmark</a>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
<section class="article-section" id="audit">
|
<section class="article-section" id="audit">
|
||||||
<p class="eyebrow"><span>29</span> FIGURE & TABLE AUDIT</p>
|
<p class="eyebrow"><span>30</span> FIGURE & TABLE AUDIT</p>
|
||||||
<h2>Figure 1–16、Table 1–5:每张图究竟支持什么,不能支持什么</h2>
|
<h2>Figure 1–16、Table 1–5:每张图究竟支持什么,不能支持什么</h2>
|
||||||
<div class="figure-atlas">
|
<div class="figure-atlas">
|
||||||
{k3FigureAtlas.map(([id, report, title, contract]) => (
|
{k3FigureAtlas.map(([id, report, title, contract]) => (
|
||||||
@@ -900,6 +925,28 @@ const paperGroups = [
|
|||||||
<style>
|
<style>
|
||||||
.k3-hero { border-bottom-color: var(--copper); }
|
.k3-hero { border-bottom-color: var(--copper); }
|
||||||
.anchor-alias { position: relative; top: -88px; display: block; visibility: hidden; }
|
.anchor-alias { position: relative; top: -88px; display: block; visibility: hidden; }
|
||||||
|
.artifact-callout {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: repeat(3, minmax(0, 1fr));
|
||||||
|
max-width: 1080px;
|
||||||
|
margin: 32px 0;
|
||||||
|
border-top: 1px solid var(--line);
|
||||||
|
border-left: 1px solid var(--line);
|
||||||
|
}
|
||||||
|
.artifact-callout article {
|
||||||
|
min-height: 145px;
|
||||||
|
padding: 20px;
|
||||||
|
border-right: 1px solid var(--line);
|
||||||
|
border-bottom: 1px solid var(--line);
|
||||||
|
background: var(--paper-raised);
|
||||||
|
}
|
||||||
|
.artifact-callout article.warning {
|
||||||
|
border-color: var(--signal);
|
||||||
|
background: color-mix(in srgb, var(--signal) 7%, var(--paper));
|
||||||
|
}
|
||||||
|
.artifact-callout span { color: var(--copper); font: .58rem/1.2 var(--mono); letter-spacing: .08em; }
|
||||||
|
.artifact-callout b { display: block; margin-top: 13px; font: 800 .85rem/1.3 var(--mono); }
|
||||||
|
.artifact-callout p { margin: 12px 0 0; color: var(--muted); font-size: .68rem; line-height: 1.6; }
|
||||||
.ledger-grid {
|
.ledger-grid {
|
||||||
display: grid;
|
display: grid;
|
||||||
grid-template-columns: repeat(3, minmax(0, 1fr));
|
grid-template-columns: repeat(3, minmax(0, 1fr));
|
||||||
@@ -1200,7 +1247,7 @@ const paperGroups = [
|
|||||||
.xtml-stage > i:nth-of-type(n+3) { display: none; }
|
.xtml-stage > i:nth-of-type(n+3) { display: none; }
|
||||||
}
|
}
|
||||||
@media (max-width: 720px) {
|
@media (max-width: 720px) {
|
||||||
.ledger-grid, .axis-grid, .domain-grid, .agentenv-grid, .serving-grid, .eval-axes, .case-grid,
|
.artifact-callout, .ledger-grid, .axis-grid, .domain-grid, .agentenv-grid, .serving-grid, .eval-axes, .case-grid,
|
||||||
.mechanism-steps, .protocol-grid, .comparison-grid, .layer-separation, .precision-contract,
|
.mechanism-steps, .protocol-grid, .comparison-grid, .layer-separation, .precision-contract,
|
||||||
.two-column, .environment-list, .figure-atlas, .audit-legend, .harness-parts {
|
.two-column, .environment-list, .figure-atlas, .audit-legend, .harness-parts {
|
||||||
grid-template-columns: 1fr;
|
grid-template-columns: 1fr;
|
||||||
|
|||||||
@@ -9,7 +9,7 @@ const researching = chapters.filter((chapter) => ["researching", "drafting"].inc
|
|||||||
const workstreams = [
|
const workstreams = [
|
||||||
{ label: "研究框架与规范", value: 83, next: "给 Scaling 与推理专题补逐篇图表/实验精读层级" },
|
{ label: "研究框架与规范", value: 83, next: "给 Scaling 与推理专题补逐篇图表/实验精读层级" },
|
||||||
{ label: "网站设计系统", value: 89, next: "打印样式与更多通用可视化组件" },
|
{ label: "网站设计系统", value: 89, next: "打印样式与更多通用可视化组件" },
|
||||||
{ label: "Kimi K3 深读", value: 88, next: "第三轮加入官方权重 traces、独立复现与逐图数值重绘" },
|
{ label: "Kimi K3 深读", value: 92, next: "在匹配 CUDA 12.9+ 环境执行 FlashKDA,并接入真实 hidden-state / expert-load traces" },
|
||||||
{ label: "语言模型前史", value: 78, next: "逐图精读 Kneser–Ney、LSTM 与 Bahdanau,并加入真实小语料复现" },
|
{ label: "语言模型前史", value: 78, next: "逐图精读 Kneser–Ney、LSTM 与 Bahdanau,并加入真实小语料复现" },
|
||||||
{ label: "Transformer 基础", value: 79, next: "逐图精读多头电路、Pre/Post-LN 与真实 kernel / KV 配置" },
|
{ label: "Transformer 基础", value: 79, next: "逐图精读多头电路、Pre/Post-LN 与真实 kernel / KV 配置" },
|
||||||
{ label: "表示、位置与残差高速公路", value: 81, next: "加入真实 hidden-state / norm traces、长上下文位置外推复现与更多深层稳定性消融" },
|
{ label: "表示、位置与残差高速公路", value: 81, next: "加入真实 hidden-state / norm traces、长上下文位置外推复现与更多深层稳定性消融" },
|
||||||
@@ -50,7 +50,7 @@ const workstreams = [
|
|||||||
<div><dt>OVERALL</dt><dd>专题平均 {average}%</dd></div>
|
<div><dt>OVERALL</dt><dd>专题平均 {average}%</dd></div>
|
||||||
<div><dt>READABLE</dt><dd>{published} 个首版可读专题</dd></div>
|
<div><dt>READABLE</dt><dd>{published} 个首版可读专题</dd></div>
|
||||||
<div><dt>ACTIVE</dt><dd>{researching} 个研究/写作中</dd></div>
|
<div><dt>ACTIVE</dt><dd>{researching} 个研究/写作中</dd></div>
|
||||||
<div><dt>UPDATED</dt><dd>2026-07-29 11:54 CST</dd></div>
|
<div><dt>UPDATED</dt><dd>2026-07-29 12:40 CST</dd></div>
|
||||||
<div><dt>MODE</dt><dd>持续迭代,不锁死版本</dd></div>
|
<div><dt>MODE</dt><dd>持续迭代,不锁死版本</dd></div>
|
||||||
</dl>
|
</dl>
|
||||||
</div>
|
</div>
|
||||||
@@ -97,13 +97,14 @@ const workstreams = [
|
|||||||
<article><span>✓</span><h3>K3 报告已结构化拆解</h3><p>47 页报告目录、151 条参考来源和架构/后训练/系统主线已经提取。</p></article>
|
<article><span>✓</span><h3>K3 报告已结构化拆解</h3><p>47 页报告目录、151 条参考来源和架构/后训练/系统主线已经提取。</p></article>
|
||||||
<article><span>✓</span><h3>17 专题知识图</h3><p>从语言模型基础到评测安全,包含先修依赖和三条贯穿案例。</p></article>
|
<article><span>✓</span><h3>17 专题知识图</h3><p>从语言模型基础到评测安全,包含先修依赖和三条贯穿案例。</p></article>
|
||||||
<article><span>✓</span><h3>编辑式网站系统</h3><p>响应式导航、章节模板、侧栏、进度、论文链和证据提示组件。</p></article>
|
<article><span>✓</span><h3>编辑式网站系统</h3><p>响应式导航、章节模板、侧栏、进度、论文链和证据提示组件。</p></article>
|
||||||
<article><span>✓</span><h3>六十七个原创交互视图</h3><p>K3 三轴图与八联实验、语言模型前史、Transformer、表示深度、DeepSeek 四联实验、长上下文、MoE、推理、Agent、多模态,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
|
<article><span>✓</span><h3>七十一个原创交互视图</h3><p>K3 三轴图、八联报告实验与四联开放工件实验,以及语言模型前史、Transformer、表示深度、DeepSeek、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
|
||||||
<article><span>✓</span><h3>十七篇首版长文</h3><p>K3、语言模型前史、Transformer、表示/位置/残差、DeepSeek、Scaling、数据工程、长上下文、MoE、后训练、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全专题。</p></article>
|
<article><span>✓</span><h3>十七篇首版长文</h3><p>K3、语言模型前史、Transformer、表示/位置/残差、DeepSeek、Scaling、数据工程、长上下文、MoE、后训练、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全专题。</p></article>
|
||||||
<article><span>✓</span><h3>语言模型前史深度专题</h3><p>八张独立问题账、33 个正式节点、20 段长文与概率—向量—记忆—对齐四联实验。</p></article>
|
<article><span>✓</span><h3>语言模型前史深度专题</h3><p>八张独立问题账、33 个正式节点、20 段长文与概率—向量—记忆—对齐四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>Transformer 深度专题</h3><p>十张独立问题账、40 个正式节点、21 段正文与 QKV—Mask—多头位置—Block 成本四联实验。</p></article>
|
<article><span>✓</span><h3>Transformer 深度专题</h3><p>十张独立问题账、40 个正式节点、21 段正文与 QKV—Mask—多头位置—Block 成本四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>表示、位置与残差高速公路深度专题</h3><p>二十张问题账、66 个一手节点、DeepSeek/Kimi 双谱系,以及 Token—位置—Norm—Residual/FFN 四联实验。</p></article>
|
<article><span>✓</span><h3>表示、位置与残差高速公路深度专题</h3><p>二十张问题账、66 个一手节点、DeepSeek/Kimi 双谱系,以及 Token—位置—Norm—Residual/FFN 四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>DeepSeek 技术谱系二轮深读</h3><p>二十四张问题账、十次技术转向、60 个一手/官方节点,以及稀疏容量—MLA 缓存—V3 协同—RL 偏差四联实验。</p></article>
|
<article><span>✓</span><h3>DeepSeek 技术谱系二轮深读</h3><p>二十四张问题账、十次技术转向、60 个一手/官方节点,以及稀疏容量—MLA 缓存—V3 协同—RL 偏差四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>Kimi K3 技术报告二轮深读</h3><p>三十二张问题账、Figure 1–16 / Table 1–5 审计、100 节点阅读链,以及 Delta—Decay—AttnRes—LatentMoE—SiTU—QB—MOPD—Cache 八联实验。</p></article>
|
<article><span>✓</span><h3>Kimi K3 技术报告二轮深读</h3><p>三十二张问题账、Figure 1–16 / Table 1–5 审计、100 节点阅读链,以及 Delta—Decay—AttnRes—LatentMoE—SiTU—QB—MOPD—Cache 八联实验。</p></article>
|
||||||
|
<article><span>✓</span><h3>Kimi K3 三轮开放工件里程碑</h3><p>固定官方 revisions,审计 96 个 shards、497,220 个 tensor entries 与真实 KDA / MLA / MoE / MoonViT shapes;四联实验分开显示层型、tensor anatomy、参数范围和复现边界。</p></article>
|
||||||
<article><span>✓</span><h3>Scaling Laws 深度专题</h3><p>九张账、29 个一手节点、DeepSeek/Kimi 双谱系与曲面—部署—复用—涌现四联实验。</p></article>
|
<article><span>✓</span><h3>Scaling Laws 深度专题</h3><p>九张账、29 个一手节点、DeepSeek/Kimi 双谱系与曲面—部署—复用—涌现四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>数据工程深度专题</h3><p>十二张账、31 个一手节点、DeepSeek/Kimi 双谱系与流水线—去重—混合—改写四联实验。</p></article>
|
<article><span>✓</span><h3>数据工程深度专题</h3><p>十二张账、31 个一手节点、DeepSeek/Kimi 双谱系与流水线—去重—混合—改写四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>长上下文深度专题</h3><p>五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。</p></article>
|
<article><span>✓</span><h3>长上下文深度专题</h3><p>五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。</p></article>
|
||||||
@@ -209,6 +210,8 @@ const workstreams = [
|
|||||||
<div><time>2026-07-29</time><b>K3 二轮按三十二张账重建</b><p>从架构组件摘要升级为覆盖预训练、后训练、环境、系统、评测、案例与附录的完整报告因果链。</p></div>
|
<div><time>2026-07-29</time><b>K3 二轮按三十二张账重建</b><p>从架构组件摘要升级为覆盖预训练、后训练、环境、系统、评测、案例与附录的完整报告因果链。</p></div>
|
||||||
<div><time>2026-07-29</time><b>K3 原生视觉事实纠错</b><p>MoonViT-V2 从头训练;视觉与文本从训练开始在同一个 NTP objective 中联合优化,不再沿用冻结/解冻式 post-hoc 叙述。</p></div>
|
<div><time>2026-07-29</time><b>K3 原生视觉事实纠错</b><p>MoonViT-V2 从头训练;视觉与文本从训练开始在同一个 NTP objective 中联合优化,不再沿用冻结/解冻式 post-hoc 叙述。</p></div>
|
||||||
<div><time>2026-07-29</time><b>K3 图表与实验永久分级</b><p>Figure 1–16 / Table 1–5 建立视觉契约;报告事实、原论文、确定性推导与教学模型使用 R/P/D/T 四种身份。</p></div>
|
<div><time>2026-07-29</time><b>K3 图表与实验永久分级</b><p>Figure 1–16 / Table 1–5 建立视觉契约;报告事实、原论文、确定性推导与教学模型使用 R/P/D/T 四种身份。</p></div>
|
||||||
|
<div><time>2026-07-29</time><b>K3 开放工件按五种证据身份审计</b><p>真实观测 O、确定性推导 D、本机执行 X、合成探针 S 与未决矛盾 U 分开;作者 benchmark 不冒充本站实测。</p></div>
|
||||||
|
<div><time>2026-07-29</time><b>A_log 形状冲突保持未决</b><p>checkpoint 的 [128] 与 config / remote code / FlashKDA API 期待的 [96] 并列展示;不宣布权重损坏,也不把 channel-wise 假设写成真实 forward。</p></div>
|
||||||
</div>
|
</div>
|
||||||
</section>
|
</section>
|
||||||
|
|
||||||
|
|||||||
@@ -105,9 +105,9 @@ const stages = [
|
|||||||
<p>K3 架构 → 07/06/03 → 05/04 → 11/12 → 08/14/15</p>
|
<p>K3 架构 → 07/06/03 → 05/04 → 11/12 → 08/14/15</p>
|
||||||
<ol>
|
<ol>
|
||||||
<li>沿 32 张问题账解释三维信息流</li>
|
<li>沿 32 张问题账解释三维信息流</li>
|
||||||
<li>用 8 个实验比较 KDA、MLA、AttnRes 与 LatentMoE</li>
|
<li>用 8 个报告实验与 4 个工件视图比较机制和真实 shape</li>
|
||||||
<li>分清 2.78T / 104.2B、2.5× 与 1M 的证据口径</li>
|
<li>分清 2.78T / 104.2B、2.5× 与 1M 的证据口径</li>
|
||||||
<li>读懂九专家 MOPD、AgentENV、混合缓存与评测协议</li>
|
<li>读懂九专家 MOPD、混合缓存、FlashKDA 与复现边界</li>
|
||||||
</ol>
|
</ol>
|
||||||
</article>
|
</article>
|
||||||
<article>
|
<article>
|
||||||
|
|||||||
Reference in New Issue
Block a user