feat: isolate DeepSeek history boundary token

This commit is contained in:
wuyang
2026-07-29 20:00:37 +08:00
parent f5a69170c7
commit f4d4a05dfa
14 changed files with 3676654 additions and 33 deletions
+12 -3
View File
@@ -14,7 +14,7 @@
| 表示、位置与残差高速公路 | 完成首版 | 81% | 真实 hidden-state / norm traces、长上下文位置外推与深层稳定性消融 |
| Scaling Laws | 完成首版 | 74% | 真实拟合复现、置信区间与更多模型族对照 |
| 数据工程与预训练配方 | 完成首版 | 73% | FineWeb / DCLM 逐图精读、真实去重误伤与 mixture traces |
| DeepSeek 专题 | 三轮实证进行中 | 95% | SM90 FlashMLA kernel、完整 27 层、EOS / 角色 / 多 filler / 内容正交控制、FP8/pipeline 与 R1-like RL 复现 |
| DeepSeek 专题 | 三轮实证进行中 | 96% | 角色标记与 special-token family、V2-Lite-Chat 行为、完整 27 层与固定 batch-shape 对照,再推进 SM90 FlashMLA、FP8/pipeline 与 R1-like RL 复现 |
| 指令微调与人类偏好 | 完成首版 | 75% | 真实偏好分歧、RM 长度偏置与 PPO/DPO 小模型复现 |
| 推理与测试时扩展 | 完成首版 | 76% | 真实模型采样曲线、PRM 案例与逐篇图表精读 |
| 工具使用与长程 Agent | 完成首版 | 74% | 真实环境 traces、cross-harness 对照、Agent RL 曲线与安全案例 |
@@ -41,7 +41,7 @@
- [x] 完成 486 篇关键论文索引,覆盖 16 个标签专题与 Kimi/DeepSeek 聚光主线。
- [x] 完成可检索、可按专题筛选的论文库页面。
- [x] 完成 K3、语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全十七篇首版长文。
- [x] 完成 K3 三轴架构、八联报告实验与四联开放工件实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 十三联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等八十个原创交互视图。
- [x] 完成 K3 三轴架构、八联报告实验与四联开放工件实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 十四联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等八十一个原创交互视图。
- [x] 完成长上下文首版:五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。
- [x] 核验 FlashAttention、DeepSeek-V2/V3.2/V4、Kimi Linear/K3 等六份论文原文,并建立长上下文研究账本。
- [x] 核验 Switch、ST-MoE、DeepSeekMoE、Loss-Free、V3、LatentMoE 与 K3 原文,并建立 MoE 研究账本。
@@ -210,11 +210,17 @@
- [x] 六格正式运行与独立复跑 SHA-256 均为 `423a095d…e648e`,约 41 MiB JSON byte-exact;网站第九个真实工件页签联动 layer、scope、aggregation 与两个 TV 台阶 depth map。
- [x] 等长历史控制本地闸门通过:69 个 Astro 文件零诊断、21 个页面、1,151 个站内引用、12 个跨页锚点零失败,十六套真实 Chrome 回归全部通过,9 页签桌面端与 390px 移动端均无文档级横向溢出。
- [x] DeepSeek 等长历史控制里程碑以源提交 `8130e37`、不可变镜像 `20260729T110301Z-8130e37` 发布;OCI digest `sha256:e4b79166…5c0846`,NAS / VPS / NPM / DNS / TLS / HTTP2 / gzip / 门户与十六套生产 Chrome 回归全通过;保留 `20260729T100729Z-efe3ffc` 回滚。
- [x] 历史边界单 token 控制固定官方 V2-Lite template:在 system 0/1 × EOS/x/句点/换行八格中保留重复词元历史、`Assistant:` / `User:`、长度、目标绝对位置、attention mask 与同一 32-row batch,只将 assistant EOS `100001` 替换为一个普通 token ID。
- [x] RTX 5090 执行 1,024 个输入变体、56,784 个输入 token,新增 2,044,224 次真实 top-6 路由,使公开语料累计达到 5,157,072 次;八格各有相同的 2,874 个精确对齐目标 token,768 / 768 个反事实序列恰好只改一个 input ID。
- [x] 目标内容的 system-edge TV 均值为 EOS `.0374`、x `.0539`、句点 `.0553`、换行 `.0492`;x / 句点在 24 / 24 格点估计高于 EOS、23 / 24 配对区间完全高于零,换行为 23 / 24 与 16 / 24。结论只命名为“EOS token identity 在固定协议中的路由效应”,不升级为回合理解、能力或 Chat 模型行为。
- [x] 直接 EOS↔control 的 S1−S0 interaction 与 ΔCV 均跨层跨域混合;完整输入差异也远弱于目标内容主结果。网站明确保留 `User:` role marker、base ≠ Chat/SFT、反事实非官方合法 chat、没有生成/准确率指标等边界。
- [x] 八格正式运行与独立复跑各 54,254,445 bytes,SHA-256 均为 `9bb93834…b9c37` 且 byte-exact;确定性 compact 生成器把前端载荷降至约 1.2 MiB,同时在产物内钉住两份完整来源 hash。
- [x] 历史边界控制本地闸门通过:70 个 Astro 文件零诊断、21 个页面、1,151 个站内引用、12 个跨页锚点零失败;十页签桌面与 390px 移动端浏览器回归通过,无运行时异常或文档级横向溢出。
## 正在进行
- [ ] K3 三轮下一闸门:获得真实 token hidden states、expert load 与 cache traces,解释或修订 `A_log [128]` 工件冲突,再做 Figure 3/4/5 数值重绘和独立小模型复现。
- [ ] DeepSeek 三轮下一闸门:在官方支持的 SM90 环境执行 FlashMLA 优化 kernel;扩到完整 27 层并继续拆分 EOS、角色、多 filler、示例内容与 batch shape,再推进 FP8 / pipeline traces 与 R1-like RL 小模型复现。
- [ ] DeepSeek 三轮下一闸门:拆分 `User:` / `Assistant:` 角色标记与 special-token family,加入 V2-Lite-Chat 生成/行为对照;扩到完整 27 层并固定 batch shape 审计,再推进 SM90 FlashMLA、FP8 / pipeline traces 与 R1-like RL 小模型复现。
- [ ] 表示、位置与残差二轮:真实 hidden-state / norm traces、长上下文位置外推复现与 mHC / AttnRes 深层稳定性消融。
- [ ] 评测安全二轮:真实 cross-harness / pass@k 复跑、Judge 元评测、动态污染与过拒案例。
- [ ] 推理服务二轮:真实 GPU kernel / workload traces、功耗与成本、跨 vLLM / SGLang / TensorRT-LLM 复现。
@@ -357,6 +363,9 @@
| 2026-07-29 | 等长历史的两个 TV 台阶分开报告 | none→filler 回答“历史结构是否足以复现缓冲”;filler→demo 回答固定协议字段后的文本替换;两者都不升级为能力或示例正确性 |
| 2026-07-29 | batch contract 是 BF16 路由复现的一部分 | 跨实验 512/512 token-ID 合同 exact,但矩阵形状改变会让深层临界 gate hash 分化;正式结论只使用同一次六格 batch 内对比 |
| 2026-07-29 | DeepSeek 等长历史控制以 `20260729T110301Z-8130e37` 发布 | OCI digest `sha256:e4b79166…5c0846`;复用 NAS 12010→8080、NPM 31 / cert 41、门户 order 180;十六套生产 Chrome 回归通过,保留上一不可变镜像回滚 |
| 2026-07-29 | 历史 assistant 边界使用单 token ID 替换,不删除 token | EOS→x/句点/换行逐条保持长度、目标位置、角色标记、mask 与 batch;识别的是一个输入 ID 的干预,不是假装移除了全部回合边界 |
| 2026-07-29 | 边界控制主结论只使用精确对齐目标内容 TV | x/句点/换行相对 EOS 的 system edge 更大,但 direct interaction、CV 与完整输入方向更混合;不命名为路由更优、回合理解或能力变化 |
| 2026-07-29 | base checkpoint 与 Chat 行为永久分证据层 | 当前 V2-Lite base 虽使用官方 tokenizer/template 构造协议,却没有 Chat/SFT 行为身份;后续必须另跑 Chat checkpoint 与生成/任务指标 |
| 2026-07-29 | K3 二轮按 32 张对象账与完整报告顺序重建 | total/active、2.5×、KDA state、深度来源、专家路由、视觉目标、轨迹、缓存与评测协议不再压成一页组件摘要 |
| 2026-07-29 | K3 原生视觉事实回到 §2.4 / §3.3 核验 | 删除“先冻结语言模型再解冻”旧表述;明确 MoonViT-V2 从头训练,视觉/文本从开始共同 NTP |
| 2026-07-29 | K3 Figure 1–16 / Table 1–5 全部建立课程视觉契约 | 每张图同时写支持范围与不可外推项;作者报告、论文、推导与 toy model 使用 R/P/D/T 标签 |
+10 -4
View File
@@ -19,7 +19,7 @@
当前里程碑包含 17 专题学习地图、486 篇关键论文索引、Kimi K3 完整导读,
语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 技术谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、工具使用与长程 Agent、原生多模态、训练系统、推理服务、数值优化,以及评测与安全深度专题,
以及 80 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、
以及 81 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、
8 个交互实验和 100 个一手/官方节点,完整覆盖架构、预训练、后训练、系统、评测、案例与附录。
第三轮已完成开放工件与首个真实 kernel 里程碑:固定官方模型与 FlashKDA revisions,审计 96 个 checkpoint shards、
497,220 个 tensor entries、真实 KDA / MLA / MoE / MoonViT shapes 与小范围参数统计,并用 4 个新视图
@@ -28,7 +28,7 @@
[K3_ARTIFACT_AUDIT.md](./research/K3_ARTIFACT_AUDIT.md) 与
[checkpoint_probe.py](./experiments/k3/checkpoint_probe.py)、[FlashKDA probe](./experiments/k3/flashkda/)。
DeepSeek 三轮专题以 24 张问题账、10 次技术转向、
13 个交互实验和 60 个一手/官方节点,串起 Dense、MoE、MLA、V3 协同、R1 与 V4;
14 个交互实验和 60 个一手/官方节点,串起 Dense、MoE、MLA、V3 协同、R1 与 V4;
并固定官方 V2-Lite revision,在 RTX 5090 上连续执行 7/27 层,记录 3,240 次真实专家选择、
MLA/HF eager cache shapes 与 `31/31` exact 独立复跑;进一步用真实 layer-1 权重执行官方 V3
naive/absorb 路径,实际写入 576 元素 latent cache,并以 FP32 将两种结合顺序的最大误差压到
@@ -46,7 +46,12 @@ TV 在 24 / 24 个 layer×domain 中都下降,均值从 `0.073` 降至 `0.019`
越界解释为示例语义或能力提升。最新一步再加入与原 one-shot 精确等长的重复词元 filler,
新增 1,376,496 次真实路由:system-edge TV 呈 `0.0738→0.0378→0.0188`
的 none→filler→demo 两级阶梯,两个台阶都在 24 / 24 格下降且 paired 区间完全低于零;
但 filler 仍是学习过的 token,不能冒充纯距离因果。当前累计 3,112,848 次公开语料路由。FlashMLA 的
但 filler 仍是学习过的 token,不能冒充纯距离因果。随后在固定 filler 历史中只把 assistant
后的官方 EOS ID 替换为 `x`、句点或换行,新增 2,044,224 次真实路由;三种替换逐条同长度、
同目标位置且相对官方序列恰好只改一个 ID。目标内容的平均 system-edge TV 为
`.0374 / .0539 / .0553 / .0492`(EOS / x / 句点 / 换行);X 与句点在 24 / 24 格高于
EOS,换行为 23 / 24,但 base checkpoint、非法反事实序列和无行为指标的边界被明确保留。
当前累计 5,157,072 次公开语料路由。FlashMLA 的
SM90/SM100 官方支持矩阵与本机 SM120 边界单独记账。详见
[DEEPSEEK_V2_LITE_TRACE.md](./research/DEEPSEEK_V2_LITE_TRACE.md) 与
[DEEPSEEK_MLA_ABSORB_AUDIT.md](./research/DEEPSEEK_MLA_ABSORB_AUDIT.md)、
@@ -54,7 +59,8 @@ SM90/SM100 官方支持矩阵与本机 SM120 边界单独记账。详见
[DEEPSEEK_ROUTING_LENGTH_CONTROL_AUDIT.md](./research/DEEPSEEK_ROUTING_LENGTH_CONTROL_AUDIT.md) 与
[DEEPSEEK_ROUTING_TEMPLATE_AUDIT.md](./research/DEEPSEEK_ROUTING_TEMPLATE_AUDIT.md)、
[DEEPSEEK_ROUTING_HISTORY_FACTORIAL_AUDIT.md](./research/DEEPSEEK_ROUTING_HISTORY_FACTORIAL_AUDIT.md)、
[DEEPSEEK_ROUTING_HISTORY_DISTANCE_CONTROL_AUDIT.md](./research/DEEPSEEK_ROUTING_HISTORY_DISTANCE_CONTROL_AUDIT.md)。
[DEEPSEEK_ROUTING_HISTORY_DISTANCE_CONTROL_AUDIT.md](./research/DEEPSEEK_ROUTING_HISTORY_DISTANCE_CONTROL_AUDIT.md) 与
[DEEPSEEK_ROUTING_HISTORY_BOUNDARY_TOKEN_AUDIT.md](./research/DEEPSEEK_ROUTING_HISTORY_BOUNDARY_TOKEN_AUDIT.md)。
其余专题按进度账本持续扩建。
## 本地开发
+52
View File
@@ -319,3 +319,55 @@ in 24/24 layer×domain cells, with all 24 paired intervals below zero. See
`research/DEEPSEEK_ROUTING_HISTORY_DISTANCE_CONTROL_AUDIT.md` for the full
table, token-level route stability, BF16 batch-shape boundary, literature
context, and non-claims.
## History-boundary single-token control
`v2_lite_routing_history_boundary_token_control.py` keeps the repeated-token
history from the preceding probe and changes exactly one token ID at the
completed assistant boundary:
```text
system off/on × official EOS / x / period / newline
```
The official template places EOS between the filler assistant content and the
next `User:` marker. The three controls replace only that EOS ID after official
tokenization. They are explicit counterfactual token sequences, not valid
official chat serializations. All eight conditions preserve sequence length,
target position, role markers, attention mask, and within-run batch shape.
```bash
PYTHONPATH=/path/to/transformers-4.41.2-deps:/usr/lib/python3/dist-packages \
python -B experiments/deepseek/v2_lite_routing_history_boundary_token_control.py \
--artifact-dir /path/to/deepseek-v2-lite \
--human-eval /path/to/HumanEval.jsonl.gz \
--gsm8k /path/to/gsm8k/test.jsonl \
--tnews /path/to/tnews/test.json \
--tnews-archive /path/to/tnews_public.zip \
--wikitext /path/to/wikitext-validation.parquet \
--output src/data/deepseek-v2-lite-routing-history-boundary-token-control.json \
--per-domain 32 \
--content-tokens 23 \
--batch-prompts 4 \
--layers 7 \
--bootstrap 2000 \
--seed 20260729 \
--captured-at 2026-07-29T11:12:00+00:00
```
The eight cells add 2,044,224 real top-6 route selections. All 256
source×system groups are equal-length and equal-position; all 768
counterfactual cells differ from their official sequence at exactly one input
ID. The committed run and independent rerun are byte-exact:
```text
9bb93834ffd8536aeebe4325e45d2179ba590554c6ff6b6fcceaba2f499b9c37
```
For exact target content under prompt-balanced aggregation, mean system-edge
TV is `.0374 / .0539 / .0553 / .0492` for EOS / x / period / newline. X and
period exceed EOS in 24/24 layer×domain cells, while newline does so in 23/24.
See `research/DEEPSEEK_ROUTING_HISTORY_BOUNDARY_TOKEN_AUDIT.md` for all paired
intervals, per-token route alignment, scope split, BF16 batch-shape audit,
primary literature, and the boundary between an input-ID intervention and a
chat-turn semantic claim.
@@ -0,0 +1,774 @@
#!/usr/bin/env python3
"""Run a 2 x 4 history-boundary token control on DeepSeek-V2-Lite.
Every condition contains the same repeated-token user/assistant history and
the same target user content. The system factor is off/on. The second factor
changes exactly one token ID at the completed assistant-turn boundary:
eos: the official-template EOS token
x: ordinary one-token content control
period: ordinary one-token punctuation control
newline: ordinary one-token formatting control
The three non-EOS conditions are explicit token-ID counterfactuals rather than
official serialized chats. They preserve sequence length, target position,
role-marker tokens, attention mask, and padded batch shape. This identifies
the routing effect of replacing the learned EOS ID at this one boundary; it
does not identify the effect of removing all turn-boundary information.
The official BF16 forward path is reused from the audited message-history
runner. This file owns the eight-cell renderer, shared-bootstrap statistics,
counterfactual contract, and final result schema.
"""
from __future__ import annotations
import contextlib
import hashlib
import json
import os
import sys
from pathlib import Path
from typing import Any
import numpy as np
import v2_lite_routing_history_distance_control as prior
base = prior.base
SYSTEM_MESSAGE = prior.SYSTEM_MESSAGE
FILLER_USER = prior.FILLER_USER
FILLER_ASSISTANT = prior.FILLER_ASSISTANT
BOUNDARY_LEVELS = ("eos", "x", "period", "newline")
BOUNDARY_TEXT = {
"x": "x",
"period": ".",
"newline": "\n",
}
CONDITIONS = tuple(
f"s{system}_{boundary}"
for boundary in BOUNDARY_LEVELS
for system in (0, 1)
)
FACTORS = {
condition: {
"system": int(condition[1]),
"history": "filler",
"boundary": condition.split("_", 1)[1],
}
for condition in CONDITIONS
}
SYSTEM_CELLS = {
boundary: (f"s0_{boundary}", f"s1_{boundary}")
for boundary in BOUNDARY_LEVELS
}
SYSTEM_EDGE_CONTRASTS = {
f"{boundary}_minus_eos": ("eos", boundary)
for boundary in BOUNDARY_LEVELS
if boundary != "eos"
}
COMPARISONS = tuple(
[
(
f"system_{boundary}",
f"s0_{boundary}",
f"s1_{boundary}",
)
for boundary in BOUNDARY_LEVELS
]
+ [
(
f"{boundary}_at_s{system}",
f"s{system}_eos",
f"s{system}_{boundary}",
)
for boundary in BOUNDARY_LEVELS
if boundary != "eos"
for system in (0, 1)
]
)
ALIGNMENT_COMPARISONS = COMPARISONS
RENDER_AUDIT: list[dict[str, Any]] = []
BOUNDARY_TOKEN_IDS: dict[str, int] = {}
def condition_messages(
content: str,
condition: str,
) -> list[dict[str, str]]:
factors = FACTORS[condition]
messages: list[dict[str, str]] = []
if factors["system"]:
messages.append({"role": "system", "content": SYSTEM_MESSAGE})
messages.extend(
[
{"role": "user", "content": FILLER_USER},
{"role": "assistant", "content": FILLER_ASSISTANT},
{"role": "user", "content": content},
]
)
return messages
def boundary_token_ids(tokenizer: Any) -> dict[str, int]:
if tokenizer.eos_token_id is None:
raise RuntimeError("tokenizer has no EOS token")
ids = {"eos": int(tokenizer.eos_token_id)}
for name, text in BOUNDARY_TEXT.items():
encoded = list(
tokenizer(text, add_special_tokens=False).input_ids
)
if len(encoded) != 1:
raise RuntimeError(
f"{name} boundary control is not one token: {encoded}"
)
if encoded[0] in tokenizer.all_special_ids:
raise RuntimeError(
f"{name} boundary control unexpectedly uses a special token"
)
ids[name] = int(encoded[0])
if len(set(ids.values())) != len(ids):
raise RuntimeError(f"boundary token IDs are not distinct: {ids}")
return ids
def render_boundary_variant(
tokenizer: Any,
content: str,
condition: str,
) -> dict[str, Any]:
messages = condition_messages(content, condition)
rendered = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
official_ids = list(
tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
)
)
token_ids, offsets = base.tokenize_with_offsets(
tokenizer,
rendered,
add_special_tokens=False,
)
if token_ids != official_ids:
raise RuntimeError(
f"{condition} rendered token IDs differ from apply_chat_template"
)
ids = boundary_token_ids(tokenizer)
BOUNDARY_TOKEN_IDS.update(ids)
boundary_positions = [
index
for index, token_id in enumerate(official_ids)
if token_id == ids["eos"]
]
if len(boundary_positions) != 1:
raise RuntimeError(
f"{condition} expected one history EOS; got {boundary_positions}"
)
boundary_position = boundary_positions[0]
boundary = FACTORS[condition]["boundary"]
token_ids = list(token_ids)
token_ids[boundary_position] = ids[boundary]
differing_ids = sum(
left != right
for left, right in zip(official_ids, token_ids, strict=True)
)
expected_differences = 0 if boundary == "eos" else 1
if differing_ids != expected_differences:
raise RuntimeError(
f"{condition} changed {differing_ids} IDs; "
f"expected {expected_differences}"
)
content_start = rendered.rfind(content)
if content_start < 0:
raise RuntimeError(
f"{condition} target content is absent after rendering"
)
content_end = content_start + len(content)
positions, records, crossing = base.content_positions(
token_ids,
offsets,
content_start,
content_end,
)
if not positions:
raise RuntimeError(f"{condition} has no target-content tokens")
if boundary_position >= min(positions):
raise RuntimeError(
f"{condition} history boundary is not before target content"
)
decoded = tokenizer.decode(
token_ids,
skip_special_tokens=False,
clean_up_tokenization_spaces=False,
)
RENDER_AUDIT.append(
{
"content_sha256": base.text_sha256(content),
"condition": condition,
"boundary": boundary,
"boundary_position": boundary_position,
"boundary_token_id": ids[boundary],
"official_eos_token_id": ids["eos"],
"input_tokens": len(token_ids),
"target_first_position": min(positions),
"changed_token_ids_vs_official": differing_ids,
"counterfactual_decoded_sha256": base.text_sha256(decoded),
}
)
return {
"condition": condition,
"messages_sha256": base.canonical_hash(messages),
"rendered_sha256": base.text_sha256(rendered),
"token_ids": token_ids,
"tokens": len(token_ids),
"token_ids_sha256": base.canonical_hash(token_ids),
"content_positions": positions,
"content_records": records,
"content_tokens": len(positions),
"wrapper_tokens": len(token_ids) - len(positions),
"boundary_crossing_tokens": crossing,
}
def boundary_control_domain(
loads: dict[str, np.ndarray],
mode: str,
replicates: int,
seed: int,
scope: str,
) -> dict[str, Any]:
"""Compute all eight cells with one shared source-bootstrap matrix."""
shapes = {value.shape for value in loads.values()}
if len(shapes) != 1:
raise ValueError(f"boundary-control shape mismatch: {sorted(shapes)}")
rows = next(iter(loads.values())).shape[0]
rng = np.random.default_rng(base.scoped_seed(seed, scope))
sampled = rng.integers(
0,
rows,
size=(replicates, rows),
endpoint=False,
)
point = {
condition: base.distribution(value, mode)
for condition, value in loads.items()
}
boot = {
condition: base.bootstrap_distributions(value, mode, sampled)
for condition, value in loads.items()
}
point_metrics = {
condition: base.metric_vector(value)
for condition, value in point.items()
}
boot_metrics = {
condition: base.metric_vector(value)
for condition, value in boot.items()
}
def system_edges(
values: dict[str, np.ndarray],
) -> dict[str, np.ndarray]:
return {
boundary: values[after] - values[before]
for boundary, (before, after) in SYSTEM_CELLS.items()
}
def edge_contrasts(
edges: dict[str, np.ndarray],
) -> dict[str, np.ndarray]:
return {
name: edges[after] - edges[before]
for name, (before, after) in SYSTEM_EDGE_CONTRASTS.items()
}
metric_system_edges: dict[str, Any] = {}
metric_system_edge_contrasts: dict[str, Any] = {}
for metric in point_metrics["s0_eos"]:
point_edges = system_edges(
{
condition: point_metrics[condition][metric]
for condition in CONDITIONS
}
)
boot_edges = system_edges(
{
condition: boot_metrics[condition][metric]
for condition in CONDITIONS
}
)
metric_system_edges[metric] = {
boundary: {
"point": float(point_edges[boundary][0]),
"ci95": base.interval(boot_edges[boundary]),
}
for boundary in BOUNDARY_LEVELS
}
point_contrasts = edge_contrasts(point_edges)
boot_contrasts = edge_contrasts(boot_edges)
metric_system_edge_contrasts[metric] = {
name: {
"point": float(point_contrasts[name][0]),
"ci95": base.interval(boot_contrasts[name]),
}
for name in SYSTEM_EDGE_CONTRASTS
}
point_vectors = system_edges(point)
boot_vectors = system_edges(boot)
point_vector_contrasts = edge_contrasts(point_vectors)
boot_vector_contrasts = edge_contrasts(boot_vectors)
distribution_system_edge_contrasts = {}
for name in SYSTEM_EDGE_CONTRASTS:
point_magnitude = 0.5 * np.abs(
point_vector_contrasts[name]
).sum()
boot_magnitude = 0.5 * np.abs(
boot_vector_contrasts[name]
).sum(axis=1)
distribution_system_edge_contrasts[name] = {
"half_l1_magnitude": {
"point": float(point_magnitude),
"ci95": base.interval(boot_magnitude),
},
"expert_share_difference_in_system_edges": (
point_vector_contrasts[name].tolist()
),
"expert_share_difference_in_system_edges_ci95": (
base.interval(boot_vector_contrasts[name])
),
}
point_tv: dict[str, float] = {}
boot_tv: dict[str, np.ndarray] = {}
point_jsd: dict[str, float] = {}
boot_jsd: dict[str, np.ndarray] = {}
system_edge_distances: dict[str, Any] = {}
for boundary, (before, after) in SYSTEM_CELLS.items():
point_delta = point[after] - point[before]
point_tv[boundary] = float(0.5 * np.abs(point_delta).sum())
boot_tv[boundary] = 0.5 * np.abs(
boot[after] - boot[before]
).sum(axis=1)
point_jsd[boundary] = float(
base.js_divergence(point[before], point[after])[0]
)
boot_jsd[boundary] = base.js_divergence(
boot[before],
boot[after],
)
system_edge_distances[boundary] = {
"total_variation": {
"point": point_tv[boundary],
"ci95": base.interval(boot_tv[boundary]),
},
"js_divergence": {
"point": point_jsd[boundary],
"ci95": base.interval(boot_jsd[boundary]),
"unit": "nats",
},
}
system_edge_distance_contrasts = {}
for name, (before, after) in SYSTEM_EDGE_CONTRASTS.items():
system_edge_distance_contrasts[name] = {
"total_variation_delta": {
"point": point_tv[after] - point_tv[before],
"ci95": base.interval(boot_tv[after] - boot_tv[before]),
},
"js_divergence_delta": {
"point": point_jsd[after] - point_jsd[before],
"ci95": base.interval(boot_jsd[after] - boot_jsd[before]),
"unit": "nats",
},
}
direct_substitutions = {}
for boundary in BOUNDARY_LEVELS:
if boundary == "eos":
continue
direct_substitutions[boundary] = {}
direct_point_tv: dict[int, float] = {}
direct_boot_tv: dict[int, np.ndarray] = {}
direct_point_jsd: dict[int, float] = {}
direct_boot_jsd: dict[int, np.ndarray] = {}
for system in (0, 1):
before = f"s{system}_eos"
after = f"s{system}_{boundary}"
point_delta = point[after] - point[before]
tv_boot = 0.5 * np.abs(
boot[after] - boot[before]
).sum(axis=1)
jsd_boot = base.js_divergence(boot[before], boot[after])
direct_point_tv[system] = float(
0.5 * np.abs(point_delta).sum()
)
direct_boot_tv[system] = tv_boot
direct_point_jsd[system] = float(
base.js_divergence(
point[before],
point[after],
)[0]
)
direct_boot_jsd[system] = jsd_boot
direct_substitutions[boundary][f"at_s{system}"] = {
"total_variation": {
"point": direct_point_tv[system],
"ci95": base.interval(tv_boot),
},
"js_divergence": {
"point": direct_point_jsd[system],
"ci95": base.interval(jsd_boot),
"unit": "nats",
},
}
direct_substitutions[boundary]["s1_minus_s0"] = {
"total_variation_delta": {
"point": direct_point_tv[1] - direct_point_tv[0],
"ci95": base.interval(
direct_boot_tv[1] - direct_boot_tv[0]
),
},
"js_divergence_delta": {
"point": direct_point_jsd[1] - direct_point_jsd[0],
"ci95": base.interval(
direct_boot_jsd[1] - direct_boot_jsd[0]
),
"unit": "nats",
},
}
return {
"metric_system_edges": metric_system_edges,
"metric_system_edge_contrasts": metric_system_edge_contrasts,
"system_edge_distances": system_edge_distances,
"system_edge_distance_contrasts": (
system_edge_distance_contrasts
),
"distribution_system_edge_contrasts": (
distribution_system_edge_contrasts
),
"direct_substitutions": direct_substitutions,
}
def layer_statistics(
prompt_rows: list[dict[str, Any]],
replicates: int,
seed: int,
layer_index: int,
) -> dict[str, Any]:
scopes = {}
for load_scope, load_key in (
("full_input", "full_load"),
("target_content", "content_load"),
):
modes = {}
for mode in ("token_weighted", "prompt_balanced"):
conditions: dict[str, Any] = {}
domain_loads: dict[str, dict[str, np.ndarray]] = {}
for domain in base.DOMAIN_ORDER:
rows = [
row
for row in prompt_rows
if row["domain"] == domain
]
domain_loads[domain] = {
condition: np.asarray(
[
row["conditions"][condition][load_key]
for row in rows
],
dtype=np.int64,
)
for condition in CONDITIONS
}
for condition in CONDITIONS:
conditions.setdefault(condition, {})[domain] = (
base.bootstrap_domain(
domain_loads[domain][condition],
mode,
replicates,
seed,
(
f"layer={layer_index}|scope={load_scope}|"
f"mode={mode}|condition={condition}|"
f"domain={domain}"
),
)
)
comparisons: dict[str, Any] = {}
for comparison, before, after in COMPARISONS:
comparisons[comparison] = {}
for domain in base.DOMAIN_ORDER:
comparisons[comparison][domain] = base.paired_domain(
domain_loads[domain][before],
domain_loads[domain][after],
mode,
replicates,
seed,
(
f"layer={layer_index}|scope={load_scope}|"
f"mode={mode}|comparison={comparison}|"
f"domain={domain}"
),
)
control = {
domain: boundary_control_domain(
domain_loads[domain],
mode,
replicates,
seed,
(
f"layer={layer_index}|scope={load_scope}|"
f"mode={mode}|boundary_control|domain={domain}"
),
)
for domain in base.DOMAIN_ORDER
}
modes[mode] = {
"conditions": conditions,
"comparisons": comparisons,
"boundary_control": control,
}
scopes[load_scope] = {"modes": modes}
return scopes
def install_control_contract() -> None:
"""Install the eight-cell renderer/statistics into the audited runner."""
base.CONDITIONS = CONDITIONS
base.FACTORS = FACTORS
base.COMPARISONS = COMPARISONS
base.ALIGNMENT_COMPARISONS = ALIGNMENT_COMPARISONS
base.condition_messages = condition_messages
base.render_variant = render_boundary_variant
base.layer_statistics = layer_statistics
def output_path_from_argv() -> Path:
try:
return Path(sys.argv[sys.argv.index("--output") + 1])
except (ValueError, IndexError) as error:
raise ValueError("--output is required") from error
def render_contract_summary() -> dict[str, Any]:
if not RENDER_AUDIT:
raise RuntimeError("render audit is empty")
by_content: dict[str, list[dict[str, Any]]] = {}
for row in RENDER_AUDIT:
by_content.setdefault(row["content_sha256"], []).append(row)
if len(by_content) != 128 and "--per-domain" not in sys.argv:
raise RuntimeError(
f"expected 128 selected contents; observed {len(by_content)}"
)
equal_lengths = 0
equal_target_positions = 0
one_id_replacements = 0
official_eos_exact = 0
for rows in by_content.values():
for system in (0, 1):
cells = [
row
for row in rows
if FACTORS[row["condition"]]["system"] == system
]
if len(cells) != len(BOUNDARY_LEVELS):
raise RuntimeError(
"render audit lacks one or more boundary cells"
)
if len({row["input_tokens"] for row in cells}) == 1:
equal_lengths += 1
if len({row["target_first_position"] for row in cells}) == 1:
equal_target_positions += 1
official_eos_exact += sum(
row["boundary"] == "eos"
and row["changed_token_ids_vs_official"] == 0
for row in cells
)
one_id_replacements += sum(
row["boundary"] != "eos"
and row["changed_token_ids_vs_official"] == 1
for row in cells
)
systems = 2 * len(by_content)
replacements = 3 * systems
return {
"selected_contents": len(by_content),
"system_groups": systems,
"equal_input_length_groups": equal_lengths,
"equal_target_position_groups": equal_target_positions,
"official_eos_cells_exact": official_eos_exact,
"one_id_counterfactual_cells_exact": one_id_replacements,
"expected_one_id_counterfactual_cells": replacements,
"all_group_lengths_equal": equal_lengths == systems,
"all_target_positions_equal": equal_target_positions == systems,
"all_counterfactuals_change_exactly_one_id": (
one_id_replacements == replacements
),
}
def finalize_result(path: Path) -> dict[str, Any]:
result = json.loads(path.read_text(encoding="utf-8"))
result["schema_version"] = 3
result["evidence_identity"] = (
"X / official BF16 weights and tokenizer; official EOS sequence plus "
"three paired one-token boundary-ID counterfactuals on local "
"truncated forward"
)
boundary = result["boundary"]
boundary.pop("factorial_claim", None)
boundary.update(
{
"boundary_token_identity_control": True,
"official_serialization_by_boundary": {
"eos": True,
"x": False,
"period": False,
"newline": False,
},
"single_input_id_intervention": True,
"turn_boundary_removed": False,
"role_markers_held_fixed": True,
"target_position_held_fixed": True,
"task_performance": False,
"causal_boundary": (
"EOS-to-control comparisons causally intervene on exactly "
"one prior input token ID within this fixed forward contract; "
"they do not remove the following User role marker, establish "
"general EOS semantics, or measure answer quality"
),
}
)
old_contract = result.pop("message_history_contract")
render_validation = render_contract_summary()
result["history_boundary_token_contract"] = {
"chat_template_revision": base.MODEL_REVISION,
"chat_template": old_contract["chat_template"],
"chat_template_sha256": old_contract["chat_template_sha256"],
"official_assistant_boundary": (
"Assistant: {content} + eos_token + User:"
),
"system_message": SYSTEM_MESSAGE,
"system_message_sha256": base.text_sha256(SYSTEM_MESSAGE),
"filler_user": FILLER_USER,
"filler_user_sha256": base.text_sha256(FILLER_USER),
"filler_assistant": FILLER_ASSISTANT,
"filler_assistant_sha256": base.text_sha256(FILLER_ASSISTANT),
"boundary_levels": list(BOUNDARY_LEVELS),
"boundary_token_ids": BOUNDARY_TOKEN_IDS,
"boundary_text_controls": BOUNDARY_TEXT,
"conditions": FACTORS,
"comparisons": [
{"name": name, "before": before, "after": after}
for name, before, after in COMPARISONS
],
"system_edge_contrasts": {
name: {
"before_boundary": before,
"after_boundary": after,
"definition": (
f"system edge at {after} minus system edge at {before}"
),
}
for name, (before, after) in (
SYSTEM_EDGE_CONTRASTS.items()
)
},
"render_validation": render_validation,
"target_role": "user",
"add_generation_prompt": True,
"scope_split": {
"full_input": (
"all rendered or counterfactually edited BOS, system/history, "
"target, newline, and generation-prompt token IDs"
),
"target_content": (
"exact intersection of (relative character span, token ID) "
"inside target user content across all eight conditions"
),
},
}
inference = result["inference_contract"]
inference["batch_grouping"] = (
"all eight boundary variants of one source prompt execute in the "
"same right-padded batch"
)
statistical = result["statistical_contract"]
statistical["paired_indices"] = (
"one sampled source-prompt index matrix is reused across all eight "
"cells for every boundary-token contrast within each "
"domain/layer/scope/mode"
)
statistical.pop("interaction_distribution_magnitude", None)
statistical["system_edge_contrast_distribution_magnitude"] = (
"0.5 * L1 norm of the signed difference between two system-edge "
"expert-share vectors; this is not labeled standard TV"
)
path.write_text(
json.dumps(result, indent=2, ensure_ascii=False) + "\n",
encoding="utf-8",
)
return result
def main() -> None:
install_control_contract()
output = output_path_from_argv()
with open(os.devnull, "w", encoding="utf-8") as sink:
with contextlib.redirect_stdout(sink):
base.main()
result = finalize_result(output)
payload = output.read_bytes()
print(
json.dumps(
{
"output": str(output),
"sha256": hashlib.sha256(payload).hexdigest(),
"bytes": len(payload),
"source_prompts": result["inference_contract"][
"total_source_prompts"
],
"prompt_variants": result["inference_contract"][
"total_prompt_variants"
],
"input_tokens_by_condition": result[
"inference_contract"
]["input_tokens_by_condition"],
"total_routes": result["inference_contract"][
"total_routes_all_conditions_all_moe_layers"
],
"render_validation": result[
"history_boundary_token_contract"
]["render_validation"],
},
indent=2,
ensure_ascii=False,
)
)
if __name__ == "__main__":
main()
+1
View File
@@ -9,6 +9,7 @@
"build": "astro build",
"preview": "astro preview --host 0.0.0.0",
"check": "astro check",
"build:data:deepseek-boundary": "node scripts/build-deepseek-boundary-compact.mjs",
"check:site": "node scripts/check-site.mjs",
"check:moe-browser": "node scripts/check-moe-browser.mjs",
"check:reasoning-browser": "node scripts/check-reasoning-browser.mjs",
@@ -0,0 +1,875 @@
# DeepSeek-V2-Lite 历史边界 Token 控制审计
> 状态:真实官方权重执行(X)<br />
> 模型:`deepseek-ai/DeepSeek-V2-Lite` **base checkpoint**<br />
> revision:`604d5664dddd88a0433dbae533b7fe9472482de0`<br />
> 执行边界:layer 0–6;观测 MoE layer 1–6<br />
> 样本:WikiText-2 / TNEWS / HumanEval / GSM8K 各 32 条<br />
> 正式运行与独立复跑:byte-exact<br />
> 完整 JSON SHA-256:`9bb93834ffd8536aeebe4325e45d2179ba590554c6ff6b6fcceaba2f499b9c37`
## 0. 一句话先说结论
在固定 DeepSeek-V2-Lite base 权重、固定 128 个公开样本、固定重复词元
历史、固定角色标记、固定长度、固定目标位置和固定八格 batch 时:
> 把历史 assistant 后面的官方 EOS token 换成 `x`、句点或换行这三个
> 单 token 对照,会让后续目标内容对 system 开关表现出更大的专家路由差异。
目标内容、prompt-balanced 口径下,六个 MoE 层 × 四个域的平均
system-edge total variation(TV)为:
```text
官方 EOS .0374
x .0539
句点 .0553
换行 .0492
```
但这句话必须与下面四条边界一起读:
1. 本实验使用的是 **base checkpoint**,不是 Chat/SFT checkpoint;
2. 三个替换条件不是官方有效对话序列,而是精确的 token-ID 反事实;
3. 下一条 `User:` 角色标记仍然存在,所以没有“移除全部回合边界”;
4. 本实验没有生成答案,因此没有证明 EOS 让回答更好、更稳或更正确。
最准确的命名是:
> **历史 assistant 边界上的单 token 身份路由控制。**
---
## 1. 为什么要继续拆上一轮
上一轮把 system 开/关与三种历史组合成了 2×3:
```text
none → repeated-token filler → fixed demo
```
其中 filler 与 demo:
- 都增加 17 个官方模板 token;
- 都包含一个 user turn;
- 都包含一个 assistant turn;
- 都在 assistant 内容后插入 EOS;
- 都把目标 user 推到同一绝对位置。
这已经把“具体示例文本”从“历史结构”里拆出了一层。
可是 none → filler 仍然同时增加了:
- 历史文本;
- user / assistant 角色切换;
- assistant EOS;
- 目标距离;
- 一整段新的隐藏状态演化。
因此上一轮无法回答:
> 历史缓冲模式中,assistant 后面的那个 EOS token 本身是不是一个特殊边界输入?
本轮只处理这个更窄、也更可辨识的问题。
---
## 2. 先看官方模板到底做了什么
固定 revision 的官方 `tokenizer_config.json` 中,核心模板是:
```jinja
{% if message['role'] == 'user' %}
{{ 'User: ' + message['content'] + '\n\n' }}
{% elif message['role'] == 'assistant' %}
{{ 'Assistant: ' + message['content'] + eos_token }}
{% elif message['role'] == 'system' %}
{{ message['content'] + '\n\n' }}
{% endif %}
```
所以 repeated-token 历史与目标 user 被序列化为:
```text
<BOS>
User: x x x x x x x x x
Assistant: x
<EOS>
User: TARGET
Assistant:
```
这里的 `<EOS>` 不是我们根据字符串猜出来的边界。
在固定 tokenizer 中:
| 名称 | 文本 / special token | token ID | token 数 |
|---|---|---:|---:|
| 官方 EOS | `<|end▁of▁sentence|>` | `100001` | 1 |
| 内容对照 | `x` | `87` | 1 |
| 标点对照 | `.` | `13` | 1 |
| 格式对照 | `\n` | `185` | 1 |
四个 ID 完全不同;后三者不是 special token。
官方模板事实可以在
[固定模型工件](https://huggingface.co/deepseek-ai/DeepSeek-V2-Lite/blob/604d5664dddd88a0433dbae533b7fe9472482de0/tokenizer_config.json)
中直接检查。
---
## 3. 为什么不直接“删掉 EOS”
最直觉的 EOS 消融是:
```text
有 EOS vs 没有 EOS
```
但直接删除会同时改变:
- 总 token 数;
- 目标 user 的绝对位置;
- RoPE 位置;
- padded sequence;
- 矩阵形状或有效 attention mask;
- 后续所有 token 在批中的列号。
那样看到差异时,无法知道它来自 EOS 身份还是“一格位置移动”。
因此本轮采用一条更干净的干预:
```text
官方 token IDs
... Assistant : x EOS User : ...
反事实 token IDs
... Assistant : x x User : ...
... Assistant : x . User : ...
... Assistant : x \n User : ...
```
每次只改一个历史 token ID。
这样可以固定:
- 序列长度;
- 目标绝对位置;
- user / assistant 文本;
- `User:` / `Assistant:` 角色标记;
- attention mask;
- generation prompt;
- 同一次 forward 的 batch shape。
代价是:
> 三个替换条件不再是官方 chat template 能自然产生的合法序列。
所以它们必须叫 **token-ID counterfactuals**,不能叫三种官方模板。
---
## 4. 2×4 设计
两个因子是:
- `S`:system 关闭 / 开启;
- `B`:边界 ID 为 EOS / `x` / `.` / `\n`。
八格如下:
| 条件 | system | 历史边界 | 官方序列 |
|---|---:|---|---:|
| `S0 / EOS` | 0 | EOS `100001` | 是 |
| `S1 / EOS` | 1 | EOS `100001` | 是 |
| `S0 / X` | 0 | `x` ID `87` | 否,单 ID 反事实 |
| `S1 / X` | 1 | `x` ID `87` | 否,单 ID 反事实 |
| `S0 / PERIOD` | 0 | `.` ID `13` | 否,单 ID 反事实 |
| `S1 / PERIOD` | 1 | `.` ID `13` | 否,单 ID 反事实 |
| `S0 / NEWLINE` | 0 | `\n` ID `185` | 否,单 ID 反事实 |
| `S1 / NEWLINE` | 1 | `\n` ID `185` | 否,单 ID 反事实 |
同一个 source prompt 的八格进入同一个 right-padded batch。
正式运行采用:
```text
batch-prompts = 4
rows per batch = 4 prompts × 8 conditions = 32
```
---
## 5. 三个问题,三个不同统计量
### 5.1 system 对目标路由影响多大
设:
```text
P(s, b, l, d)
```
为 system 水平 `s`、边界 `b`、层 `l`、域 `d` 的目标内容专家份额分布。
边界 `b` 下的 system edge 是:
```text
Δ_b = P(S1, b) - P(S0, b)
```
实际两分布之间的 standard TV 是:
```text
TV_b = 1/2 × Σ_e |P(S1, b, e) - P(S0, b, e)|
```
它回答:
> 开关固定 system 消息,目标 token 的总体专家份额移动了多少?
### 5.2 替换 EOS 后,system edge 变大还是变小
对每个普通 token 对照:
```text
C_x = TV_x - TV_EOS
C_period = TV_period - TV_EOS
C_newline = TV_newline - TV_EOS
```
若 `C_x > 0`,只表示:
> 在这个协议里,把 EOS 换成 `x` 后,system 开关造成的目标路由距离更大。
它不表示:
- EOS 更正确;
- EOS 更有语义;
- EOS 提高准确率;
- `x` 是噪声;
- system 的内容被“记住”或“忘掉”。
### 5.3 EOS 替换本身改了多少路由
固定 system 水平,定义:
```text
D_s,b = TV(P(s, EOS), P(s, b))
```
再比较:
```text
I_b = D_1,b - D_0,b
```
它回答:
> system 的存在是否会改变目标路由对边界 ID 替换的敏感度?
`I_b` 不是传统线性模型中的标量主效应,而是两个实际分布距离的差。
---
## 6. 为什么 bootstrap 必须八格共用
每个域有 32 个 source prompts。
本轮使用 2,000 次 source-paired bootstrap:
1. 对一个 layer × domain × scope × aggregation 生成一张
`2000 × 32` source index 矩阵;
2. 同一张矩阵同时重采八个条件;
3. 每次先重算八个专家分布;
4. 再计算 TV、JSD、CV 与对比;
5. 最后取 2.5% / 97.5% 分位数。
这样:
- 八格的 source 构成逐次完全相同;
- `TV_x - TV_EOS` 是配对差;
- `D_1,x - D_0,x` 也是配对差;
- 不会把 token 当成独立样本。
24 个 layer × domain 区间仍然是描述性格子:
> 没有进行家族级多重检验校正。
---
## 7. 正式执行账
### 7.1 语料与变体
| 项目 | 数量 |
|---|---:|
| 源 prompts | 128 |
| 每域 | 32 |
| 条件数 | 8 |
| prompt variants | 1,024 |
| canonical target tokens | 每条 23 |
| 六层逐条件精确对齐目标 token | 2,874 |
### 7.2 输入 token
| 条件 | 输入 token |
|---|---:|
| `S0 / EOS` | 6,074 |
| `S1 / EOS` | 8,122 |
| `S0 / X` | 6,074 |
| `S1 / X` | 8,122 |
| `S0 / PERIOD` | 6,074 |
| `S1 / PERIOD` | 8,122 |
| `S0 / NEWLINE` | 6,074 |
| `S1 / NEWLINE` | 8,122 |
| **合计** | **56,784** |
每个 source × system 分组内,四种边界逐条同长度。
### 7.3 路由量
```text
56,784 input tokens
× 6 observed MoE layers
× top-6 routed experts
= 2,044,224 route selections
```
加上前序公开实验:
```text
3,112,848 + 2,044,224 = 5,157,072
```
截至本轮,网站公开账本累计记录 **5,157,072 次真实 top-6 路由**。
### 7.4 单 ID 合同
| 校验 | 结果 |
|---|---:|
| source × system 分组 | 256 |
| 四格逐组同输入长度 | 256 / 256 |
| 四格逐组同目标起始位置 | 256 / 256 |
| 官方 EOS 格相对官方模板改 0 个 ID | 256 / 256 |
| 三个反事实格相对官方模板改 1 个 ID | 768 / 768 |
---
## 8. 主结果:24 张目标内容问题账
口径:
- scope:`target_content`
- aggregation:`prompt_balanced`
- 数字:system-edge TV
- 括号:replacement − EOS 的 source-paired 95% interval
| 层 | 域 | EOS | X | 句点 | 换行 | X−EOS | 句点−EOS | 换行−EOS |
|---:|---|---:|---:|---:|---:|---|---|---|
| 1 | 英文 | .0559 | .0836 | .0807 | .0802 | +.0277 [.0187,.0353] | +.0248 [.0166,.0310] | +.0243 [.0168,.0295] |
| 1 | 中文 | .0276 | .0505 | .0517 | .0510 | +.0229 [.0168,.0310] | +.0241 [.0169,.0319] | +.0234 [.0168,.0319] |
| 1 | 代码 | .0571 | .0804 | .0810 | .0917 | +.0233 [.0172,.0301] | +.0239 [.0173,.0298] | +.0346 [.0280,.0407] |
| 1 | 数学 | .0784 | .1069 | .1055 | .1056 | +.0285 [.0211,.0360] | +.0271 [.0197,.0342] | +.0272 [.0194,.0345] |
| 2 | 英文 | .0305 | .0432 | .0427 | .0463 | +.0127 [.0059,.0208] | +.0122 [.0047,.0202] | +.0158 [.0076,.0227] |
| 2 | 中文 | .0207 | .0272 | .0281 | .0286 | +.0066 [.0029,.0152] | +.0075 [.0029,.0143] | +.0079 [.0029,.0147] |
| 2 | 代码 | .0327 | .0532 | .0509 | .0468 | +.0205 [.0152,.0286] | +.0182 [.0126,.0252] | +.0141 [.0088,.0222] |
| 2 | 数学 | .0404 | .0522 | .0547 | .0516 | +.0118 [.0060,.0195] | +.0144 [.0082,.0222] | +.0113 [.0057,.0189] |
| 3 | 英文 | .0312 | .0432 | .0445 | .0423 | +.0120 [.0042,.0198] | +.0133 [.0052,.0189] | +.0111 [.0036,.0180] |
| 3 | 中文 | .0229 | .0317 | .0277 | .0263 | +.0088 [.0034,.0161] | +.0048 [.0003,.0126] | +.0034 [−.0008,.0122] |
| 3 | 代码 | .0403 | .0473 | .0446 | .0432 | +.0070 [.0018,.0125] | +.0043 [.0000,.0091] | +.0029 [−.0010,.0082] |
| 3 | 数学 | .0333 | .0360 | .0360 | .0422 | +.0027 [−.0014,.0100] | +.0027 [−.0008,.0104] | +.0089 [.0035,.0155] |
| 4 | 英文 | .0324 | .0510 | .0515 | .0453 | +.0187 [.0119,.0263] | +.0191 [.0128,.0267] | +.0129 [.0068,.0210] |
| 4 | 中文 | .0293 | .0404 | .0422 | .0331 | +.0111 [.0063,.0192] | +.0129 [.0085,.0204] | +.0039 [−.0002,.0109] |
| 4 | 代码 | .0413 | .0609 | .0750 | .0591 | +.0197 [.0134,.0266] | +.0337 [.0284,.0401] | +.0178 [.0133,.0259] |
| 4 | 数学 | .0512 | .0705 | .0705 | .0607 | +.0193 [.0132,.0289] | +.0192 [.0130,.0283] | +.0094 [.0045,.0168] |
| 5 | 英文 | .0342 | .0503 | .0508 | .0409 | +.0161 [.0109,.0225] | +.0165 [.0120,.0222] | +.0067 [.0012,.0137] |
| 5 | 中文 | .0256 | .0365 | .0447 | .0276 | +.0109 [.0077,.0195] | +.0191 [.0134,.0281] | +.0021 [−.0014,.0110] |
| 5 | 代码 | .0389 | .0501 | .0515 | .0384 | +.0113 [.0058,.0181] | +.0126 [.0067,.0185] | −.0004 [−.0054,.0059] |
| 5 | 数学 | .0442 | .0555 | .0530 | .0455 | +.0113 [.0049,.0195] | +.0088 [.0022,.0181] | +.0013 [−.0057,.0089] |
| 6 | 英文 | .0300 | .0562 | .0607 | .0434 | +.0262 [.0165,.0355] | +.0307 [.0220,.0395] | +.0134 [.0073,.0231] |
| 6 | 中文 | .0258 | .0428 | .0469 | .0292 | +.0170 [.0114,.0263] | +.0211 [.0152,.0315] | +.0034 [−.0007,.0118] |
| 6 | 代码 | .0320 | .0603 | .0634 | .0546 | +.0283 [.0207,.0345] | +.0314 [.0240,.0383] | +.0226 [.0137,.0307] |
| 6 | 数学 | .0427 | .0643 | .0690 | .0470 | +.0216 [.0154,.0304] | +.0264 [.0201,.0350] | +.0043 [−.0003,.0135] |
---
## 9. 把 24 张账压成可读摘要
### 9.1 平均 system-edge TV
| 边界 | 平均 TV | 相对 EOS |
|---|---:|---:|
| EOS | .03744 | — |
| X | .05393 | +.01649,约 +44% |
| 句点 | .05531 | +.01787,约 +48% |
| 换行 | .04920 | +.01176,约 +31% |
### 9.2 方向计数
| 对比 | 点估计 replacement > EOS | 区间完全高于 0 | 跨 0 |
|---|---:|---:|---:|
| X − EOS | 24 / 24 | 23 / 24 | 1 / 24 |
| 句点 − EOS | 24 / 24 | 23 / 24 | 1 / 24 |
| 换行 − EOS | 23 / 24 | 16 / 24 | 8 / 24 |
token-weighted 口径得到几乎相同的总体图景:
```text
EOS .03742
X .05392
句点 .05531
换行 .04919
```
因此主现象不是由某几个长 prompt 在 token-weighted 口径里取得更高权重造成的。
---
## 10. 深度与领域不是平的
### 10.1 按层平均
| MoE 层 | EOS | X | 句点 | 换行 |
|---:|---:|---:|---:|---:|
| 1 | .0548 | .0803 | .0797 | .0821 |
| 2 | .0311 | .0440 | .0441 | .0433 |
| 3 | .0319 | .0396 | .0382 | .0385 |
| 4 | .0385 | .0557 | .0598 | .0496 |
| 5 | .0357 | .0481 | .0500 | .0381 |
| 6 | .0326 | .0559 | .0600 | .0436 |
边界替换效应不是简单随深度单调放大或衰减。
Layer 3 的差距最小;Layer 6 的 X / 句点差距又明显扩大。
### 10.2 按域平均
| 域 | EOS | X | 句点 | 换行 |
|---|---:|---:|---:|---:|
| 英文 | .0357 | .0546 | .0552 | .0498 |
| 中文 | .0253 | .0382 | .0402 | .0326 |
| 代码 | .0404 | .0587 | .0611 | .0556 |
| 数学 | .0484 | .0642 | .0648 | .0588 |
四个域都保留 EOS 较小的总体模式,但幅度不同。
---
## 11. 逐 token 专家集合给出第二条证据
分布 TV 可能掩盖“哪些 token 改路由”。
因此本轮同时对 2,874 个精确对齐目标 token 聚合:
- ordered top-6 exact;
- unordered top-6 set exact;
- mean top-6 Jaccard;
- mean overlap。
下面报告 system 开/关两格之间的 set exact 与 Jaccard:
| 层 | EOS exact / Jaccard | X | 句点 | 换行 |
|---:|---:|---:|---:|---:|
| 1 | .565 / .864 | .437 / .815 | .441 / .817 | .441 / .815 |
| 2 | .695 / .905 | .551 / .855 | .554 / .856 | .566 / .860 |
| 3 | .682 / .901 | .573 / .865 | .573 / .865 | .588 / .868 |
| 4 | .673 / .895 | .522 / .839 | .511 / .834 | .538 / .846 |
| 5 | .656 / .887 | .522 / .839 | .525 / .840 | .575 / .859 |
| 6 | .651 / .888 | .486 / .827 | .474 / .821 | .519 / .845 |
六层中,EOS 条件的 system-pair top-6 set exact 和 Jaccard 都高于三个普通 token 对照。
这与“EOS 下 system-edge TV 更小”同向,但不是同一个统计量:
- TV 看聚合专家份额移动;
- exact / Jaccard 看逐 token 专家集合是否保持。
---
## 12. system 会不会改变对 EOS 替换的敏感度
目标内容、prompt-balanced 下,EOS → 普通 token 的直接 TV 均值:
| 替换 | system 关 | system 开 | S1 − S0 |
|---|---:|---:|---:|
| EOS → X | .03655 | .04396 | +.00741 |
| EOS → 句点 | .03630 | .04525 | +.00896 |
| EOS → 换行 | .04094 | .04652 | +.00557 |
点估计 `S1 > S0` 的格数:
```text
X 20 / 24
句点 21 / 24
换行 18 / 24
```
但配对区间完全高于零的格数只有:
```text
X 11 / 24
句点 12 / 24
换行 8 / 24
```
所以可以报告平均交互模式,不能写成每层每域都成立的规律。
---
## 13. target content 与 full input 必须分开
### 13.1 目标内容
```text
EOS .0374 → X .0539 / 句点 .0553 / 换行 .0492
```
方向高度一致。
### 13.2 完整输入
```text
EOS .13894
X .14159
句点 .14026
换行 .14076
```
replacement > EOS 的点估计仅为:
```text
X 16 / 24
句点 15 / 24
换行 15 / 24
```
完整输入包含:
- system 本身;
- filler 历史;
- 被替换的边界 token;
- `User:` / `Assistant:` 包装;
- 目标内容;
- generation prompt。
一个 ID 的替换在完整输入总体计数里很容易被稀释。
因此本轮最有辨识力的结论应限定在:
> **完全相同的后续目标内容 token。**
---
## 14. 负载“更均衡”仍然不是结论
若把每个条件的专家份额压成 CV,replacement − EOS 的 system-edge
CV 对比方向为:
| 替换 | 点估计负 | 点估计正 | 区间负 / 正 / 跨零 |
|---|---:|---:|---:|
| X | 6 | 18 | 4 / 8 / 12 |
| 句点 | 6 | 18 | 4 / 10 / 10 |
| 换行 | 6 | 18 | 2 / 9 / 13 |
它不像 TV 那样高度一致。
因此本轮不能推出:
- EOS 让专家更均衡;
- ordinary token 让专家更集中;
- 更小 system-edge TV 等于更优负载均衡。
TV 衡量两条件间的路由距离,CV 衡量单条件内的份额离散程度。
两者回答不同问题。
---
## 15. BF16 batch shape 再次显示为真实复现边界
新实验的官方 EOS 条件与上一轮 filler 条件具有逐条完全相同的 token IDs:
```text
256 / 256 condition sequences token-ID hash exact
```
但上一轮 batch 是:
```text
5 prompts × 6 conditions = 30 rows
```
本轮 batch 是:
```text
4 prompts × 8 conditions = 32 rows
```
跨两次运行比较六层、128 prompts、system 开关两格:
| 对象 | exact / 1,536 |
|---|---:|
| 完整 ordered top-6 route hash | 359 |
| 目标内容 ordered top-6 route hash | 758 |
| 完整 expert-load vector | 577 |
| 目标内容 expert-load vector | 996 |
逐层完整 route hash exact:
```text
Layer 1 242 / 256
Layer 2 57 / 256
Layer 3 30 / 256
Layer 4 2 / 256
Layer 5 11 / 256
Layer 6 17 / 256
```
这不是 token 合同漂移,而是矩阵形状改变后,BF16 前向中的微小数值差异
沿层传播,并让接近 top-k 边界的 gate 排序分叉。
因此正式主结论只使用:
> 同一次 32-row 八格 batch 内的配对对比。
不能把上一轮 `.0378` 与本轮 `.0374` 当成“EOS 效应复现误差”做统计比较。
---
## 16. 这次能说“因果”到哪一层
本轮比观察性相关更强,因为在每个 EOS → control 对比中:
- source prompt 相同;
- system 相同;
- token 数相同;
- attention mask 相同;
- 目标位置相同;
- 角色标记相同;
- batch 相同;
- 只有一个历史 input ID 被替换。
因此在固定模型与固定 forward 合同内,可以说:
> 这个单 input-ID 干预造成了观测到的后续路由差异。
但不能继续跳到:
> “EOS 的语义导致模型正确关闭回合。”
缺失的证据包括:
- Chat/SFT checkpoint 对照;
- 行为生成与任务指标;
- attention / residual / gate-logit 中介;
- EOS embedding 与多个 special-token 对照;
- 去掉 `User:` 角色标记的独立实验;
- 完整 27 层。
这是“输入干预的路由因果”,不是“语义机制的完整因果解释”。
---
## 17. 文献脉络:哪些来源支持什么
### 17.1 DeepSeek 官方工件
[DeepSeek-V2-Lite tokenizer config](https://huggingface.co/deepseek-ai/DeepSeek-V2-Lite/blob/604d5664dddd88a0433dbae533b7fe9472482de0/tokenizer_config.json)
直接规定 assistant 内容后拼接 `eos_token`。
它支持:
- 本实验边界位置来自官方工件;
- 不是自行发明的字符串模板。
它不支持:
- base checkpoint 已经学习了聊天角色层级;
- EOS 的内部功能就是“回合关闭”。
### 17.2 Chat template 是序列化合同
[Hugging Face Chat Templates](https://huggingface.co/docs/transformers/en/chat_templating)
明确说明 chat 消息最终仍会变成一条线性 token 序列,角色和边界通常通过
模型特定的 control tokens 注入。
它支持:
- 角色字段本身不会被 Transformer 神奇地看见;
- 真正进入模型的是序列化 token。
它不支持:
- 任意模型都共享相同 role / EOS 语义。
### 17.3 EOT 可以成为监督目标
[LIMA: Less Is More for Alignment](https://arxiv.org/abs/2305.11206)
在用户与 assistant 说话者之间加入专门 EOT token,并训练模型在回答结束时
预测它。
它支持:
- 对话边界 token 可以通过监督训练学成。
它不支持:
- DeepSeek-V2-Lite base 的 EOS 等同于 LIMA 的 EOT;
- 本轮结果来自 instruction tuning。
### 17.4 EOS 会改变隐藏状态动力学,但任务不同
[The EOS Decision and Length Extrapolation](https://aclanthology.org/2020.blackboxnlp-1.26/)
比较训练时预测 / 不预测 EOS 的序列模型,发现 EOS 决策会伴随 length
manifolds 与 length attractors,并影响长度外推。
它支持:
- EOS 不是只能在 decoder 外部解释的“停止按钮”;
- EOS 训练目标可能改变内部表示动力学。
它不支持:
- 把 LSTM / SCAN 的结论直接迁移到 DeepSeekMoE;
- 解释本轮具体专家路由。
### 17.5 跨模型的 end-of-turn 规范
[Meta Llama 3 official prompt format](https://github.com/meta-llama/llama3)
规定每条消息以 `<|eot_id|>` 结束。
它支持:
- 现代 chat 模型确实把 message boundary 编成特殊 token。
它不支持:
- Llama 的 EOT 与 DeepSeek base EOS 具有相同内部机制。
### 17.6 irrelevant context 不是 routing 论文
[Large Language Models Can Be Easily Distracted by Irrelevant Context](https://proceedings.mlr.press/v202/shi23a.html)
与
[Lost in the Middle](https://aclanthology.org/2024.tacl-1.9/)
说明额外上下文和位置会影响行为表现。
它们支持:
- 历史内容、位置与边界值得独立控制。
它们不支持:
- route TV 必然与准确率同向;
- repeated `x` 就是一般 irrelevant context。
### 17.7 干预方法也有边界
[Towards Best Practices of Activation Patching](https://arxiv.org/abs/2309.16042)
说明干预结论会受到 metric 与 corruption method 选择影响。
本轮因此:
- 不用一个 CV 指标代表全部路由现象;
- 同报 TV、JSD、逐 token exact/Jaccard;
- 明确普通 token 对照不是自然语言合法模板;
- 不把单 ID 输入干预升级为完整内部机制定位。
---
## 18. 正式运行与独立复跑
主结果:
```text
src/data/deepseek-v2-lite-routing-history-boundary-token-control.json
```
独立复跑:
```text
src/data/deepseek-v2-lite-routing-history-boundary-token-control-repro.json
```
两者:
```text
bytes 54,254,445
SHA256 9bb93834ffd8536aeebe4325e45d2179ba590554c6ff6b6fcceaba2f499b9c37
cmp byte-exact
```
执行命令:
```bash
PYTHONPATH=/path/to/transformers-4.41.2-deps:/usr/lib/python3/dist-packages \
python -B experiments/deepseek/v2_lite_routing_history_boundary_token_control.py \
--artifact-dir /path/to/deepseek-v2-lite \
--human-eval /path/to/HumanEval.jsonl.gz \
--gsm8k /path/to/gsm8k/test.jsonl \
--tnews /path/to/tnews/test.json \
--tnews-archive /path/to/tnews_public.zip \
--wikitext /path/to/wikitext-validation.parquet \
--output src/data/deepseek-v2-lite-routing-history-boundary-token-control.json \
--per-domain 32 \
--content-tokens 23 \
--batch-prompts 4 \
--layers 7 \
--bootstrap 2000 \
--seed 20260729 \
--captured-at 2026-07-29T11:12:00+00:00
```
---
## 19. 已证明、未证明、下一步
### 已证明
- 官方历史 EOS 在固定模板中是独立 token ID;
- 三个反事实条件逐条只替换该一个 ID;
- 八格逐组同长度、同目标位置、同 batch;
- EOS 下的目标内容 system-edge TV 平均小于三个普通 token 对照;
- X / 句点的方向在 24 / 24 格一致;
- 逐 token set exact / Jaccard 与分布 TV 同向;
- 正式运行与独立复跑 byte-exact。
### 未证明
- EOS “理解”了回合结束;
- base checkpoint 具有 chat role hierarchy;
- EOS 让回答更正确;
- 更小 TV 就是更优;
- `x` / 句点 / 换行代表全部普通 token;
- 现象覆盖全部 27 层;
- 现象覆盖 DeepSeek-V2-Lite-Chat;
- 现象在不同 batch shape 下 route hash 不变。
### 下一步
1. **角色标记控制**:固定内容与长度,替换 / 移除下一条 `User:` 标记;
2. **special-token 家族控制**:加入 BOS、其他已训练 special token 与多个普通 token;
3. **Chat checkpoint 对照**:比较 base 与 V2-Lite-Chat;
4. **行为联结**:生成答案,测任务正确率、格式终止与 route TV 的关系;
5. **gate-logit 中介**:记录替换后 router logits、top-k margin 与翻转位置;
6. **完整 27 层**:获得其余 shards 后扩展;
7. **固定 batch-shape 复跑**:用 dummy rows 锁定矩阵形状,验证跨实验可比性。
---
## 20. 最终课程表述
可以写:
> 在 DeepSeek-V2-Lite base 的固定前六个 MoE 层中,官方 assistant EOS
> 与三个单 token 反事实对照产生了可重复的后续路由差异;EOS 条件下,
> system 开关对精确目标内容的专家分布影响更小,且逐 token top-6 集合
> 更稳定。
不可以写:
> DeepSeek 用 EOS 理解并关闭了对话,所以回答更稳定。
前一句是当前证据。
后一句仍需要 Chat 权重、行为指标与内部中介实验。
+148
View File
@@ -0,0 +1,148 @@
import { createHash } from "node:crypto";
import { readFileSync, statSync, writeFileSync } from "node:fs";
import { resolve } from "node:path";
const root = resolve(import.meta.dirname, "..");
const mainPath = resolve(
root,
"src/data/deepseek-v2-lite-routing-history-boundary-token-control.json",
);
const reproPath = resolve(
root,
"src/data/deepseek-v2-lite-routing-history-boundary-token-control-repro.json",
);
const outputPath = resolve(
root,
"src/data/deepseek-v2-lite-routing-history-boundary-token-control-compact.json",
);
const sha256 = (path) => createHash("sha256")
.update(readFileSync(path))
.digest("hex");
const mainSha256 = sha256(mainPath);
const reproSha256 = sha256(reproPath);
const mainBytes = statSync(mainPath).size;
const reproBytes = statSync(reproPath).size;
const exact = mainSha256 === reproSha256 && mainBytes === reproBytes;
if (!exact) {
throw new Error("boundary-token formal run and rerun are not byte-exact");
}
const boundary = JSON.parse(readFileSync(mainPath, "utf8"));
const boundaryEdges = [
"system_eos",
"system_x",
"system_period",
"system_newline",
"x_at_s0",
"x_at_s1",
"period_at_s0",
"period_at_s1",
"newline_at_s0",
"newline_at_s1",
];
const aggregateAlignment = (layer, domain, edge) => {
const rows = layer.prompts
.filter((prompt) => prompt.domain === domain)
.map((prompt) => prompt.alignments[edge]);
const aligned = rows.reduce(
(sum, row) => sum + row.aligned_tokens,
0,
);
const setExact = rows.reduce(
(sum, row) => sum + row.set_topk_exact,
0,
);
const orderedExact = rows.reduce(
(sum, row) => sum + row.ordered_topk_exact,
0,
);
const weightedJaccard = rows.reduce(
(sum, row) => sum + row.mean_jaccard * row.aligned_tokens,
0,
);
return {
aligned,
setExactRate: setExact / aligned,
orderedExactRate: orderedExact / aligned,
meanJaccard: weightedJaccard / aligned,
};
};
const compact = {
schemaVersion: 1,
source: {
mainSha256,
reproSha256,
mainBytes,
reproBytes,
exact,
},
domains: boundary.corpus_contract.domains,
labels: boundary.corpus_contract.domain_labels,
inference: boundary.inference_contract,
contract: {
tokenIds: boundary.history_boundary_token_contract.boundary_token_ids,
validation: boundary.history_boundary_token_contract.render_validation,
official: boundary.boundary.official_serialization_by_boundary,
},
layers: boundary.layers.slice(1).map((layer) => ({
layer: layer.layer,
alignment: Object.fromEntries(
boundary.corpus_contract.domains.map((domain) => [
domain,
Object.fromEntries(
boundaryEdges.map((edge) => [
edge,
aggregateAlignment(layer, domain, edge),
]),
),
]),
),
scopes: Object.fromEntries(
["target_content", "full_input"].map((scope) => [
scope,
{
modes: Object.fromEntries(
["prompt_balanced", "token_weighted"].map((mode) => {
const control = (
layer.statistics[scope].modes[mode].boundary_control
);
return [
mode,
Object.fromEntries(
boundary.corpus_contract.domains.map((domain) => [
domain,
{
distances: control[domain].system_edge_distances,
contrasts: (
control[domain].system_edge_distance_contrasts
),
cvEdges: control[domain].metric_system_edges.cv,
cvContrasts: (
control[domain].metric_system_edge_contrasts.cv
),
direct: control[domain].direct_substitutions,
},
]),
),
];
}),
),
},
]),
),
})),
};
writeFileSync(
outputPath,
`${JSON.stringify(compact, null, 2)}\n`,
"utf8",
);
process.stdout.write(
`${outputPath}\n${mainSha256}\n${mainBytes} bytes source → `
+ `${statSync(outputPath).size} bytes compact\n`,
);
+95 -4
View File
@@ -519,6 +519,68 @@ await evaluate(`(() => {
await pause(120);
await screenshot("/tmp/llm-atlas-deepseek-distance-results-desktop.png");
const artifactBoundary = await evaluate(`(() => {
const root = document.querySelector("[data-dsv2-lab]");
root.querySelector('[data-artifact-tab="boundary"]').click();
const read = () => ({
panel: root.querySelector("[data-artifact-panel]:not([hidden])").dataset.artifactPanel,
tokenCards: root.querySelectorAll(".boundary-token-grid > article").length,
domainCards: root.querySelectorAll("[data-boundary-domain-grid] > article").length,
domains: [...root.querySelectorAll("[data-boundary-domain-grid] > article")].map((node) => ({
label: node.querySelector(":scope > span").textContent.trim(),
values: [...node.querySelectorAll(".boundary-tv-ladder > b")].map((cell) => ({
name: cell.querySelector("small").textContent.trim(),
value: cell.querySelector("strong").textContent.trim(),
})),
effect: node.querySelector(":scope > strong").textContent.trim(),
className: node.querySelector(":scope > strong").className,
ci: node.querySelector(":scope > p").textContent.trim(),
stability: node.querySelector(":scope > small").textContent.trim(),
substitution: node.querySelector(":scope > em").textContent.trim(),
interaction: node.querySelector(":scope > u").textContent.trim(),
cv: node.querySelector(":scope > i").textContent.trim(),
})),
summary: [...root.querySelectorAll('[data-artifact-panel="boundary"] .boundary-summary article b')].map((node) => node.textContent.trim()),
depthRows: root.querySelectorAll("[data-boundary-depth-map] > div").length,
depthCells: root.querySelectorAll("[data-boundary-depth-map] > div > span").length,
depthTitle: root.querySelector("[data-boundary-depth-title]").textContent.trim(),
exact: root.querySelector(".boundary-ledger .exact b").textContent.trim(),
note: root.querySelector("[data-boundary-note]").textContent.trim(),
trackToken: root.querySelector("[data-boundary-track-token]").textContent.trim(),
activeLayer: root.querySelector("[data-boundary-layer].active").textContent.trim(),
activeScope: root.querySelector('[data-boundary-scope][aria-pressed="true"]').dataset.boundaryScope,
activeMode: root.querySelector('[data-boundary-mode][aria-pressed="true"]').dataset.boundaryMode,
activeContrast: root.querySelector('[data-boundary-contrast][aria-pressed="true"]').dataset.boundaryContrast,
});
const layer1X = read();
root.querySelector('[data-boundary-layer="4"]').click();
const layer4X = read();
root.querySelector('[data-boundary-contrast="period_minus_eos"]').click();
const layer4Period = read();
root.querySelector('[data-boundary-contrast="newline_minus_eos"]').click();
const layer4Newline = read();
root.querySelector('[data-boundary-scope="full_input"]').click();
const layer4Full = read();
root.querySelector('[data-boundary-mode="token_weighted"]').click();
const layer4FullToken = read();
root.querySelector('[data-boundary-layer="1"]').click();
root.querySelector('[data-boundary-scope="target_content"]').click();
root.querySelector('[data-boundary-mode="prompt_balanced"]').click();
root.querySelector('[data-boundary-contrast="x_minus_eos"]').click();
return { layer1X, layer4X, layer4Period, layer4Newline, layer4Full, layer4FullToken, restored: read() };
})()`);
await evaluate(`(() => {
document.querySelector("[data-dsv2-lab]").scrollIntoView({ block: "start", behavior: "instant" });
window.scrollBy(0, -82);
})()`);
await pause(180);
await screenshot("/tmp/llm-atlas-deepseek-boundary-desktop.png");
await evaluate(`(() => {
document.querySelector(".boundary-domain-grid").scrollIntoView({ block: "center", behavior: "instant" });
})()`);
await pause(120);
await screenshot("/tmp/llm-atlas-deepseek-boundary-results-desktop.png");
const artifactEvidence = await evaluate(`(() => {
const root = document.querySelector("[data-dsv2-lab]");
root.querySelector('[data-artifact-tab="evidence"]').click();
@@ -609,6 +671,13 @@ const mobile = await evaluate(`(() => {
distanceContrasts: artifact.querySelectorAll("[data-distance-contrast]").length,
distanceDomainCards: artifact.querySelectorAll("[data-distance-domain-grid] > article").length,
distanceDepthCells: artifact.querySelectorAll("[data-distance-depth-map] > div > span").length,
boundaryLayers: artifact.querySelectorAll("[data-boundary-layer]").length,
boundaryScopes: artifact.querySelectorAll("[data-boundary-scope]").length,
boundaryModes: artifact.querySelectorAll("[data-boundary-mode]").length,
boundaryContrasts: artifact.querySelectorAll("[data-boundary-contrast]").length,
boundaryTokenCards: artifact.querySelectorAll(".boundary-token-grid > article").length,
boundaryDomainCards: artifact.querySelectorAll("[data-boundary-domain-grid] > article").length,
boundaryDepthCells: artifact.querySelectorAll("[data-boundary-depth-map] > div > span").length,
offenders: [...document.querySelectorAll("body *")]
.filter((node) => !node.closest(".paper-chain, .advantage-table, .precision-table, .mapping-table, [data-deepseek-lab], [data-dsv2-lab]"))
.filter((node) => node.getBoundingClientRect().right > document.documentElement.clientWidth + 1)
@@ -678,8 +747,22 @@ await evaluate(`(() => {
})()`);
await pause(120);
await screenshot("/tmp/llm-atlas-deepseek-distance-results-mobile.png");
await evaluate(`(() => {
const artifact = document.querySelector("[data-dsv2-lab]");
artifact.querySelector('[data-artifact-tab="boundary"]').click();
artifact.scrollIntoView({ block: "start", behavior: "instant" });
window.scrollBy(0, -70);
})()`);
await pause(180);
await screenshot("/tmp/llm-atlas-deepseek-boundary-mobile.png");
await evaluate(`(() => {
document.querySelector(".boundary-domain-grid").scrollIntoView({ block: "start", behavior: "instant" });
window.scrollBy(0, -72);
})()`);
await pause(120);
await screenshot("/tmp/llm-atlas-deepseek-boundary-results-mobile.png");
const report = { overview, capacity, cache, codesign, rl, artifactRoute, artifactLoad, artifactCache, artifactAbsorb, artifactCorpus, artifactTemplate, artifactHistory, artifactDistance, artifactEvidence, home, papers, mobile, exceptions };
const report = { overview, capacity, cache, codesign, rl, artifactRoute, artifactLoad, artifactCache, artifactAbsorb, artifactCorpus, artifactTemplate, artifactHistory, artifactDistance, artifactBoundary, artifactEvidence, home, papers, mobile, exceptions };
console.log(JSON.stringify(report, null, 2));
const numeric = (text) => Number.parseFloat(text.replaceAll(",", ""));
@@ -689,8 +772,8 @@ if (overview.sections !== 26 || overview.tocLinks !== 26) failures.push("二十
if (overview.ledgers !== 24 || overview.waves !== 10) failures.push("二十四张问题账或十次转向结构异常");
if (overview.paperLinks !== 60 || overview.branches !== 5 || overview.followups !== 1) failures.push("论文链、旁支或公开后续标记异常");
if (overview.labTabs !== 4 || overview.labPanels !== 4) failures.push("四联实验结构异常");
if (overview.artifactTabs !== 9 || overview.artifactPanels !== 9 || overview.artifactLayers !== 27) failures.push("真实权重九联实验结构异常");
if (overview.heroLabs !== "13 个可操作实验") failures.push("DeepSeek 实验总数账异常");
if (overview.artifactTabs !== 10 || overview.artifactPanels !== 10 || overview.artifactLayers !== 27) failures.push("真实权重十联实验结构异常");
if (overview.heroLabs !== "14 个可操作实验") failures.push("DeepSeek 实验总数账异常");
if (overview.navLinks !== 20 || home.navLinks !== 20 || mobile.mobileLinks !== 20 || overview.activeNav !== "DeepSeek") failures.push("全站导航未同步 DeepSeek");
if (overview.documentOverflow > 1 || mobile.documentOverflow > 1) failures.push("桌面或移动端存在文档级横向溢出");
if (capacity.initial.panel !== "capacity" || capacity.initial.total !== "32.1× FFN" || capacity.initial.active !== "1.13× FFN") failures.push("V3 稀疏容量初始账异常");
@@ -744,12 +827,20 @@ if (artifactDistance.layer4Filler.domains[1].values.map((cell) => cell.value).jo
if (artifactDistance.layer4Demo.domains[1].effect !== "DEMO − FILLER · ΔTV -0.014" || artifactDistance.layer4Demo.activeContrast !== "demo_minus_filler" || !artifactDistance.layer4Demo.depthTitle.includes("文本替换")) failures.push("等长历史文本替换 contrast 异常");
if (artifactDistance.layer4FullDemo.domains[1].effect !== "DEMO − FILLER · ΔTV -0.023" || artifactDistance.layer4FullDemo.activeScope !== "full_input" || !artifactDistance.layer4FullDemo.note.includes("完整输入")) failures.push("等长历史完整输入 scope 异常");
if (artifactDistance.layer4FullToken.activeMode !== "token_weighted" || artifactDistance.restored.activeScope !== "target_content" || artifactDistance.restored.activeMode !== "prompt_balanced" || artifactDistance.restored.activeContrast !== "filler_minus_none") failures.push("等长历史聚合口径或恢复状态异常");
if (artifactBoundary.layer1X.panel !== "boundary" || artifactBoundary.layer1X.tokenCards !== 4 || artifactBoundary.layer1X.domainCards !== 4 || artifactBoundary.layer1X.depthRows !== 4 || artifactBoundary.layer1X.depthCells !== 24 || artifactBoundary.layer1X.exact !== "BYTE-EXACT") failures.push("单 token 边界控制结构或独立复跑闸门异常");
if (artifactBoundary.layer1X.domains[0].values.map((cell) => cell.value).join("/") !== "0.056/0.084/0.081/0.080" || artifactBoundary.layer1X.domains[0].effect !== "X − EOS · ΔTV +0.028" || !artifactBoundary.layer1X.domains[0].ci.includes("+0.019, +0.035")) failures.push("L1 英文边界替换统计异常");
if (artifactBoundary.layer1X.summary.join("|") !== "24 / 24 ↑|24 / 24 ↑|23 / 24 ↑|.037 → .054 / .055 / .049" || artifactBoundary.layer1X.trackToken !== "X" || !artifactBoundary.layer1X.domains[0].stability.includes("EOS 50.7%")) failures.push("边界控制总账、协议轨或逐 token 稳定性异常");
if (artifactBoundary.layer4X.domains[0].values.map((cell) => cell.value).join("/") !== "0.032/0.051/0.051/0.045" || artifactBoundary.layer4X.domains[0].effect !== "X − EOS · ΔTV +0.019") failures.push("L4 英文 x 边界替换异常");
if (artifactBoundary.layer4Period.domains[2].effect !== "PERIOD − EOS · ΔTV +0.034" || artifactBoundary.layer4Period.activeContrast !== "period_minus_eos" || artifactBoundary.layer4Period.trackToken !== ".") failures.push("L4 代码句点边界替换异常");
if (artifactBoundary.layer4Newline.activeContrast !== "newline_minus_eos" || artifactBoundary.layer4Newline.trackToken !== "↵" || !artifactBoundary.layer4Newline.depthTitle.includes("换行")) failures.push("换行边界替换切换异常");
if (artifactBoundary.layer4Full.activeScope !== "full_input" || !artifactBoundary.layer4Full.note.includes("完整输入") || artifactBoundary.layer4Full.domains[0].values[0].value === artifactBoundary.layer4Newline.domains[0].values[0].value) failures.push("边界控制完整输入 scope 异常");
if (artifactBoundary.layer4FullToken.activeMode !== "token_weighted" || artifactBoundary.restored.activeLayer !== "L1" || artifactBoundary.restored.activeScope !== "target_content" || artifactBoundary.restored.activeMode !== "prompt_balanced" || artifactBoundary.restored.activeContrast !== "x_minus_eos") failures.push("边界控制聚合口径或恢复状态异常");
if (artifactEvidence.panel !== "evidence" || artifactEvidence.layers !== 27 || artifactEvidence.executed !== 7 || artifactEvidence.split !== 1 || artifactEvidence.unloaded !== 19 || artifactEvidence.exact !== "31 / 31") failures.push("真实工件执行边界或复跑闸门异常");
if (!artifactEvidence.dependency.includes("Transformers 5.5") || !artifactEvidence.dependency.includes("4.41.2") || !artifactEvidence.boundary.includes("完整 27 层生成")) failures.push("依赖版本或未覆盖边界异常");
if (artifactEvidence.keyboardSelected !== "load" || artifactEvidence.keyboardVisible !== "load") failures.push("真实工件实验键盘 tab 导航异常");
if (home.releaseCards !== 17 || !home.firstRelease.includes("47 页不再压成摘要") || home.firstHref !== "/k3/" || home.paperCount !== "486") failures.push("首页 DeepSeek 首发入口或论文数异常");
if (papers.total !== 486 || !papers.hasFilter || papers.visible < 20 || !papers.hasCoder || !papers.hasEngram) failures.push("论文库 DeepSeek 聚光异常");
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4 || mobile.artifactTabs !== 9 || mobile.artifactHeatCells !== 64 || mobile.corpusCohorts !== 3 || mobile.lengthDeltaCards !== 4 || mobile.templateLayers !== 6 || mobile.templateScopes !== 2 || mobile.templateModes !== 2 || mobile.templateDomainCards !== 4 || mobile.templateDepthCells !== 24 || mobile.historyLayers !== 6 || mobile.historyScopes !== 2 || mobile.historyModes !== 2 || mobile.historyEffects !== 3 || mobile.historyDomainCards !== 4 || mobile.historyDepthCells !== 24 || mobile.distanceLayers !== 6 || mobile.distanceScopes !== 2 || mobile.distanceModes !== 2 || mobile.distanceContrasts !== 2 || mobile.distanceDomainCards !== 4 || mobile.distanceDepthCells !== 24) failures.push("移动端导航或实验异常");
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4 || mobile.artifactTabs !== 10 || mobile.artifactHeatCells !== 64 || mobile.corpusCohorts !== 3 || mobile.lengthDeltaCards !== 4 || mobile.templateLayers !== 6 || mobile.templateScopes !== 2 || mobile.templateModes !== 2 || mobile.templateDomainCards !== 4 || mobile.templateDepthCells !== 24 || mobile.historyLayers !== 6 || mobile.historyScopes !== 2 || mobile.historyModes !== 2 || mobile.historyEffects !== 3 || mobile.historyDomainCards !== 4 || mobile.historyDepthCells !== 24 || mobile.distanceLayers !== 6 || mobile.distanceScopes !== 2 || mobile.distanceModes !== 2 || mobile.distanceContrasts !== 2 || mobile.distanceDomainCards !== 4 || mobile.distanceDepthCells !== 24 || mobile.boundaryLayers !== 6 || mobile.boundaryScopes !== 2 || mobile.boundaryModes !== 2 || mobile.boundaryContrasts !== 3 || mobile.boundaryTokenCards !== 4 || mobile.boundaryDomainCards !== 4 || mobile.boundaryDepthCells !== 24) failures.push("移动端导航或实验异常");
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
+627 -12
View File
@@ -16,6 +16,7 @@ import rawHistory from "@/data/deepseek-v2-lite-routing-history-factorial.json";
import rawHistoryRepro from "@/data/deepseek-v2-lite-routing-history-factorial-repro.json";
import rawDistance from "@/data/deepseek-v2-lite-routing-history-distance-control.json";
import rawDistanceRepro from "@/data/deepseek-v2-lite-routing-history-distance-control-repro.json";
import rawBoundaryCompact from "@/data/deepseek-v2-lite-routing-history-boundary-token-control-compact.json";
const trace = rawTrace as any;
const repro = rawRepro as any;
@@ -34,6 +35,7 @@ const history = rawHistory as any;
const historyRepro = rawHistoryRepro as any;
const distance = rawDistance as any;
const distanceRepro = rawDistanceRepro as any;
const boundaryCompact = rawBoundaryCompact as any;
const absorbExact = JSON.stringify(absorb) === JSON.stringify(absorbRepro);
const corpusExact = JSON.stringify(corpus) === JSON.stringify(corpusRepro);
const matched16Exact = JSON.stringify(matched16) === JSON.stringify(matched16Repro);
@@ -41,6 +43,7 @@ const matched24Exact = JSON.stringify(matched24) === JSON.stringify(matched24Rep
const templateExact = JSON.stringify(template) === JSON.stringify(templateRepro);
const historyExact = JSON.stringify(history) === JSON.stringify(historyRepro);
const distanceExact = JSON.stringify(distance) === JSON.stringify(distanceRepro);
const boundaryExact = boundaryCompact.source.exact;
const bytes = (value: number) => value >= 1024
? `${(value / 1024).toFixed(2)} KiB`
: `${value.toLocaleString()} B`;
@@ -329,6 +332,7 @@ const distanceCompact = {
})),
};
const distanceCompactJson = JSON.stringify(distanceCompact).replaceAll("<", "\\u003c");
const boundaryCompactJson = JSON.stringify(boundaryCompact).replaceAll("<", "\\u003c");
const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoint_tensor_bytes;
---
@@ -340,7 +344,7 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
</div>
<p>
固定官方 revision、tokenizer、模型代码和 BF16 第一分片;RTX 5090 连续执行 layer 0–6,
从 3,240 次 token 显微轨迹扩到 3,112,848 次公开语料路由,并让 layer-1 权重继续走入官方吸收式 cache。
从 3,240 次 token 显微轨迹扩到 5,157,072 次公开语料路由,并让 layer-1 权重继续走入官方吸收式 cache。
所有结论都带证据身份与停止线。
</p>
</header>
@@ -377,8 +381,11 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
<button type="button" role="tab" data-artifact-tab="distance" aria-selected="false" tabindex="-1">
<span>08</span><b>等长历史控制</b><small>none → filler → demo</small>
</button>
<button type="button" role="tab" data-artifact-tab="boundary" aria-selected="false" tabindex="-1">
<span>09</span><b>边界单词元控制</b><small>EOS ↔ x / . / ↵</small>
</button>
<button type="button" role="tab" data-artifact-tab="evidence" aria-selected="false" tabindex="-1">
<span>09</span><b>证据断面</b><small>revision · shards · rerun</small>
<span>10</span><b>证据断面</b><small>revision · shards · rerun</small>
</button>
</div>
@@ -1164,6 +1171,147 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
</div>
</section>
<section class="artifact-panel" data-artifact-panel="boundary" hidden>
<div class="panel-lead">
<div><span>X / SINGLE-ID BOUNDARY CONTROL</span><h4>只换 assistant 后面的一个 ID:EOS 不是普通占位符</h4></div>
<p>
重复词元历史、角色标记、长度、目标绝对位置与 32-row batch 全部不动;
只把官方 EOS `100001` 分别换成单 token 的 x、句点或换行。
</p>
</div>
<div class="history-ledger boundary-ledger">
<article><span>SOURCE PROMPTS</span><b>128</b><p>四个公开域各 32 条;同一 cohort</p></article>
<article><span>2×4 VARIANTS</span><b>1,024</b><p>system 0/1 × 四种边界 ID</p></article>
<article><span>INPUT TOKENS</span><b>56,784</b><p>八格逐组完全同长度</p></article>
<article><span>REAL ROUTES</span><b>2,044,224</b><p>八格 × 前六个 MoE 层</p></article>
<article><span>ALIGNED TARGET</span><b>2,874 × 8</b><p>相同字符跨度、位置与 token ID</p></article>
<article class="exact"><span>INDEPENDENT RERUN</span><b>{boundaryExact ? "BYTE-EXACT" : "MISMATCH"}</b><p>完整 JSON SHA-256 9bb93834…b9c37</p></article>
</div>
<div class="boundary-protocol" aria-label="历史 assistant 边界的单 token ID 控制">
<div class="boundary-track">
<span>FIXED PREFIX</span>
<b>Assistant:</b>
<i>x</i>
<mark data-boundary-track-token>EOS</mark>
<b>User:</b>
<i>TARGET</i>
<span>FIXED SUFFIX</span>
</div>
<div class="boundary-token-grid">
<article class="official">
<span>OFFICIAL</span><b>EOS</b><code>ID 100001</code>
<p>官方 template 自然产生;唯一的合法序列格。</p>
</article>
<article>
<span>COUNTERFACTUAL A</span><b>x</b><code>ID 87</code>
<p>普通内容 token;相对官方序列只改一个 ID。</p>
</article>
<article>
<span>COUNTERFACTUAL B</span><b>.</b><code>ID 13</code>
<p>普通标点 token;长度与后续位置完全不变。</p>
</article>
<article>
<span>COUNTERFACTUAL C</span><b>↵</b><code>ID 185</code>
<p>普通换行 token;保留紧随其后的 `User:` 标记。</p>
</article>
</div>
<p>
每个 source 在 S0 / S1 内都通过 4 / 4 同长度、4 / 4 同目标位置;768 / 768
反事实格相对官方 token 序列恰好只改一个 ID。
</p>
</div>
<div class="boundary-controls">
<div>
<span>MOE LAYER</span>
<div class="layer-switch boundary-layer-switch" role="group" aria-label="选择边界 token 控制层">
{[1, 2, 3, 4, 5, 6].map((layer) => (
<button type="button" data-boundary-layer={layer} class={layer === 1 ? "active" : ""}>L{layer}</button>
))}
</div>
</div>
<div>
<span>MEASUREMENT SCOPE</span>
<div class="boundary-scope-switch" role="group" aria-label="选择边界 token 统计范围">
<button type="button" data-boundary-scope="target_content" aria-pressed="true">目标内容</button>
<button type="button" data-boundary-scope="full_input" aria-pressed="false">完整输入</button>
</div>
</div>
<div>
<span>AGGREGATION</span>
<div class="boundary-mode-switch" role="group" aria-label="选择边界 token 聚合口径">
<button type="button" data-boundary-mode="prompt_balanced" aria-pressed="true">prompt 等权</button>
<button type="button" data-boundary-mode="token_weighted" aria-pressed="false">token 加权</button>
</div>
</div>
<div>
<span>REPLACEMENT − EOS</span>
<div class="boundary-contrast-switch" role="group" aria-label="选择 EOS 的单 token 替换">
<button type="button" data-boundary-contrast="x_minus_eos" aria-pressed="true">x − EOS</button>
<button type="button" data-boundary-contrast="period_minus_eos" aria-pressed="false">. − EOS</button>
<button type="button" data-boundary-contrast="newline_minus_eos" aria-pressed="false">↵ − EOS</button>
</div>
</div>
<p data-boundary-note>
目标内容:八格只比较完全相同的后续内容 token;正 ΔTV 表示替换 EOS 后 system edge 更大,不表示能力更差。
</p>
</div>
<div class="boundary-domain-grid" data-boundary-domain-grid></div>
<div class="history-buffer-summary boundary-summary">
<article><span>X CONTROL</span><b>24 / 24 ↑</b><p>system-edge TV 全部高于 EOS;23 / 24 配对区间完全高于零。</p></article>
<article><span>PERIOD CONTROL</span><b>24 / 24 ↑</b><p>点估计全部高于 EOS;同样 23 / 24 区间完全高于零。</p></article>
<article><span>NEWLINE CONTROL</span><b>23 / 24 ↑</b><p>16 / 24 区间完全高于零;比另两个对照更依赖层与域。</p></article>
<article><span>MEAN TARGET TV</span><b>.037 → .054 / .055 / .049</b><p>EOS / x / 句点 / 换行;不是准确率或优劣排名。</p></article>
</div>
<div class="boundary-depth">
<div>
<span>DEPTH MAP / REPLACEMENT − EOS</span>
<h5 data-boundary-depth-title>x − EOS:只替换历史边界的一个 input ID</h5>
<p>红色为替换后 system-edge TV 更大,绿色为更小;每格使用八格共享 source-bootstrap。</p>
</div>
<div data-boundary-depth-map></div>
</div>
<div class="boundary-interpretation">
<article>
<span>WHAT IS CAUSAL</span>
<b>一个历史 input ID</b>
<p>在固定模型、目标、位置、mask 与 batch 内,EOS→control 的路由差异来自这一个输入干预。</p>
</article>
<article>
<span>WHAT REMAINS</span>
<b>`User:` 边界仍在</b>
<p>实验没有删除全部回合结构,只识别 EOS token identity;三个反事实也不是合法官方 chat。</p>
</article>
<article>
<span>CHECKPOINT BOUNDARY</span>
<b>BASE ≠ CHAT / SFT</b>
<p>不能把较小 TV 命名为“理解回合结束”;仍需 V2-Lite-Chat 与行为生成对照。</p>
</article>
</div>
<div class="evidence-links">
<a href="https://huggingface.co/deepseek-ai/DeepSeek-V2-Lite/blob/604d5664dddd88a0433dbae533b7fe9472482de0/tokenizer_config.json" rel="noreferrer">固定官方 tokenizer_config ↗</a>
<a href="https://huggingface.co/docs/transformers/en/chat_templating" rel="noreferrer">HF Chat Templates ↗</a>
<a href="https://arxiv.org/abs/2305.11206" rel="noreferrer">LIMA · EOT supervision ↗</a>
<a href="https://aclanthology.org/2020.blackboxnlp-1.26/" rel="noreferrer">EOS Decision ↗</a>
<a href="https://arxiv.org/abs/2309.16042" rel="noreferrer">Activation Patching Limits ↗</a>
</div>
<div class="artifact-boundary">
<b>SINGLE-ID ROUTING CAUSALITY, NOT TURN-SEMANTIC OR CAPABILITY PROOF</b>
<p>
EOS 条件下后续目标路由对 system 开关更稳定,但本实验既未生成答案,也未覆盖 Chat 权重;
更小 TV 不等于更正确。下一步要拆 `User:` 角色标记、special-token 家族与行为指标。
</p>
</div>
</section>
<section class="artifact-panel" data-artifact-panel="evidence" hidden>
<div class="panel-lead">
<div><span>O + X / EVIDENCE SLICE</span><h4>为什么执行到 layer 6 就停,而不是把“部分下载”写成“完整复现”</h4></div>
@@ -1261,7 +1409,9 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
<code>experiments/deepseek/v2_lite_routing_history_factorial_probe.py</code> ·
<code>research/DEEPSEEK_ROUTING_HISTORY_FACTORIAL_AUDIT.md</code> ·
<code>experiments/deepseek/v2_lite_routing_history_distance_control.py</code> ·
<code>research/DEEPSEEK_ROUTING_HISTORY_DISTANCE_CONTROL_AUDIT.md</code>
<code>research/DEEPSEEK_ROUTING_HISTORY_DISTANCE_CONTROL_AUDIT.md</code> ·
<code>experiments/deepseek/v2_lite_routing_history_boundary_token_control.py</code> ·
<code>research/DEEPSEEK_ROUTING_HISTORY_BOUNDARY_TOKEN_AUDIT.md</code>
</figcaption>
<script is:inline type="application/json" data-dsv2-trace set:html={compactJson}></script>
@@ -1269,6 +1419,7 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
<script is:inline type="application/json" data-dsv2-template set:html={templateCompactJson}></script>
<script is:inline type="application/json" data-dsv2-history set:html={historyCompactJson}></script>
<script is:inline type="application/json" data-dsv2-distance set:html={distanceCompactJson}></script>
<script is:inline type="application/json" data-dsv2-boundary set:html={boundaryCompactJson}></script>
</figure>
<script>
@@ -1284,18 +1435,21 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
const templateNode = one<HTMLScriptElement>("[data-dsv2-template]");
const historyNode = one<HTMLScriptElement>("[data-dsv2-history]");
const distanceNode = one<HTMLScriptElement>("[data-dsv2-distance]");
const boundaryNode = one<HTMLScriptElement>("[data-dsv2-boundary]");
if (
!payloadNode?.textContent
|| !corpusNode?.textContent
|| !templateNode?.textContent
|| !historyNode?.textContent
|| !distanceNode?.textContent
|| !boundaryNode?.textContent
) return;
const data = JSON.parse(payloadNode.textContent);
const corpusData = JSON.parse(corpusNode.textContent);
const templateData = JSON.parse(templateNode.textContent);
const historyData = JSON.parse(historyNode.textContent);
const distanceData = JSON.parse(distanceNode.textContent);
const boundaryData = JSON.parse(boundaryNode.textContent);
const tabs = all<HTMLButtonElement>("[data-artifact-tab]");
const panels = all<HTMLElement>("[data-artifact-panel]");
@@ -2188,6 +2342,172 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
});
});
renderDistance();
let boundaryLayerNumber = 1;
let boundaryScope = "target_content";
let boundaryMode = "prompt_balanced";
let boundaryContrast = "x_minus_eos";
const boundaryContrastLabels: Record<string, string> = {
x_minus_eos: "X − EOS",
period_minus_eos: "PERIOD − EOS",
newline_minus_eos: "NEWLINE − EOS",
};
const boundaryLevelLabels: Record<string, string> = {
eos: "EOS",
x: "X",
period: ".",
newline: "↵",
};
const boundaryReplacement = () => boundaryContrast.replace("_minus_eos", "");
const renderBoundary = () => {
all<HTMLButtonElement>("[data-boundary-layer]").forEach((button) => {
button.classList.toggle(
"active",
Number(button.dataset.boundaryLayer) === boundaryLayerNumber,
);
});
all<HTMLButtonElement>("[data-boundary-scope]").forEach((button) => {
button.setAttribute(
"aria-pressed",
String(button.dataset.boundaryScope === boundaryScope),
);
});
all<HTMLButtonElement>("[data-boundary-mode]").forEach((button) => {
button.setAttribute(
"aria-pressed",
String(button.dataset.boundaryMode === boundaryMode),
);
});
all<HTMLButtonElement>("[data-boundary-contrast]").forEach((button) => {
button.setAttribute(
"aria-pressed",
String(button.dataset.boundaryContrast === boundaryContrast),
);
});
set(
"[data-boundary-note]",
boundaryScope === "target_content"
? "目标内容:八格只比较完全相同的后续内容 token;正 ΔTV 表示替换 EOS 后 system edge 更大,不表示能力更差。"
: "完整输入:system、历史、被替换边界、角色包装与目标全部计入;一个 ID 的后续效应会被整段协议流量稀释。",
);
const replacement = boundaryReplacement();
set(
"[data-boundary-track-token]",
boundaryLevelLabels[replacement],
);
const currentLayer = boundaryData.layers.find(
(item: any) => item.layer === boundaryLayerNumber,
);
const view = currentLayer.scopes[boundaryScope].modes[boundaryMode];
const grid = one<HTMLElement>("[data-boundary-domain-grid]");
if (grid) {
grid.replaceChildren(...boundaryData.domains.map((domain: string) => {
const data = view[domain];
const contrast = data.contrasts[boundaryContrast].total_variation_delta;
const direct = data.direct[replacement];
const interaction = direct.s1_minus_s0.total_variation_delta;
const eosAlignment = currentLayer.alignment[domain].system_eos;
const replacementAlignment = currentLayer.alignment[domain][`system_${replacement}`];
const card = document.createElement("article");
const label = document.createElement("span");
const ladder = document.createElement("div");
const primary = document.createElement("strong");
const ci = document.createElement("p");
const stability = document.createElement("small");
const substitution = document.createElement("em");
const interactionLine = document.createElement("u");
const cv = document.createElement("i");
label.textContent = corpusLabels[domain];
ladder.className = "boundary-tv-ladder";
["eos", "x", "period", "newline"].forEach((boundary) => {
const cell = document.createElement("b");
const name = document.createElement("small");
const value = document.createElement("strong");
name.textContent = boundaryLevelLabels[boundary];
value.textContent = data.distances[boundary].total_variation.point.toFixed(3);
cell.classList.toggle("selected", boundary === replacement);
cell.classList.toggle("official", boundary === "eos");
cell.append(name, value);
ladder.append(cell);
});
primary.textContent = `${boundaryContrastLabels[boundaryContrast]} · ΔTV ${signed(contrast.point)}`;
primary.className = deltaClass(contrast.ci95);
ci.textContent = `source-paired 95% ${formatSignedCi(contrast.ci95)}`;
stability.textContent = `system top-6 set exact EOS ${(eosAlignment.setExactRate * 100).toFixed(1)}% · ${boundaryLevelLabels[replacement]} ${(replacementAlignment.setExactRate * 100).toFixed(1)}% · J ${eosAlignment.meanJaccard.toFixed(3)} → ${replacementAlignment.meanJaccard.toFixed(3)}`;
substitution.textContent = `direct EOS↔${boundaryLevelLabels[replacement]} TV · S0 ${direct.at_s0.total_variation.point.toFixed(3)} · S1 ${direct.at_s1.total_variation.point.toFixed(3)}`;
interactionLine.textContent = `direct S1−S0 ${signed(interaction.point)} · 95% ${formatSignedCi(interaction.ci95)}`;
const cvContrast = data.cvContrasts[boundaryContrast];
cv.textContent = `system-edge ΔCV ${signed(cvContrast.point)} · 95% ${formatSignedCi(cvContrast.ci95)}`;
card.append(
label,
ladder,
primary,
ci,
stability,
substitution,
interactionLine,
cv,
);
return card;
}));
}
const titles: Record<string, string> = {
x_minus_eos: "x − EOS:普通内容 token 替换官方边界",
period_minus_eos: "句点 − EOS:普通标点 token 替换官方边界",
newline_minus_eos: "换行 − EOS:普通格式 token 替换官方边界",
};
set("[data-boundary-depth-title]", titles[boundaryContrast]);
const depth = one<HTMLElement>("[data-boundary-depth-map]");
if (depth) {
depth.replaceChildren(...boundaryData.domains.map((domain: string) => {
const row = document.createElement("div");
const label = document.createElement("b");
label.textContent = corpusLabels[domain];
row.append(label);
boundaryData.layers.forEach((layer: any) => {
const contrast = layer.scopes[boundaryScope].modes[boundaryMode]
[domain].contrasts[boundaryContrast].total_variation_delta;
const cell = document.createElement("span");
cell.className = deltaClass(contrast.ci95);
cell.style.setProperty(
"--strength",
String(Math.min(1, Math.abs(contrast.point) / 0.04)),
);
cell.textContent = `L${layer.layer} ${signed(contrast.point)}`;
cell.title = `${corpusLabels[domain]} · L${layer.layer} · ${boundaryContrastLabels[boundaryContrast]} Δ system-edge TV ${signed(contrast.point)} · paired 95% ${formatSignedCi(contrast.ci95)}`;
row.append(cell);
});
return row;
}));
}
};
all<HTMLButtonElement>("[data-boundary-layer]").forEach((button) => {
button.addEventListener("click", () => {
boundaryLayerNumber = Number(button.dataset.boundaryLayer);
renderBoundary();
});
});
all<HTMLButtonElement>("[data-boundary-scope]").forEach((button) => {
button.addEventListener("click", () => {
boundaryScope = button.dataset.boundaryScope ?? "target_content";
renderBoundary();
});
});
all<HTMLButtonElement>("[data-boundary-mode]").forEach((button) => {
button.addEventListener("click", () => {
boundaryMode = button.dataset.boundaryMode ?? "prompt_balanced";
renderBoundary();
});
});
all<HTMLButtonElement>("[data-boundary-contrast]").forEach((button) => {
button.addEventListener("click", () => {
boundaryContrast = button.dataset.boundaryContrast ?? "x_minus_eos";
renderBoundary();
});
});
renderBoundary();
});
</script>
@@ -2309,7 +2629,7 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
.artifact-status b { color: var(--ink); font-size: .72rem; }
.artifact-tabs {
display: grid;
grid-template-columns: repeat(9, 1fr);
grid-template-columns: repeat(10, 1fr);
background: var(--ink);
}
.artifact-tabs button {
@@ -3691,6 +4011,281 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
font-size: .62rem;
line-height: 1.45;
}
.boundary-protocol {
margin-top: .8rem;
padding: 1rem;
border: 1px solid rgba(32,32,39,.15);
background:
linear-gradient(115deg, rgba(98,105,155,.09), transparent 50%),
#fffdf8;
}
.boundary-track {
display: flex;
align-items: center;
justify-content: center;
flex-wrap: wrap;
gap: .25rem;
padding: .8rem;
background: var(--ink);
color: white;
}
.boundary-track > * {
padding: .36rem .48rem;
font: 650 .61rem/1 var(--font-mono);
font-style: normal;
}
.boundary-track span {
color: rgba(255,255,255,.48);
font-size: .52rem;
}
.boundary-track b { background: rgba(255,255,255,.1); }
.boundary-track i { color: #d5d8ef; }
.boundary-track mark {
min-width: 3.6rem;
background: var(--amber);
color: white;
text-align: center;
}
.boundary-token-grid {
display: grid;
grid-template-columns: repeat(4, 1fr);
margin-top: .55rem;
border: 1px solid rgba(32,32,39,.12);
}
.boundary-token-grid article {
display: grid;
gap: .34rem;
padding: .75rem;
border-right: 1px solid rgba(32,32,39,.12);
background: #f3eee5;
}
.boundary-token-grid article:last-child { border-right: 0; }
.boundary-token-grid article.official {
background: rgba(57,120,110,.11);
}
.boundary-token-grid span {
color: var(--blue);
font: 700 .52rem/1 var(--font-mono);
}
.boundary-token-grid b {
font: 760 1.05rem/1 var(--font-display);
}
.boundary-token-grid code {
width: max-content;
padding: .22rem .3rem;
background: rgba(32,32,39,.08);
font-size: .57rem;
}
.boundary-token-grid p,
.boundary-protocol > p {
margin: 0;
color: rgba(32,32,39,.58);
font-size: .61rem;
line-height: 1.45;
}
.boundary-protocol > p {
margin-top: .7rem;
text-align: center;
}
.boundary-controls {
display: grid;
grid-template-columns: auto .9fr .9fr 1.55fr;
gap: .8rem;
align-items: end;
margin-top: .8rem;
padding: .85rem;
border: 1px solid rgba(32,32,39,.14);
background: #e8e2d7;
}
.boundary-controls > div { display: grid; gap: .45rem; }
.boundary-controls .layer-switch { margin: 0; }
.boundary-scope-switch,
.boundary-mode-switch,
.boundary-contrast-switch { display: flex; }
.boundary-scope-switch button,
.boundary-mode-switch button,
.boundary-contrast-switch button {
padding: .58rem .66rem;
border: 1px solid rgba(32,32,39,.22);
background: #fffdf8;
color: var(--ink);
font: 650 .6rem/1 var(--font-mono);
cursor: pointer;
}
.boundary-scope-switch button + button,
.boundary-mode-switch button + button,
.boundary-contrast-switch button + button { border-left: 0; }
.boundary-scope-switch button[aria-pressed="true"],
.boundary-mode-switch button[aria-pressed="true"],
.boundary-contrast-switch button[aria-pressed="true"] {
border-color: var(--blue);
background: var(--blue);
color: white;
}
.boundary-controls > p {
grid-column: 1 / -1;
margin: 0;
padding-top: .75rem;
border-top: 1px solid rgba(32,32,39,.12);
color: rgba(32,32,39,.62);
font-size: .69rem;
line-height: 1.5;
}
.boundary-domain-grid {
display: grid;
grid-template-columns: repeat(4, 1fr);
margin-top: .8rem;
border: 1px solid rgba(32,32,39,.14);
background: #fffdf8;
}
.boundary-domain-grid > :global(article) {
min-width: 0;
padding: .85rem;
border-right: 1px solid rgba(32,32,39,.12);
}
.boundary-domain-grid > :global(article:last-child) { border-right: 0; }
.boundary-domain-grid :global(.boundary-tv-ladder) {
display: grid;
grid-template-columns: repeat(4, 1fr);
gap: .18rem;
margin-top: .55rem;
}
.boundary-domain-grid :global(.boundary-tv-ladder > b) {
display: grid;
gap: .18rem;
min-width: 0;
padding: .34rem;
background: #e8e2d7;
}
.boundary-domain-grid :global(.boundary-tv-ladder > b.official) {
background: rgba(57,120,110,.12);
}
.boundary-domain-grid :global(.boundary-tv-ladder > b.selected) {
outline: 1px solid var(--blue);
outline-offset: -1px;
}
.boundary-domain-grid :global(.boundary-tv-ladder small) {
color: rgba(32,32,39,.5);
font: 650 .46rem/1 var(--font-mono);
}
.boundary-domain-grid :global(.boundary-tv-ladder strong) {
white-space: nowrap;
font: 720 .58rem/1 var(--font-mono);
letter-spacing: -.025em;
}
.boundary-domain-grid > :global(article > strong) {
display: inline-block;
margin-top: .48rem;
padding: .26rem .38rem;
font: 750 .62rem/1 var(--font-mono);
}
.boundary-domain-grid > :global(article > strong.down),
.boundary-depth :global(span.down) {
background: rgba(57,120,110,.13);
color: var(--teal);
}
.boundary-domain-grid > :global(article > strong.up),
.boundary-depth :global(span.up) {
background: rgba(161,77,77,.12);
color: var(--red);
}
.boundary-domain-grid > :global(article > strong.neutral),
.boundary-depth :global(span.neutral) {
background: rgba(186,118,44,.12);
color: var(--amber);
}
.boundary-domain-grid > :global(article > p),
.boundary-domain-grid > :global(article > small),
.boundary-domain-grid > :global(article > em),
.boundary-domain-grid > :global(article > u),
.boundary-domain-grid > :global(article > i) {
display: block;
margin: .38rem 0 0;
color: rgba(32,32,39,.57);
overflow-wrap: anywhere;
font: .54rem/1.4 var(--font-mono);
font-style: normal;
text-decoration: none;
}
.boundary-domain-grid > :global(article > em),
.boundary-domain-grid > :global(article > u),
.boundary-domain-grid > :global(article > i) {
padding-top: .34rem;
border-top: 1px solid rgba(32,32,39,.1);
}
.boundary-depth {
display: grid;
grid-template-columns: .52fr 1.48fr;
gap: 1rem;
margin-top: .8rem;
padding: 1rem;
border: 1px solid rgba(32,32,39,.14);
}
.boundary-depth h5 {
margin: .4rem 0;
font: 720 1rem/1.15 var(--font-display);
}
.boundary-depth p {
margin: 0;
color: rgba(32,32,39,.58);
font-size: .66rem;
line-height: 1.5;
}
.boundary-depth > :global([data-boundary-depth-map]) {
display: grid;
gap: .35rem;
}
.boundary-depth :global([data-boundary-depth-map] > div) {
display: grid;
grid-template-columns: 5.5rem repeat(6, 1fr);
gap: .25rem;
}
.boundary-depth :global([data-boundary-depth-map] > div > b),
.boundary-depth :global([data-boundary-depth-map] > div > span) {
display: grid;
align-items: center;
min-height: 2.2rem;
padding: .35rem;
font: 650 .55rem/1.2 var(--font-mono);
}
.boundary-depth :global([data-boundary-depth-map] > div > b) {
color: var(--blue);
}
.boundary-depth :global([data-boundary-depth-map] > div > span.down) {
background: color-mix(in srgb, var(--teal) calc(var(--strength) * 55%), #eef0e9);
color: var(--ink);
}
.boundary-depth :global([data-boundary-depth-map] > div > span.up) {
background: color-mix(in srgb, var(--red) calc(var(--strength) * 48%), #f3ebe6);
color: var(--ink);
}
.boundary-depth :global([data-boundary-depth-map] > div > span.neutral) {
background: rgba(186,118,44,.1);
color: var(--ink);
}
.boundary-interpretation {
display: grid;
grid-template-columns: repeat(3, 1fr);
margin-top: .8rem;
border: 1px solid rgba(32,32,39,.14);
background: #e8e2d7;
}
.boundary-interpretation article {
padding: .9rem;
border-right: 1px solid rgba(32,32,39,.12);
}
.boundary-interpretation article:last-child { border-right: 0; }
.boundary-interpretation b {
display: block;
margin-top: .4rem;
font: 730 .78rem/1.2 var(--font-display);
}
.boundary-interpretation p {
margin: .4rem 0 0;
color: rgba(32,32,39,.58);
font-size: .62rem;
line-height: 1.45;
}
.observed-cache {
display: grid;
grid-template-columns: 1fr auto 1.25fr;
@@ -3934,7 +4529,9 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
.history-buffer,
.history-depth,
.distance-controls,
.distance-depth { grid-template-columns: 1fr; }
.distance-depth,
.boundary-controls,
.boundary-depth { grid-template-columns: 1fr; }
.artifact-status { grid-template-columns: 1fr 1fr; }
.artifact-tabs { grid-template-columns: 1fr 1fr; }
.route-controls { grid-template-columns: 1fr 1fr; }
@@ -3951,9 +4548,13 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
.history-ledger { grid-template-columns: repeat(3, 1fr); }
.history-domain-grid,
.history-buffer-summary,
.distance-domain-grid { grid-template-columns: 1fr 1fr; }
.distance-interpretation { grid-template-columns: 1fr; }
.distance-interpretation article { border-right: 0; border-bottom: 1px solid rgba(32,32,39,.12); }
.distance-domain-grid,
.boundary-domain-grid,
.boundary-token-grid { grid-template-columns: 1fr 1fr; }
.distance-interpretation,
.boundary-interpretation { grid-template-columns: 1fr; }
.distance-interpretation article,
.boundary-interpretation article { border-right: 0; border-bottom: 1px solid rgba(32,32,39,.12); }
.template-protocol > i { transform: rotate(90deg); justify-self: center; }
.length-delta-grid { grid-template-columns: 1fr 1fr; }
.corpus-heat-head p { text-align: left; }
@@ -3994,7 +4595,10 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
.history-domain-grid,
.history-buffer-summary,
.distance-domain-grid,
.distance-interpretation { grid-template-columns: 1fr; }
.distance-interpretation,
.boundary-domain-grid,
.boundary-token-grid,
.boundary-interpretation { grid-template-columns: 1fr; }
.route-metrics article,
.cache-ratio article,
.load-lessons article,
@@ -4010,7 +4614,10 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
.history-domain-grid > :global(article),
.history-buffer-summary article,
.distance-domain-grid > :global(article),
.distance-interpretation article { border-right: 0; border-bottom: 1px solid rgba(32,32,39,.12); }
.distance-interpretation article,
.boundary-domain-grid > :global(article),
.boundary-token-grid article,
.boundary-interpretation article { border-right: 0; border-bottom: 1px solid rgba(32,32,39,.12); }
.corpus-mode-switch,
.corpus-cohort-switch,
.template-scope-switch,
@@ -4020,7 +4627,10 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
.history-effect-switch,
.distance-scope-switch,
.distance-mode-switch,
.distance-contrast-switch { display: grid; grid-template-columns: 1fr; }
.distance-contrast-switch,
.boundary-scope-switch,
.boundary-mode-switch,
.boundary-contrast-switch { display: grid; grid-template-columns: 1fr; }
.corpus-mode-switch button + button,
.corpus-cohort-switch button + button,
.template-scope-switch button + button,
@@ -4030,7 +4640,10 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
.history-effect-switch button + button,
.distance-scope-switch button + button,
.distance-mode-switch button + button,
.distance-contrast-switch button + button { border-left: 1px solid rgba(32,32,39,.22); border-top: 0; }
.distance-contrast-switch button + button,
.boundary-scope-switch button + button,
.boundary-mode-switch button + button,
.boundary-contrast-switch button + button { border-left: 1px solid rgba(32,32,39,.22); border-top: 0; }
.length-delta-grid > :global(article),
.length-pair-summary article { border-right: 0; border-bottom: 1px solid rgba(32,32,39,.11); }
.artifact-boundary { grid-template-columns: 1fr; }
@@ -4050,6 +4663,8 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
.history-buffer > :global([data-history-buffer-grid]) { grid-template-columns: 1fr; }
.history-depth { overflow-x: auto; }
.history-depth > :global([data-history-depth-map]) { min-width: 620px; }
.boundary-depth { overflow-x: auto; }
.boundary-depth > :global([data-boundary-depth-map]) { min-width: 620px; }
.layer-evidence { grid-template-columns: repeat(7, 1fr); }
.repro-gate { grid-template-columns: 1fr; }
.repro-gate > p { grid-column: auto; }
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+5 -5
View File
@@ -37,7 +37,7 @@ const toc = [
<BaseLayout
title="DeepSeek 技术谱系与真实权重深读:从 Dense、MoE、MLA 到 R1 与 V4"
description="用二十四张问题账、十次技术转向、十三个交互实验、真实 V2-Lite 权重、公开语料路由区间、官方模板、消息历史因子与等长 filler 控制、吸收式缓存 trace 和六十个一手节点,完整理解 DeepSeek 的 MoE、MLA、FP8、DualPipe、GRPO、R1、V3.2 与 V4。"
description="用二十四张问题账、十次技术转向、十四个交互实验、真实 V2-Lite 权重、公开语料路由区间、官方模板、消息历史、等长 filler 与单 EOS-ID 边界控制、吸收式缓存 trace 和六十个一手节点,完整理解 DeepSeek 的 MoE、MLA、FP8、DualPipe、GRPO、R1、V3.2 与 V4。"
section="deepseek"
>
<header class="page-hero deepseek-hero">
@@ -55,7 +55,7 @@ const toc = [
<div><dt>SPAN</dt><dd>2024.01 → 2026.06</dd></div>
<div><dt>LEDGERS</dt><dd>24 张问题账</dd></div>
<div><dt>LINEAGE</dt><dd>10 次技术转向</dd></div>
<div><dt>LABS</dt><dd>13 个可操作实验</dd></div>
<div><dt>LABS</dt><dd>14 个可操作实验</dd></div>
<div><dt>EVIDENCE</dt><dd>60 个一手 / 官方节点</dd></div>
<div><dt>STATUS</dt><dd>三轮 · 真实权重执行</dd></div>
</dl>
@@ -768,15 +768,15 @@ const toc = [
<p class="eyebrow"><span>22</span> OFFICIAL WEIGHTS / EXECUTED</p>
<h2>从“MLA 与 MoE 的概念”再往前一步:让官方 V2-Lite 权重真的跑起来</h2>
<p class="lede">
前面的四联实验负责建立公式与角色合同;下面的九联工件实验固定官方 revision、tokenizer、
前面的四联实验负责建立公式与角色合同;下面的十联工件实验固定官方 revision、tokenizer、
模型代码和 checkpoint 第一分片,在 RTX 5090 上连续执行 layer 0–6。它把真实观测、shape 推导、
吸收式 latent cache、长度对照、官方 chat-template 扰动、实现差距和未覆盖范围放在同一张证据图里。
</p>
<div class="artifact-callout">
<article><span>X / FORWARD</span><b>7 / 27 layers</b><p>1 个 dense 层 + 6 个 MoE 层;layer 7 因跨分片停止。</p></article>
<article><span>X / ROUTES</span><b>3,112,848</b><p>三档长度、模板、system × one-shot 与 none/filler/demo 六格的真实 top-6 选择。</p></article>
<article><span>X / ROUTES</span><b>5,157,072</b><p>三档长度、模板、消息历史、等长 filler 与单 EOS-ID 八格控制的真实 top-6 选择。</p></article>
<article><span>X / ABSORB CACHE</span><b>266,240 → 29,952 B</b><p>同一真实 layer-1 权重的 naive / absorb active buffers。</p></article>
<article><span>X / RERUN</span><b>6 / 6 EXACT</b><p>三档长度、官方模板、历史因子与等长 filler 控制均 byte-exact;比较使用 paired prompt bootstrap。</p></article>
<article><span>X / RERUN</span><b>7 / 7 EXACT</b><p>三档长度、官方模板、历史因子、等长 filler 与边界 ID 控制均 byte-exact;比较使用 paired prompt bootstrap。</p></article>
</div>
<DeepSeekArtifactLab />
</section>
+7 -5
View File
@@ -15,7 +15,7 @@ const workstreams = [
{ label: "表示、位置与残差高速公路", value: 81, next: "加入真实 hidden-state / norm traces、长上下文位置外推复现与更多深层稳定性消融" },
{ label: "Scaling Laws", value: 74, next: "加入真实拟合复现、置信区间与更多模型族对照" },
{ label: "数据工程与预训练配方", value: 73, next: "逐图精读 FineWeb / DCLM,加入真实去重与 mixture traces" },
{ label: "DeepSeek 专题", value: 95, next: "SM90 FlashMLA kernel、完整 27 层、EOS / 角色 / 多 filler / 内容与 batch-shape 控制、FP8/pipeline 与 R1-like RL 复现" },
{ label: "DeepSeek 专题", value: 96, next: "角色标记与 special-token family、V2-Lite-Chat 行为、完整 27 层与固定 batch shape,再推进 SM90 FlashMLA、FP8/pipeline 与 R1-like RL" },
{ label: "指令微调与人类偏好", value: 75, next: "加入真实偏好分歧样本、RM 长度偏置与 PPO/DPO 小模型复现" },
{ label: "推理与测试时扩展", value: 76, next: "真实模型采样曲线、PRM 案例与逐篇图表精读" },
{ label: "工具使用与长程 Agent", value: 74, next: "补真实环境 traces、cross-harness 对照、Agent RL 训练曲线与安全案例" },
@@ -50,7 +50,7 @@ const workstreams = [
<div><dt>OVERALL</dt><dd>专题平均 {average}%</dd></div>
<div><dt>READABLE</dt><dd>{published} 个首版可读专题</dd></div>
<div><dt>ACTIVE</dt><dd>{researching} 个研究/写作中</dd></div>
<div><dt>UPDATED</dt><dd>2026-07-29 19:05 CST</dd></div>
<div><dt>UPDATED</dt><dd>2026-07-29 19:55 CST</dd></div>
<div><dt>MODE</dt><dd>持续迭代,不锁死版本</dd></div>
</dl>
</div>
@@ -97,12 +97,12 @@ const workstreams = [
<article><span>✓</span><h3>K3 报告已结构化拆解</h3><p>47 页报告目录、151 条参考来源和架构/后训练/系统主线已经提取。</p></article>
<article><span>✓</span><h3>17 专题知识图</h3><p>从语言模型基础到评测安全,包含先修依赖和三条贯穿案例。</p></article>
<article><span>✓</span><h3>编辑式网站系统</h3><p>响应式导航、章节模板、侧栏、进度、论文链和证据提示组件。</p></article>
<article><span>✓</span><h3>八十个原创交互视图</h3><p>K3 三轴图、八联报告实验与四联开放工件实验,DeepSeek 四联公式实验与九联真实权重实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
<article><span>✓</span><h3>八十一个原创交互视图</h3><p>K3 三轴图、八联报告实验与四联开放工件实验,DeepSeek 四联公式实验与十联真实权重实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
<article><span>✓</span><h3>十七篇首版长文</h3><p>K3、语言模型前史、Transformer、表示/位置/残差、DeepSeek、Scaling、数据工程、长上下文、MoE、后训练、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全专题。</p></article>
<article><span>✓</span><h3>语言模型前史深度专题</h3><p>八张独立问题账、33 个正式节点、20 段长文与概率—向量—记忆—对齐四联实验。</p></article>
<article><span>✓</span><h3>Transformer 深度专题</h3><p>十张独立问题账、40 个正式节点、21 段正文与 QKV—Mask—多头位置—Block 成本四联实验。</p></article>
<article><span>✓</span><h3>表示、位置与残差高速公路深度专题</h3><p>二十张问题账、66 个一手节点、DeepSeek/Kimi 双谱系,以及 Token—位置—Norm—Residual/FFN 四联实验。</p></article>
<article><span>✓</span><h3>DeepSeek 三轮真实权重里程碑</h3><p>在二十四张问题账、十次转向与四联公式实验上,新增 V2-Lite 7/27 层连续 forward、官方 V3 absorb、长度/模板、system × one-shot 与等长 filler 控制;累计 3,112,848 次真实路由,none→filler→demo 的两个 TV 台阶均在 24 / 24 格下降,六份运行结果均 byte-exact 独立复跑。</p></article>
<article><span>✓</span><h3>DeepSeek 三轮真实权重里程碑</h3><p>在二十四张问题账、十次转向与四联公式实验上,新增 V2-Lite 7/27 层连续 forward、官方 V3 absorb、长度/模板、system × one-shot、等长 filler 与 EOS 单 token 边界控制;累计 5,157,072 次真实路由。最新八格实验固定长度、位置、角色与 batch,仅替换一个 input ID,两份 54,254,445-byte JSON 的 SHA-256 同为 9bb93834…b9c37。</p></article>
<article><span>✓</span><h3>Kimi K3 技术报告二轮深读</h3><p>三十二张问题账、Figure 1–16 / Table 1–5 审计、100 节点阅读链,以及 Delta—Decay—AttnRes—LatentMoE—SiTU—QB—MOPD—Cache 八联实验。</p></article>
<article><span>✓</span><h3>Kimi K3 三轮开放工件里程碑</h3><p>固定官方 revisions,审计 96 个 shards、497,220 个 tensor entries 与真实 KDA / MLA / MoE / MoonViT shapes;四联实验分开显示层型、tensor anatomy、参数范围和复现边界。</p></article>
<article><span>✓</span><h3>FlashKDA RTX 5090 执行闸门</h3><p>隔离 CUDA 13.0 / glibc 2.39 编译 sm_120a wheel;6/6 官方参考逐元素相等,并完成 fixed / varlen、三种 state mode 的 1,800 个 CUDA Event samples。</p></article>
@@ -134,7 +134,7 @@ const workstreams = [
<div class="queue-table">
<div class="head"><b>优先级</b><b>专题</b><b>本轮交付</b><b>完成闸门</b></div>
<div><span>P0</span><strong>K3 三轮</strong><p>开放权重 traces → FlashKDA / AttnRes / MoE 真实行为 → Figure 1–16 数值重绘与独立复现</p><em>运行证据 + 逐图复现</em></div>
<div><span>P0</span><strong>DeepSeek 三轮</strong><p>SM90 FlashMLA kernel / 完整 27 层 / EOS、角色、多 filler、示例内容与 batch shape 控制 → FP8 / pipeline traces → R1-like RL 小模型复现</p><em>运行证据 + 独立复现</em></div>
<div><span>P0</span><strong>DeepSeek 三轮</strong><p>角色标记 / special-token family / V2-Lite-Chat 行为 → 完整 27 层与固定 batch shape → SM90 FlashMLA / FP8 / pipeline traces → R1-like RL 小模型复现</p><em>运行证据 + 独立复现</em></div>
<div><span>P0</span><strong>Transformer 二轮</strong><p>多头电路逐图 → Pre/Post-LN 真实 traces → Flash/KV 配置与 kernel 对照</p><em>逐图笔记 + 实测边界</em></div>
<div><span>P0</span><strong>表示、位置与残差二轮</strong><p>真实 hidden-state / norm traces → 长上下文位置外推 → mHC / AttnRes 深层稳定性消融</p><em>可复现实验 + 逐图笔记</em></div>
<div><span>P0</span><strong>语言模型前史二轮</strong><p>Kneser–Ney / LSTM / Bahdanau 逐图 → 真实小语料复现 → tokenizer 公平性</p><em>可复现实验 + 逐图笔记</em></div>
@@ -218,6 +218,8 @@ const workstreams = [
<div><time>2026-07-29</time><b>32-token 对照改为同源 16→24</b><p>TNEWS 只有 105/10,000 条达到 32 tokens,强行统一会落入约 1% 极端长尾;24-token eligibility 仍保留 1,609 条中文候选。</p></div>
<div><time>2026-07-29</time><b>长度敏感性必须成对重采样</b><p>16-token 输入严格是 24-token 输入前缀,2,000 次 bootstrap 共用 prompt indices;结果只描述固定 cohort 的长度敏感性。</p></div>
<div><time>2026-07-29</time><b>三类 cohort 永久分身份</b><p>自然长度回答本批样本如何路由;matched-16 / 24 回答同一 prompt 多看 8 tokens 后如何变化,不把二者混成内容因果。</p></div>
<div><time>2026-07-29</time><b>历史边界用单 ID 替换而非删除</b><p>EOS→x / 句点 / 换行保持长度、目标位置、角色标记、mask 与同一 batch;识别一个输入 ID 的干预,不冒充移除了全部回合结构。</p></div>
<div><time>2026-07-29</time><b>base、Chat 与行为永久分层</b><p>V2-Lite base 的路由 TV 不是回合理解或能力指标;`User:` 仍在,反事实不是官方合法 chat,后续另跑 Chat checkpoint 与生成指标。</p></div>
</div>
</section>