feat: add DeepSeek Chat behavior evidence
This commit is contained in:
+12
-3
@@ -14,7 +14,7 @@
|
|||||||
| 表示、位置与残差高速公路 | 完成首版 | 81% | 真实 hidden-state / norm traces、长上下文位置外推与深层稳定性消融 |
|
| 表示、位置与残差高速公路 | 完成首版 | 81% | 真实 hidden-state / norm traces、长上下文位置外推与深层稳定性消融 |
|
||||||
| Scaling Laws | 完成首版 | 74% | 真实拟合复现、置信区间与更多模型族对照 |
|
| Scaling Laws | 完成首版 | 74% | 真实拟合复现、置信区间与更多模型族对照 |
|
||||||
| 数据工程与预训练配方 | 完成首版 | 73% | FineWeb / DCLM 逐图精读、真实去重误伤与 mixture traces |
|
| 数据工程与预训练配方 | 完成首版 | 73% | FineWeb / DCLM 逐图精读、真实去重误伤与 mixture traces |
|
||||||
| DeepSeek 专题 | 三轮实证进行中 | 98% | V2-Lite-Chat 生成/行为对照、完整 27 层与固定 batch-content 对照,再推进 SM90 FlashMLA、FP8/pipeline 与 R1-like RL 复现 |
|
| DeepSeek 专题 | 四轮实证进行中 | 98% | completion-aware 生成评测、完整 27 层与固定 batch-content 对照,再推进 SM90 FlashMLA、FP8/pipeline 与 R1-like RL 复现 |
|
||||||
| 指令微调与人类偏好 | 完成首版 | 75% | 真实偏好分歧、RM 长度偏置与 PPO/DPO 小模型复现 |
|
| 指令微调与人类偏好 | 完成首版 | 75% | 真实偏好分歧、RM 长度偏置与 PPO/DPO 小模型复现 |
|
||||||
| 推理与测试时扩展 | 完成首版 | 76% | 真实模型采样曲线、PRM 案例与逐篇图表精读 |
|
| 推理与测试时扩展 | 完成首版 | 76% | 真实模型采样曲线、PRM 案例与逐篇图表精读 |
|
||||||
| 工具使用与长程 Agent | 完成首版 | 74% | 真实环境 traces、cross-harness 对照、Agent RL 曲线与安全案例 |
|
| 工具使用与长程 Agent | 完成首版 | 74% | 真实环境 traces、cross-harness 对照、Agent RL 曲线与安全案例 |
|
||||||
@@ -41,7 +41,7 @@
|
|||||||
- [x] 完成 486 篇关键论文索引,覆盖 16 个标签专题与 Kimi/DeepSeek 聚光主线。
|
- [x] 完成 486 篇关键论文索引,覆盖 16 个标签专题与 Kimi/DeepSeek 聚光主线。
|
||||||
- [x] 完成可检索、可按专题筛选的论文库页面。
|
- [x] 完成可检索、可按专题筛选的论文库页面。
|
||||||
- [x] 完成 K3、语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全十七篇首版长文。
|
- [x] 完成 K3、语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全十七篇首版长文。
|
||||||
- [x] 完成 K3 三轴架构、八联报告实验与四联开放工件实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 十七联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等八十四个原创交互视图。
|
- [x] 完成 K3 三轴架构、八联报告实验与四联开放工件实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 十八联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等八十五个原创交互视图。
|
||||||
- [x] 完成长上下文首版:五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。
|
- [x] 完成长上下文首版:五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。
|
||||||
- [x] 核验 FlashAttention、DeepSeek-V2/V3.2/V4、Kimi Linear/K3 等六份论文原文,并建立长上下文研究账本。
|
- [x] 核验 FlashAttention、DeepSeek-V2/V3.2/V4、Kimi Linear/K3 等六份论文原文,并建立长上下文研究账本。
|
||||||
- [x] 核验 Switch、ST-MoE、DeepSeekMoE、Loss-Free、V3、LatentMoE 与 K3 原文,并建立 MoE 研究账本。
|
- [x] 核验 Switch、ST-MoE、DeepSeekMoE、Loss-Free、V3、LatentMoE 与 K3 原文,并建立 MoE 研究账本。
|
||||||
@@ -234,11 +234,17 @@
|
|||||||
- [x] 完整角色块正式运行与独立复跑各 66,975,110 bytes、SHA-256 均为 `a703dddb…7e82`,byte-exact;新增 2,044,224 次真实路由,使累计达到 11,289,744 次。两组 compact 产物分别约 1.4 / 3.3 MiB,累计正式实验集达到 10 / 10 byte-exact。
|
- [x] 完整角色块正式运行与独立复跑各 66,975,110 bytes、SHA-256 均为 `a703dddb…7e82`,byte-exact;新增 2,044,224 次真实路由,使累计达到 11,289,744 次。两组 compact 产物分别约 1.4 / 3.3 MiB,累计正式实验集达到 10 / 10 byte-exact。
|
||||||
- [x] special family + 完整角色块本地闸门通过:73 个 Astro 文件零诊断,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;十三页签桌面与 390px 移动端无运行时异常或文档级横向溢出,十六套真实 Chrome 专题回归全部通过。
|
- [x] special family + 完整角色块本地闸门通过:73 个 Astro 文件零诊断,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;十三页签桌面与 390px 移动端无运行时异常或文档级横向溢出,十六套真实 Chrome 专题回归全部通过。
|
||||||
- [x] DeepSeek special family + 完整角色块以源提交 `1d0a331`、不可变镜像 `20260729T134520Z-1d0a331` 发布;OCI digest `sha256:0c922769…3bfd5`,复用 NAS `12010→8080`、NPM host 31 / cert 41、门户 `LLM ATLAS / projects / 180`,HTTPS 与十六套生产 Chrome 回归全通过;保留 `20260729T124516Z-ea49ed0` 回滚。
|
- [x] DeepSeek special family + 完整角色块以源提交 `1d0a331`、不可变镜像 `20260729T134520Z-1d0a331` 发布;OCI digest `sha256:0c922769…3bfd5`,复用 NAS `12010→8080`、NPM host 31 / cert 41、门户 `LLM ATLAS / projects / 180`,HTTPS 与十六套生产 Chrome 回归全通过;保留 `20260729T124516Z-ea49ed0` 回滚。
|
||||||
|
- [x] DeepSeek Round 04 固定官方 `DeepSeek-V2-Lite-Chat` revision `85864749…f64c7`,核验 4 个 safetensors 分片与索引中的 `31,412,968,448` tensor bytes;完整 BF16 权重采用 GPU 25 层 + CPU 2 层、norm 与 lm_head 的显式 offload,设备字节账与官方 index 精确闭合。
|
||||||
|
- [x] 以 English / 中文 / code / math 各 4 个完整公开来源,执行 system 0/1 × EOS/BOS/x/句点 8 条件、共 128 个 greedy 输出;十条成对边分别报告 exact、共同前缀、编辑距离与 token similarity,不把单个特殊词元干预写成能力、遵循或路由解释。
|
||||||
|
- [x] 128 个输出中 31 个自然命中 EOS、97 个在 `max_new_tokens=128` 处截断;所有聚合永久同时显示 completion 状态,math 上的 2/4 system-EOS exact 仅是固定样本描述,不升级为领域能力结论。
|
||||||
|
- [x] 独立每域 1-source 复跑的 32 / 32 prompt hash、生成 token IDs、文本与 EOS 状态 exact;formal / repro 完整 JSON SHA-256 为 `54496955…b0e6196e` / `b74c3160…025b0c72`,模型配置、代码、tokenizer 与四个权重分片 hash 一并固化。
|
||||||
|
- [x] 第十八个 DeepSeek 交互实验将逐来源双输出、十边分歧图、29 单元设备切分与证据阶梯放到同一页;compact 构建器同时校验固定 revision、官方工件 hash 与 32 / 32 长复现合同。
|
||||||
|
- [x] DeepSeek Chat 行为层本地闸门通过:75 个 Astro 文件零诊断,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;十四页签桌面与 390px 移动端无运行时异常或文档级横向溢出,十六套真实 Chrome 专题回归全部通过。
|
||||||
|
|
||||||
## 正在进行
|
## 正在进行
|
||||||
|
|
||||||
- [ ] K3 三轮下一闸门:获得真实 token hidden states、expert load 与 cache traces,解释或修订 `A_log [128]` 工件冲突,再做 Figure 3/4/5 数值重绘和独立小模型复现。
|
- [ ] K3 三轮下一闸门:获得真实 token hidden states、expert load 与 cache traces,解释或修订 `A_log [128]` 工件冲突,再做 Figure 3/4/5 数值重绘和独立小模型复现。
|
||||||
- [ ] DeepSeek 三轮下一闸门:加入 V2-Lite-Chat 生成/行为对照;扩到完整 27 层并固定 batch content 审计,再推进 SM90 FlashMLA、FP8 / pipeline traces 与 R1-like RL 小模型复现。
|
- [ ] DeepSeek 四轮下一闸门:把 31 / 128 completion 提升为长度分层、可执行 evaluator 与更长生成预算;扩到完整 27 层并固定 batch content 审计,再推进 SM90 FlashMLA、FP8 / pipeline traces 与 R1-like RL 小模型复现。
|
||||||
- [ ] 表示、位置与残差二轮:真实 hidden-state / norm traces、长上下文位置外推复现与 mHC / AttnRes 深层稳定性消融。
|
- [ ] 表示、位置与残差二轮:真实 hidden-state / norm traces、长上下文位置外推复现与 mHC / AttnRes 深层稳定性消融。
|
||||||
- [ ] 评测安全二轮:真实 cross-harness / pass@k 复跑、Judge 元评测、动态污染与过拒案例。
|
- [ ] 评测安全二轮:真实 cross-harness / pass@k 复跑、Judge 元评测、动态污染与过拒案例。
|
||||||
- [ ] 推理服务二轮:真实 GPU kernel / workload traces、功耗与成本、跨 vLLM / SGLang / TensorRT-LLM 复现。
|
- [ ] 推理服务二轮:真实 GPU kernel / workload traces、功耗与成本、跨 vLLM / SGLang / TensorRT-LLM 复现。
|
||||||
@@ -393,6 +399,9 @@
|
|||||||
| 2026-07-29 | 角色块按 head × delimiter 分解 | `User:`/`Assistant:` 都是两个普通 IDs;head、delimiter、interaction 与四条 direct edges 分账,不把单位置作用冒充完整角色语义 |
|
| 2026-07-29 | 角色块按 head × delimiter 分解 | `User:`/`Assistant:` 都是两个普通 IDs;head、delimiter、interaction 与四条 direct edges 分账,不把单位置作用冒充完整角色语义 |
|
||||||
| 2026-07-29 | target-content 与 full-input 永久分因果身份 | 精确对齐目标只观察编辑后的下游传播;完整输入包含被编辑 token 自身,因此方向变化不是主结果矛盾 |
|
| 2026-07-29 | target-content 与 full-input 永久分因果身份 | 精确对齐目标只观察编辑后的下游传播;完整输入包含被编辑 token 自身,因此方向变化不是主结果矛盾 |
|
||||||
| 2026-07-29 | DeepSeek 十三联工件实验以 `20260729T134520Z-1d0a331` 发布 | OCI digest `sha256:0c922769…3bfd5`;复用 NAS 12010→8080、NPM 31 / cert 41、门户 order 180;十六套生产 Chrome 回归通过,保留上一不可变镜像回滚 |
|
| 2026-07-29 | DeepSeek 十三联工件实验以 `20260729T134520Z-1d0a331` 发布 | OCI digest `sha256:0c922769…3bfd5`;复用 NAS 12010→8080、NPM 31 / cert 41、门户 order 180;十六套生产 Chrome 回归通过,保留上一不可变镜像回滚 |
|
||||||
|
| 2026-07-29 | Chat checkpoint、生成行为与能力评测分三层 | 官方 SFT Chat 权重可以支撑真实生成;成对输出分歧只证明干预传播,未配任务 evaluator、样本置信区间与足够 completion 前不写成能力结论 |
|
||||||
|
| 2026-07-29 | 生成实验把 completion 设为首要审计字段 | 31 / 128 自然 EOS 与 97 / 128 长度截断分开报告;截断输出不被当作完整答案参与越界判断 |
|
||||||
|
| 2026-07-29 | CPU offload 是执行拓扑而非模型属性 | 31.4 GB BF16 工件不因本机 25-GPU-layer / 2-CPU-layer 切分而改变模型身份;设备映射、参数字节与峰值显存都进入复现合同 |
|
||||||
| 2026-07-29 | K3 二轮按 32 张对象账与完整报告顺序重建 | total/active、2.5×、KDA state、深度来源、专家路由、视觉目标、轨迹、缓存与评测协议不再压成一页组件摘要 |
|
| 2026-07-29 | K3 二轮按 32 张对象账与完整报告顺序重建 | total/active、2.5×、KDA state、深度来源、专家路由、视觉目标、轨迹、缓存与评测协议不再压成一页组件摘要 |
|
||||||
| 2026-07-29 | K3 原生视觉事实回到 §2.4 / §3.3 核验 | 删除“先冻结语言模型再解冻”旧表述;明确 MoonViT-V2 从头训练,视觉/文本从开始共同 NTP |
|
| 2026-07-29 | K3 原生视觉事实回到 §2.4 / §3.3 核验 | 删除“先冻结语言模型再解冻”旧表述;明确 MoonViT-V2 从头训练,视觉/文本从开始共同 NTP |
|
||||||
| 2026-07-29 | K3 Figure 1–16 / Table 1–5 全部建立课程视觉契约 | 每张图同时写支持范围与不可外推项;作者报告、论文、推导与 toy model 使用 R/P/D/T 标签 |
|
| 2026-07-29 | K3 Figure 1–16 / Table 1–5 全部建立课程视觉契约 | 每张图同时写支持范围与不可外推项;作者报告、论文、推导与 toy model 使用 R/P/D/T 标签 |
|
||||||
|
|||||||
@@ -19,7 +19,7 @@
|
|||||||
|
|
||||||
当前里程碑包含 17 专题学习地图、486 篇关键论文索引、Kimi K3 完整导读,
|
当前里程碑包含 17 专题学习地图、486 篇关键论文索引、Kimi K3 完整导读,
|
||||||
语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 技术谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、工具使用与长程 Agent、原生多模态、训练系统、推理服务、数值优化,以及评测与安全深度专题,
|
语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 技术谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、工具使用与长程 Agent、原生多模态、训练系统、推理服务、数值优化,以及评测与安全深度专题,
|
||||||
以及 84 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、
|
以及 85 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、
|
||||||
8 个交互实验和 100 个一手/官方节点,完整覆盖架构、预训练、后训练、系统、评测、案例与附录。
|
8 个交互实验和 100 个一手/官方节点,完整覆盖架构、预训练、后训练、系统、评测、案例与附录。
|
||||||
第三轮已完成开放工件与首个真实 kernel 里程碑:固定官方模型与 FlashKDA revisions,审计 96 个 checkpoint shards、
|
第三轮已完成开放工件与首个真实 kernel 里程碑:固定官方模型与 FlashKDA revisions,审计 96 个 checkpoint shards、
|
||||||
497,220 个 tensor entries、真实 KDA / MLA / MoE / MoonViT shapes 与小范围参数统计,并用 4 个新视图
|
497,220 个 tensor entries、真实 KDA / MLA / MoE / MoonViT shapes 与小范围参数统计,并用 4 个新视图
|
||||||
@@ -27,8 +27,8 @@
|
|||||||
`sm_120a` wheel,在 RTX 5090 上完成 6/6 官方参考 exact-match 和 K3 fixed / varlen 形状计时。详见
|
`sm_120a` wheel,在 RTX 5090 上完成 6/6 官方参考 exact-match 和 K3 fixed / varlen 形状计时。详见
|
||||||
[K3_ARTIFACT_AUDIT.md](./research/K3_ARTIFACT_AUDIT.md) 与
|
[K3_ARTIFACT_AUDIT.md](./research/K3_ARTIFACT_AUDIT.md) 与
|
||||||
[checkpoint_probe.py](./experiments/k3/checkpoint_probe.py)、[FlashKDA probe](./experiments/k3/flashkda/)。
|
[checkpoint_probe.py](./experiments/k3/checkpoint_probe.py)、[FlashKDA probe](./experiments/k3/flashkda/)。
|
||||||
DeepSeek 三轮专题以 24 张问题账、10 次技术转向、
|
DeepSeek 四轮专题以 24 张问题账、10 次技术转向、
|
||||||
17 个交互实验和 60 个一手/官方节点,串起 Dense、MoE、MLA、V3 协同、R1 与 V4;
|
18 个交互实验和 60 个一手/官方节点,串起 Dense、MoE、MLA、V3 协同、R1 与 V4;
|
||||||
并固定官方 V2-Lite revision,在 RTX 5090 上连续执行 7/27 层,记录 3,240 次真实专家选择、
|
并固定官方 V2-Lite revision,在 RTX 5090 上连续执行 7/27 层,记录 3,240 次真实专家选择、
|
||||||
MLA/HF eager cache shapes 与 `31/31` exact 独立复跑;进一步用真实 layer-1 权重执行官方 V3
|
MLA/HF eager cache shapes 与 `31/31` exact 独立复跑;进一步用真实 layer-1 权重执行官方 V3
|
||||||
naive/absorb 路径,实际写入 576 元素 latent cache,并以 FP32 将两种结合顺序的最大误差压到
|
naive/absorb 路径,实际写入 576 元素 latent cache,并以 FP32 将两种结合顺序的最大误差压到
|
||||||
@@ -58,8 +58,13 @@ EOS,换行为 23 / 24,但 base checkpoint、非法反事实序列和无行
|
|||||||
则在 34,488 / 34,488 个目标 ordered top-6 上 exact,direct TV / JSD / ΔCV 全为零,
|
则在 34,488 / 34,488 个目标 ordered top-6 上 exact,direct TV / JSD / ΔCV 全为零,
|
||||||
形成严格 causal suffix 负对照。进一步穷尽 tokenizer 的 BOS/EOS special inventory,
|
形成严格 causal suffix 负对照。进一步穷尽 tokenizer 的 BOS/EOS special inventory,
|
||||||
并以 `User/Assistant × :/x` 完成两-token 角色块因子分解;两组实验各新增 2,044,224
|
并以 `User/Assistant × :/x` 完成两-token 角色块因子分解;两组实验各新增 2,044,224
|
||||||
次真实路由、完整 JSON 独立复跑 byte-exact。当前累计 11,289,744 次公开语料路由。FlashMLA 的
|
次真实路由、完整 JSON 独立复跑 byte-exact。当前累计 11,289,744 次公开语料路由。
|
||||||
SM90/SM100 官方支持矩阵与本机 SM120 边界单独记账。详见
|
Round 04 又固定官方 `DeepSeek-V2-Lite-Chat` revision,加载 31.4 GB 完整 BF16 权重,
|
||||||
|
以 GPU 25 层 + CPU 2 层和输出头的透明 offload 执行 16 个公开来源、8 种 system ×
|
||||||
|
边界条件,共 128 个 greedy 输出;四域中 31 个自然命中 EOS、97 个在 128 新 token
|
||||||
|
上限处截断,因此只把十条成对边写成“生成分歧”,不把它们冒充能力评测。每域独立抽取
|
||||||
|
1 个来源的 32 个输出再次运行,prompt hash、生成 token IDs、文本与 EOS 状态均为
|
||||||
|
`32 / 32` exact。FlashMLA 的 SM90/SM100 官方支持矩阵与本机 SM120 边界单独记账。详见
|
||||||
[DEEPSEEK_V2_LITE_TRACE.md](./research/DEEPSEEK_V2_LITE_TRACE.md) 与
|
[DEEPSEEK_V2_LITE_TRACE.md](./research/DEEPSEEK_V2_LITE_TRACE.md) 与
|
||||||
[DEEPSEEK_MLA_ABSORB_AUDIT.md](./research/DEEPSEEK_MLA_ABSORB_AUDIT.md)、
|
[DEEPSEEK_MLA_ABSORB_AUDIT.md](./research/DEEPSEEK_MLA_ABSORB_AUDIT.md)、
|
||||||
[DEEPSEEK_ROUTING_CORPUS_AUDIT.md](./research/DEEPSEEK_ROUTING_CORPUS_AUDIT.md)、
|
[DEEPSEEK_ROUTING_CORPUS_AUDIT.md](./research/DEEPSEEK_ROUTING_CORPUS_AUDIT.md)、
|
||||||
@@ -70,7 +75,8 @@ SM90/SM100 官方支持矩阵与本机 SM120 边界单独记账。详见
|
|||||||
[DEEPSEEK_ROUTING_HISTORY_BOUNDARY_TOKEN_AUDIT.md](./research/DEEPSEEK_ROUTING_HISTORY_BOUNDARY_TOKEN_AUDIT.md) 与
|
[DEEPSEEK_ROUTING_HISTORY_BOUNDARY_TOKEN_AUDIT.md](./research/DEEPSEEK_ROUTING_HISTORY_BOUNDARY_TOKEN_AUDIT.md) 与
|
||||||
[DEEPSEEK_ROUTING_ROLE_MARKER_HEAD_AUDIT.md](./research/DEEPSEEK_ROUTING_ROLE_MARKER_HEAD_AUDIT.md)、
|
[DEEPSEEK_ROUTING_ROLE_MARKER_HEAD_AUDIT.md](./research/DEEPSEEK_ROUTING_ROLE_MARKER_HEAD_AUDIT.md)、
|
||||||
[DEEPSEEK_ROUTING_SPECIAL_TOKEN_FAMILY_AUDIT.md](./research/DEEPSEEK_ROUTING_SPECIAL_TOKEN_FAMILY_AUDIT.md) 与
|
[DEEPSEEK_ROUTING_SPECIAL_TOKEN_FAMILY_AUDIT.md](./research/DEEPSEEK_ROUTING_SPECIAL_TOKEN_FAMILY_AUDIT.md) 与
|
||||||
[DEEPSEEK_ROUTING_ROLE_MARKER_BLOCK_AUDIT.md](./research/DEEPSEEK_ROUTING_ROLE_MARKER_BLOCK_AUDIT.md)。
|
[DEEPSEEK_ROUTING_ROLE_MARKER_BLOCK_AUDIT.md](./research/DEEPSEEK_ROUTING_ROLE_MARKER_BLOCK_AUDIT.md),以及
|
||||||
|
[DEEPSEEK_V2_LITE_CHAT_BEHAVIOR_AUDIT.md](./research/DEEPSEEK_V2_LITE_CHAT_BEHAVIOR_AUDIT.md)。
|
||||||
其余专题按进度账本持续扩建。
|
其余专题按进度账本持续扩建。
|
||||||
|
|
||||||
## 本地开发
|
## 本地开发
|
||||||
|
|||||||
+1
-1
@@ -156,7 +156,7 @@ pass^k、校准、动态基准、代码 Verifier、LLM Judge、Arena、Agent 最
|
|||||||
1. **Kimi K3 解剖**:二轮已完成 32 张问题账、Figure 1–16 / Table 1–5 审计、
|
1. **Kimi K3 解剖**:二轮已完成 32 张问题账、Figure 1–16 / Table 1–5 审计、
|
||||||
8 个报告实验与 100 节点阅读链;三轮首个里程碑进一步固定官方 revisions,审计 96 个 checkpoint shards、
|
8 个报告实验与 100 节点阅读链;三轮首个里程碑进一步固定官方 revisions,审计 96 个 checkpoint shards、
|
||||||
497,220 个 tensor entries、真实 KDA / MLA / MoE / MoonViT shapes,并用 4 个工件视图公开复现边界。
|
497,220 个 tensor entries、真实 KDA / MLA / MoE / MoonViT shapes,并用 4 个工件视图公开复现边界。
|
||||||
2. **DeepSeek 技术谱系**:DeepSeek LLM → DeepSeekMoE → V2/MLA → V3/FP8/MTP/DualPipe → Math/GRPO → R1 → V3.2/DSA → V4 长上下文。
|
2. **DeepSeek 技术谱系**:DeepSeek LLM → DeepSeekMoE → V2/MLA → V3/FP8/MTP/DualPipe → Math/GRPO → R1 → V3.2/DSA → V4 长上下文;四轮已经从 24 张问题账、60 个一手/官方节点和 4 个公式实验,推进到 13 个 Base 工件实验与 1 个完整 Chat 行为实验。最新生成层固定 31.4 GB 官方 V2-Lite-Chat BF16 权重,以 16 个来源 × 8 条件得到 128 个输出,并把自然 EOS、长度截断、成对分歧和 32 / 32 独立复现分账。
|
||||||
3. **“一个 Token 的旅行”**:从文本分词,经注意力、MoE、GPU 集群、后训练,再到线上推理与工具调用。
|
3. **“一个 Token 的旅行”**:从文本分词,经注意力、MoE、GPU 集群、后训练,再到线上推理与工具调用。
|
||||||
|
|
||||||
## 完成标准
|
## 完成标准
|
||||||
|
|||||||
@@ -531,3 +531,58 @@ in most cells, without a uniform system modulation. See
|
|||||||
`research/DEEPSEEK_ROUTING_ROLE_MARKER_BLOCK_AUDIT.md` for the exact factor
|
`research/DEEPSEEK_ROUTING_ROLE_MARKER_BLOCK_AUDIT.md` for the exact factor
|
||||||
coding, intervals, direct dependencies, full-input split, route alignment,
|
coding, intervals, direct dependencies, full-input split, route alignment,
|
||||||
cross-batch audit, sources, and non-claims.
|
cross-batch audit, sources, and non-claims.
|
||||||
|
|
||||||
|
## Full Chat-checkpoint generation behavior
|
||||||
|
|
||||||
|
`v2_lite_chat_special_token_behavior_probe.py` carries the special-token grid
|
||||||
|
into the full official SFT Chat checkpoint:
|
||||||
|
|
||||||
|
```text
|
||||||
|
deepseek-ai/DeepSeek-V2-Lite-Chat
|
||||||
|
@85864749cd611b4353ce1decdb286193298f64c7
|
||||||
|
|
||||||
|
system off/on × EOS / BOS / x / period
|
||||||
|
```
|
||||||
|
|
||||||
|
It reuses the pinned Base-routing source IDs and ranks but generates from the
|
||||||
|
full source text. All eight variants of one source run in one left-padded
|
||||||
|
batch. Decoding is deterministic greedy; BOS / x / period remain invalid-chat
|
||||||
|
single-ID counterfactuals.
|
||||||
|
|
||||||
|
The official BF16 checkpoint contains `31,412,968,448` tensor bytes. Because
|
||||||
|
the local RTX 5090 has less than the official 40GB single-GPU boundary, the
|
||||||
|
formal run uses Accelerate `device_map=auto`: embeddings and layers 0–24 are
|
||||||
|
placed on CUDA, while layers 25–26, final norm, and LM head are CPU-offloaded.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
PYTHONPATH=/path/to/transformers-4.41.2-deps \
|
||||||
|
python -B experiments/deepseek/v2_lite_chat_special_token_behavior_probe.py \
|
||||||
|
--artifact-dir /path/to/deepseek-v2-lite-chat \
|
||||||
|
--reference-routing-json \
|
||||||
|
src/data/deepseek-v2-lite-routing-special-token-family-control.json \
|
||||||
|
--human-eval /path/to/HumanEval.jsonl.gz \
|
||||||
|
--gsm8k /path/to/gsm8k/test.jsonl \
|
||||||
|
--tnews /path/to/tnews/test.json \
|
||||||
|
--tnews-archive /path/to/tnews_public.zip \
|
||||||
|
--wikitext /path/to/wikitext-validation.parquet \
|
||||||
|
--output src/data/deepseek-v2-lite-chat-behavior.json \
|
||||||
|
--per-domain 4 \
|
||||||
|
--max-new-tokens 128 \
|
||||||
|
--gpu-memory 29GiB \
|
||||||
|
--cpu-memory 80GiB
|
||||||
|
```
|
||||||
|
|
||||||
|
The formal artifact contains 16 sources and 128 outputs. Only 31 outputs hit
|
||||||
|
EOS; 97 reach the 128-token cap, so task-score fields are diagnostic and are
|
||||||
|
not used for capability ranking. A 4-source / 32-cell long-sequence rerun is
|
||||||
|
exact for prompt hashes, generated token IDs, decoded text, and EOS state.
|
||||||
|
|
||||||
|
```text
|
||||||
|
formal 54496955d0dd20a3116e6b758a55f56e2444c43dd99bfef94a2cb7b140e6196e
|
||||||
|
repro b74c31606c0fc70a8e52cc6c3d6135c4e916d698b4d5dae51863124d025b0c72
|
||||||
|
```
|
||||||
|
|
||||||
|
See `research/DEEPSEEK_V2_LITE_CHAT_BEHAVIOR_PROTOCOL.md` and
|
||||||
|
`research/DEEPSEEK_V2_LITE_CHAT_BEHAVIOR_AUDIT.md` for the preregistered
|
||||||
|
contract, output-divergence table, offload device map, truncation boundary,
|
||||||
|
rerun coverage, execution fixes, primary sources, and forbidden conclusions.
|
||||||
|
|||||||
File diff suppressed because it is too large
Load Diff
@@ -13,6 +13,7 @@
|
|||||||
"build:data:deepseek-role": "node scripts/build-deepseek-role-marker-compact.mjs",
|
"build:data:deepseek-role": "node scripts/build-deepseek-role-marker-compact.mjs",
|
||||||
"build:data:deepseek-special": "node scripts/build-deepseek-special-token-family-compact.mjs",
|
"build:data:deepseek-special": "node scripts/build-deepseek-special-token-family-compact.mjs",
|
||||||
"build:data:deepseek-role-block": "node scripts/build-deepseek-role-marker-block-compact.mjs",
|
"build:data:deepseek-role-block": "node scripts/build-deepseek-role-marker-block-compact.mjs",
|
||||||
|
"build:data:deepseek-chat-behavior": "node scripts/build-deepseek-chat-behavior-compact.mjs",
|
||||||
"check:site": "node scripts/check-site.mjs",
|
"check:site": "node scripts/check-site.mjs",
|
||||||
"check:moe-browser": "node scripts/check-moe-browser.mjs",
|
"check:moe-browser": "node scripts/check-moe-browser.mjs",
|
||||||
"check:reasoning-browser": "node scripts/check-reasoning-browser.mjs",
|
"check:reasoning-browser": "node scripts/check-reasoning-browser.mjs",
|
||||||
|
|||||||
@@ -0,0 +1,434 @@
|
|||||||
|
# DeepSeek-V2-Lite-Chat 最终生成行为审计
|
||||||
|
|
||||||
|
> 状态:Round 04 正式结果
|
||||||
|
>
|
||||||
|
> 正式原始 JSON:`src/data/deepseek-v2-lite-chat-behavior.json`
|
||||||
|
>
|
||||||
|
> 长序列复跑 JSON:`src/data/deepseek-v2-lite-chat-behavior-repro-1pd.json`
|
||||||
|
>
|
||||||
|
> 执行脚本:`experiments/deepseek/v2_lite_chat_special_token_behavior_probe.py`
|
||||||
|
>
|
||||||
|
> 预注册协议:`research/DEEPSEEK_V2_LITE_CHAT_BEHAVIOR_PROTOCOL.md`
|
||||||
|
|
||||||
|
## 0. 一句话结论
|
||||||
|
|
||||||
|
固定官方 SFT Chat checkpoint、greedy decoding 与同 source 八格 batch 后,一处历史边界
|
||||||
|
ID 或一条 system message 经常足以让最终 token 轨迹分叉;但本轮 128 个输出中只有 31 个
|
||||||
|
在 128-token 上限内遇到 EOS,因此它识别的是**小样本确定性输出敏感性**,不是任务能力、
|
||||||
|
采样分布稳定性,也不是“某种 token 更好”。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. 为什么上一轮还不能回答这个问题
|
||||||
|
|
||||||
|
上一组 special-token family 实验加载的是 `DeepSeek-V2-Lite` Base checkpoint,只观察
|
||||||
|
前六个 MoE 层的 hidden states 与 top-6 expert routes。它能回答:
|
||||||
|
|
||||||
|
```text
|
||||||
|
同一个 pre-target input ID 被替换以后,
|
||||||
|
后续目标 token 的专家路径怎样改变?
|
||||||
|
```
|
||||||
|
|
||||||
|
但它不能回答:
|
||||||
|
|
||||||
|
```text
|
||||||
|
完整模型最后生成了什么?
|
||||||
|
两格输出是否逐 token 相同?
|
||||||
|
任务结果有没有改变?
|
||||||
|
```
|
||||||
|
|
||||||
|
本轮因此换到官方
|
||||||
|
[`deepseek-ai/DeepSeek-V2-Lite-Chat`](https://huggingface.co/deepseek-ai/DeepSeek-V2-Lite-Chat)
|
||||||
|
SFT checkpoint,并真正调用 `generate()`。DeepSeek 官方仓库把 Lite-Chat 明确列作
|
||||||
|
16B total / 2.4B active / 32K context 的 **SFT** 模型;它与较大的 V2-Chat (RL) 不是
|
||||||
|
同一个 checkpoint 身份。
|
||||||
|
来源:
|
||||||
|
[DeepSeek-V2 官方仓库](https://github.com/deepseek-ai/DeepSeek-V2)、
|
||||||
|
[DeepSeek-V2 报告](https://arxiv.org/abs/2405.04434)。
|
||||||
|
|
||||||
|
这也意味着:
|
||||||
|
|
||||||
|
- Base routing 与 Chat generation 是两种证据层;
|
||||||
|
- 不能把二者的 effect size 直接相减;
|
||||||
|
- 不能把 Chat 输出分叉倒推成某个 expert 的因果中介;
|
||||||
|
- 不能把 SFT checkpoint 写成 R1、RL 或当前线上 DeepSeek。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. 固定了什么
|
||||||
|
|
||||||
|
### 2.1 模型工件
|
||||||
|
|
||||||
|
| 对象 | 固定值 |
|
||||||
|
|---|---:|
|
||||||
|
| 模型 | `deepseek-ai/DeepSeek-V2-Lite-Chat` |
|
||||||
|
| revision | `85864749cd611b4353ce1decdb286193298f64c7` |
|
||||||
|
| checkpoint 身份 | SFT Chat |
|
||||||
|
| dtype | BF16 |
|
||||||
|
| safetensors shards | 4 |
|
||||||
|
| index tensor bytes | `31,412,968,448` |
|
||||||
|
| shard file bytes(含 headers) | `31,413,626,576` |
|
||||||
|
| 文件数 | 12 |
|
||||||
|
| revision metadata | 12 / 12 同一 revision |
|
||||||
|
|
||||||
|
运行时参数 bytes 按 dtype 汇总也是 `31,412,968,448`,全部为 `torch.bfloat16`,与
|
||||||
|
index 的 `metadata.total_size` 精确相等。
|
||||||
|
|
||||||
|
正式原始结果:
|
||||||
|
|
||||||
|
```text
|
||||||
|
542,559 bytes
|
||||||
|
SHA-256 54496955d0dd20a3116e6b758a55f56e2444c43dd99bfef94a2cb7b140e6196e
|
||||||
|
```
|
||||||
|
长序列复跑:
|
||||||
|
|
||||||
|
```text
|
||||||
|
153,819 bytes
|
||||||
|
SHA-256 b74c31606c0fc70a8e52cc6c3d6135c4e916d698b4d5dae51863124d025b0c72
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2.2 tokenizer
|
||||||
|
|
||||||
|
Chat 与 Base 的 `tokenizer.json`、`tokenizer_config.json` 逐字节相同:
|
||||||
|
|
||||||
|
```text
|
||||||
|
tokenizer.json
|
||||||
|
41f3bf64213da8c012d8bd0871a58a1fdf70463e8f08f110ddbb1082f529f669
|
||||||
|
|
||||||
|
tokenizer_config.json
|
||||||
|
31181eaf79394ea26728d95ecb54fe7c8413e6f56085dbabc8b0818134380ec8
|
||||||
|
```
|
||||||
|
|
||||||
|
固定边界 ID:
|
||||||
|
|
||||||
|
| level | ID | 类别 | 官方合法历史边界 |
|
||||||
|
|---|---:|---|---|
|
||||||
|
| EOS | `100001` | special;同时是 PAD alias | 是 |
|
||||||
|
| BOS | `100000` | special | 否 |
|
||||||
|
| `x` | `87` | ordinary content | 否 |
|
||||||
|
| `.` | `13` | ordinary punctuation | 否 |
|
||||||
|
|
||||||
|
EOS 与 BOS 穷尽这个固定 tokenizer 的 special-token inventory。BOS / `x` / 句点三格
|
||||||
|
都不是官方合法 Chat 序列,它们只是在 official rendering 后替换一个 pre-target
|
||||||
|
input ID 的反事实控制。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. 八格设计
|
||||||
|
|
||||||
|
```text
|
||||||
|
system off/on × EOS / BOS / x / period
|
||||||
|
```
|
||||||
|
|
||||||
|
同一 system 水平内:
|
||||||
|
|
||||||
|
- 四格 token 数相同;
|
||||||
|
- target content、generation prompt 与 target position 相同;
|
||||||
|
- EOS 格相对官方序列改 0 个 ID;
|
||||||
|
- BOS / `x` / period 各恰好改 1 个 ID。
|
||||||
|
|
||||||
|
system on 比 system off 多 16 个 token。八格同 batch 时使用:
|
||||||
|
|
||||||
|
```text
|
||||||
|
PAD = EOS
|
||||||
|
left padding
|
||||||
|
attention_mask = 0 on padding
|
||||||
|
generation prefix ends at the same batch column
|
||||||
|
```
|
||||||
|
|
||||||
|
这次协议修订发生在成功加载模型之前。最初“八格全部同长度”的检查忽略了 system 文本本身
|
||||||
|
会增加长度;失败后把合同改成“同一 system 水平内等长 + batch 左填充”。这不是结果后改
|
||||||
|
统计口径,而是修复一个无法构造 tensor 的前置序列错误。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. source 合同
|
||||||
|
|
||||||
|
正式运行:
|
||||||
|
|
||||||
|
```text
|
||||||
|
4 domains × 4 sources/domain = 16 sources
|
||||||
|
16 sources × 8 conditions = 128 outputs
|
||||||
|
```
|
||||||
|
|
||||||
|
四域:
|
||||||
|
|
||||||
|
- English encyclopedia / WikiText-2 raw validation;
|
||||||
|
- Chinese news / CLUE TNEWS public test;
|
||||||
|
- Python code / HumanEval;
|
||||||
|
- Grade-school math / GSM8K test。
|
||||||
|
|
||||||
|
source ID、selection rank 与域内顺序直接复用 Base special-token 正式实验。每条 source
|
||||||
|
的文本 SHA-256 重新核验;但本轮使用**完整 source 文本**,而路由实验使用固定目标前
|
||||||
|
23 tokens。这是为了让数学题和代码 prompt 不在半句话上生成。
|
||||||
|
|
||||||
|
因此:
|
||||||
|
|
||||||
|
```text
|
||||||
|
同 source 身份 = 是
|
||||||
|
同 checkpoint = 否(Base → SFT Chat)
|
||||||
|
同输入文本长度 = 否(23-token prefix → full source)
|
||||||
|
同证据对象 = 否(routes → generated outputs)
|
||||||
|
```
|
||||||
|
|
||||||
|
每域 4 条是预定 source 顺序的前四条,不是 benchmark 的代表性抽样,也没有能力估计所需
|
||||||
|
的样本量。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. 解码合同
|
||||||
|
|
||||||
|
正式运行固定:
|
||||||
|
|
||||||
|
```text
|
||||||
|
decode = greedy
|
||||||
|
do_sample = false
|
||||||
|
temperature = unset
|
||||||
|
top_p = unset
|
||||||
|
max_new_tokens = 128
|
||||||
|
use_cache = true
|
||||||
|
```
|
||||||
|
|
||||||
|
官方 `generation_config.json` 另行记录:
|
||||||
|
|
||||||
|
```text
|
||||||
|
do_sample = true
|
||||||
|
temperature = 0.3
|
||||||
|
top_p = 0.95
|
||||||
|
```
|
||||||
|
|
||||||
|
它没有用于本轮。选择 greedy 的目的,是先建立逐 token 复跑闸门,不是宣称 greedy 比
|
||||||
|
官方 sampling 更真实。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. 完整 BF16 权重怎样装进本机
|
||||||
|
|
||||||
|
官方 Lite model card 写明 BF16 inference 需要一张 40GB GPU。本机 RTX 5090 可见
|
||||||
|
32,607 MiB,因此不能声称“单卡完整 BF16”。实际运行使用:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Accelerate device_map = auto
|
||||||
|
max_memory[CUDA:0] = 29 GiB
|
||||||
|
max_memory[CPU] = 80 GiB
|
||||||
|
attention = official eager path
|
||||||
|
```
|
||||||
|
|
||||||
|
Hugging Face 的
|
||||||
|
[Big Model Inference 文档](https://huggingface.co/docs/accelerate/main/concept_guides/big_model_inference)
|
||||||
|
说明,`hf_device_map` 是查看自动放置的正式入口;CPU-offloaded weights 会通过 hook 在
|
||||||
|
需要时进入执行设备。正式 device map:
|
||||||
|
|
||||||
|
```text
|
||||||
|
CUDA:0 embed_tokens + layers 0–24
|
||||||
|
CPU layers 25–26 + final norm + lm_head
|
||||||
|
```
|
||||||
|
|
||||||
|
| 执行账 | 观测 |
|
||||||
|
|---|---:|
|
||||||
|
| CUDA 参数 tensor bytes | `28,654,142,464` |
|
||||||
|
| offload hooks 暴露为 meta 的 tensor bytes | `2,758,825,984` |
|
||||||
|
| load peak CUDA allocation | `29,786,618,368` bytes |
|
||||||
|
| formal peak CUDA allocation | `31,578,740,224` bytes ≈ `29.41 GiB` |
|
||||||
|
| process max RSS | `18,665,132 KiB` ≈ `17.80 GiB` |
|
||||||
|
| model load | `8.58 s` |
|
||||||
|
| 16 个 source batch generation | `369.67 s` |
|
||||||
|
|
||||||
|
`29GiB max_memory` 约束参数放置,不是包含 KV cache、activation 和 CUDA runtime 的硬峰值;
|
||||||
|
因此实际 generation peak 高于 29GiB 不构成合同矛盾。
|
||||||
|
|
||||||
|
Accelerate-offloaded modules 在 forward 间可把 `parameter.device` 暴露为 `meta`;这里用
|
||||||
|
`hf_device_map` 判断 CPU placement,不能把 `meta` 字面解释为“参数不存在”。
|
||||||
|
|
||||||
|
这些延迟只描述本机 eager + CPU offload 审计,不代表 SGLang、vLLM、FlashMLA 或生产
|
||||||
|
serving 吞吐。DeepSeek 官方仓库也明确区分 Hugging Face 开源实现与优化 serving 路径。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. 正式结果:输出是否 exact
|
||||||
|
|
||||||
|
十条预注册比较边:
|
||||||
|
|
||||||
|
| edge | exact / 16 | mean common prefix | mean edit distance | mean token similarity |
|
||||||
|
|---|---:|---:|---:|---:|
|
||||||
|
| system · EOS | 2 / 16 | 32.875 | 64.625 | 48.26% |
|
||||||
|
| system · BOS | 1 / 16 | 22.125 | 74.000 | 40.95% |
|
||||||
|
| system · x | 0 / 16 | 4.875 | 95.188 | 24.15% |
|
||||||
|
| system · period | 0 / 16 | 4.500 | 94.813 | 24.82% |
|
||||||
|
| BOS − EOS · S0 | 2 / 16 | 24.563 | 75.125 | 39.89% |
|
||||||
|
| BOS − EOS · S1 | 1 / 16 | 17.500 | 77.813 | 37.49% |
|
||||||
|
| x − EOS · S0 | 0 / 16 | 12.563 | 79.188 | 37.18% |
|
||||||
|
| x − EOS · S1 | 0 / 16 | 9.813 | 85.500 | 29.82% |
|
||||||
|
| period − EOS · S0 | 1 / 16 | 22.250 | 80.375 | 35.88% |
|
||||||
|
| period − EOS · S1 | 0 / 16 | 3.688 | 85.750 | 30.30% |
|
||||||
|
|
||||||
|
`token similarity = 1 - Levenshtein distance / max(output lengths)`。
|
||||||
|
|
||||||
|
能写的最窄结论:
|
||||||
|
|
||||||
|
> 在这 16 条固定 source、固定 SFT Chat checkpoint 与 greedy contract 下,system
|
||||||
|
> message 和一个历史边界 ID 都可能改变最终生成 token 序列;完整 exact 是少数,而不是
|
||||||
|
> 默认状态。
|
||||||
|
|
||||||
|
不能写:
|
||||||
|
|
||||||
|
- system / EOS “普遍重要”;
|
||||||
|
- x / 句点“更差”;
|
||||||
|
- special token “更稳定”;
|
||||||
|
- edit distance 越大,能力变化越大;
|
||||||
|
- 输出分叉由前六层 route TV 解释。
|
||||||
|
|
||||||
|
四个 ID 既不是随机从 token 类别总体抽样,16 条 source 也不是总体样本,所以不存在可靠
|
||||||
|
的 “special vs ordinary 类别效应”。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. 分域结果最醒目的差异
|
||||||
|
|
||||||
|
### 8.1 official system edge(S0 EOS ↔ S1 EOS)
|
||||||
|
|
||||||
|
| 域 | exact / 4 | mean similarity |
|
||||||
|
|---|---:|---:|
|
||||||
|
| English encyclopedia | 0 / 4 | 51.37% |
|
||||||
|
| Chinese news | 0 / 4 | 28.57% |
|
||||||
|
| Python code | 0 / 4 | 29.10% |
|
||||||
|
| Grade-school math | 2 / 4 | 83.98% |
|
||||||
|
|
||||||
|
数学格在这个极小 cohort 上更常 exact,但不能据此说“数学更不受 system 影响”。它可能来自
|
||||||
|
source 内容、固定 greedy path、输出长度或 checkpoint 特定模式;每域只有 4 条。
|
||||||
|
|
||||||
|
### 8.2 x / period 下的 system edge
|
||||||
|
|
||||||
|
system × `x` 和 system × period 都是 0 / 16 exact,平均 similarity 分别为 24.15% /
|
||||||
|
24.82%。这说明两个固定 ordinary counterfactual 下,system on/off 的 greedy 轨迹在本
|
||||||
|
cohort 中更常早分叉;它仍不证明 ordinary token 类别导致更强 system effect。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. 截断率是主结果,不是脚注
|
||||||
|
|
||||||
|
按 condition 的 EOS 完成数:
|
||||||
|
|
||||||
|
| condition | hit EOS / 16 | max-128 truncated / 16 |
|
||||||
|
|---|---:|---:|
|
||||||
|
| S0 EOS | 2 | 14 |
|
||||||
|
| S1 EOS | 3 | 13 |
|
||||||
|
| S0 BOS | 3 | 13 |
|
||||||
|
| S1 BOS | 3 | 13 |
|
||||||
|
| S0 x | 2 | 14 |
|
||||||
|
| S1 x | 8 | 8 |
|
||||||
|
| S0 period | 2 | 14 |
|
||||||
|
| S1 period | 8 | 8 |
|
||||||
|
| **合计** | **31 / 128** | **97 / 128** |
|
||||||
|
|
||||||
|
因此本轮保存了 math final-number 与 Python AST parse 作为诊断字段,却不把它们做成能力
|
||||||
|
比较。比如某格的数学 final number 没出现,可能只是输出仍在推导;某格 AST parse 失败,
|
||||||
|
也可能只是 code fence 在第 128 token 被截断。
|
||||||
|
|
||||||
|
在公平能力表之前,下一协议至少需要:
|
||||||
|
|
||||||
|
1. 足够高的 completion contract,或明确定义 stop / answer extractor;
|
||||||
|
2. 代码实际执行与 sandbox;
|
||||||
|
3. 每个任务足够样本;
|
||||||
|
4. 对所有格使用相同完成预算与失败规则;
|
||||||
|
5. 报告 truncation、refusal、format failure;
|
||||||
|
6. 与 greedy 分开的多 seed sampling protocol。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. 复跑闸门
|
||||||
|
|
||||||
|
### 10.1 128-token 长序列复跑
|
||||||
|
|
||||||
|
复跑每域首条,共 4 source / 32 格。与正式 16-source 结果按
|
||||||
|
`source_id × condition` 对齐:
|
||||||
|
|
||||||
|
```text
|
||||||
|
prompt token hash exact 32 / 32
|
||||||
|
generated token IDs exact 32 / 32
|
||||||
|
decoded text exact 32 / 32
|
||||||
|
EOS state exact 32 / 32
|
||||||
|
```
|
||||||
|
|
||||||
|
其余 96 格没有做完整 128-token 独立复跑,网站和文档都不写成“128 / 128 rerun”。
|
||||||
|
|
||||||
|
### 10.2 smoke 与正式前缀
|
||||||
|
|
||||||
|
4 source / 32 格的 32-token smoke 做了独立复跑,32 / 32 generated token IDs 与文本
|
||||||
|
exact;正式 128-token 结果的前 32 tokens 也与 smoke 32 / 32 exact。
|
||||||
|
|
||||||
|
这支持“固定 greedy contract 可复跑”,不支持 sampling distribution 稳定。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11. 执行过程中暴露并记录的问题
|
||||||
|
|
||||||
|
### 11.1 system 长度
|
||||||
|
|
||||||
|
最初错误地检查“八格全等长”。system on 本来就多 16 tokens;改为同一 system 水平内
|
||||||
|
四个 boundary 格等长,并用显式左填充组成八格 batch。
|
||||||
|
|
||||||
|
### 11.2 Python 环境污染
|
||||||
|
|
||||||
|
早期命令把 `/usr/lib/python3/dist-packages` 混入 pyenv Python 3.10,触发 Pillow
|
||||||
|
`_imaging` 二进制不匹配。当前虚拟环境已经包含 pandas / pyarrow / Pillow,正式运行移除
|
||||||
|
系统路径。
|
||||||
|
|
||||||
|
### 11.3 `flash_attn` 动态依赖检查
|
||||||
|
|
||||||
|
Transformers 4.41 的 remote-code import scanner 把官方源码中受
|
||||||
|
`is_flash_attn_2_available()` 保护的 import 仍判成硬依赖。本机没有 flash-attn,正式运行
|
||||||
|
沿用前轮审计过的 loader,把官方 `configuration_deepseek.py` /
|
||||||
|
`modeling_deepseek.py` 作为本地 package 直接加载,并走官方 eager attention。
|
||||||
|
|
||||||
|
官方模型源码没有打补丁:
|
||||||
|
|
||||||
|
```text
|
||||||
|
modeling_deepseek.py
|
||||||
|
7d8e5221095286eea991137760893fd7ba52727c0b4ebf48ec09e8bc56b45b9c
|
||||||
|
```
|
||||||
|
|
||||||
|
### 11.4 official sampling config 的 warning
|
||||||
|
|
||||||
|
第一次 smoke 明确传 `do_sample=False`,但 model 自带 `.3 / .95` 导致 warning。正式脚本
|
||||||
|
同时把 `temperature` / `top_p` unset,消除“记录但未使用”的歧义;生成前缀复跑仍逐格
|
||||||
|
exact。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 12. 这轮最重要的非结论
|
||||||
|
|
||||||
|
1. **不是 benchmark。** 每域 4 条,且 97 / 128 截断。
|
||||||
|
2. **不是 token 类别总体效应。** 两个 special、两个人为选择的 ordinary ID。
|
||||||
|
3. **不是合法 Chat 对比。** BOS / `x` / period 是单 ID 反事实。
|
||||||
|
4. **不是 route → output 中介分析。** Base / Chat checkpoint 与输入长度都不同。
|
||||||
|
5. **不是 sampling robustness。** 只测 greedy。
|
||||||
|
6. **不是单卡 BF16。** 明确使用 CPU offload。
|
||||||
|
7. **不是生产性能。** eager + CPU offload 延迟只供执行审计。
|
||||||
|
8. **不是 RL checkpoint。** Lite-Chat 是 SFT。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 13. 下一闸门
|
||||||
|
|
||||||
|
优先级:
|
||||||
|
|
||||||
|
1. 把正式行为实验扩成 completion-aware protocol,先解决 97 / 128 截断;
|
||||||
|
2. 对数学做完整 answer extraction,对代码做 sandbox execution;
|
||||||
|
3. 独立设计多 seed sampling distribution 对照;
|
||||||
|
4. 在 Base 与 Chat 上用同一完整输入抽取 27 层 hidden / routes;
|
||||||
|
5. 预注册 route features 与 output divergence 的关联分析,并明确它仍不是中介因果;
|
||||||
|
6. 再推进支持硬件上的 FlashMLA、FP8 / pipeline traces 与 R1-like 小模型训练。
|
||||||
|
|
||||||
|
本轮真正完成的是从:
|
||||||
|
|
||||||
|
```text
|
||||||
|
“路由看起来变了”
|
||||||
|
```
|
||||||
|
|
||||||
|
前进到:
|
||||||
|
|
||||||
|
```text
|
||||||
|
“完整 SFT Chat 权重实际生成后,哪些固定格子逐 token 相同,
|
||||||
|
哪些分叉,以及证据为什么仍然不能越界。”
|
||||||
|
```
|
||||||
@@ -0,0 +1,187 @@
|
|||||||
|
# DeepSeek-V2-Lite-Chat 生成行为对照协议
|
||||||
|
|
||||||
|
> 状态:已执行;正式结果见 `DEEPSEEK_V2_LITE_CHAT_BEHAVIOR_AUDIT.md`
|
||||||
|
>
|
||||||
|
> 官方模型:`deepseek-ai/DeepSeek-V2-Lite-Chat`
|
||||||
|
>
|
||||||
|
> 固定 revision:`85864749cd611b4353ce1decdb286193298f64c7`
|
||||||
|
>
|
||||||
|
> checkpoint 身份:SFT Chat,而非 base / RL
|
||||||
|
>
|
||||||
|
> 权重:官方 BF16,四个 safetensors shards,约 31.4 GB
|
||||||
|
|
||||||
|
## 0. 为什么这一轮必须换证据层
|
||||||
|
|
||||||
|
前十组实验都观察 base checkpoint 的 hidden states 与 MoE routing。它们能回答:
|
||||||
|
|
||||||
|
```text
|
||||||
|
输入协议变化后,目标 token 的专家路径怎样改变?
|
||||||
|
```
|
||||||
|
|
||||||
|
但不能回答:
|
||||||
|
|
||||||
|
```text
|
||||||
|
模型最后生成了什么?答案是否保持一致?任务结果是否改变?
|
||||||
|
```
|
||||||
|
|
||||||
|
因此本轮不把 route TV 继续解释成能力,而是加载完整 Chat checkpoint,实际调用
|
||||||
|
`generate()`。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. 硬件边界先写在结果前面
|
||||||
|
|
||||||
|
官方 model card 明确写 BF16 inference 需要一张 40GB GPU。本机 RTX 5090 的可见
|
||||||
|
显存为 32,607 MiB,所以:
|
||||||
|
|
||||||
|
- 不能声称“单卡 BF16 完整模型运行”;
|
||||||
|
- 不能因为模型最终能生成,就隐去 CPU offload;
|
||||||
|
- 必须记录 `hf_device_map`、各设备参数 bytes、峰值 CUDA memory 和进程 RSS;
|
||||||
|
- 若 offload 失败,量化路线必须单独命名,不能与官方 BF16 混账。
|
||||||
|
|
||||||
|
首选执行:
|
||||||
|
|
||||||
|
```text
|
||||||
|
official BF16 weights
|
||||||
|
official modeling/configuration Python modules
|
||||||
|
eager attention(本机未安装 flash-attn,不伪装为 FlashAttention 路径)
|
||||||
|
GPU limit 29 GiB
|
||||||
|
CPU limit 80 GiB
|
||||||
|
Accelerate device_map=auto
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. 八格与上一轮保持同构
|
||||||
|
|
||||||
|
```text
|
||||||
|
system off/on × EOS / BOS / x / period
|
||||||
|
```
|
||||||
|
|
||||||
|
每个 source 的八格:
|
||||||
|
|
||||||
|
- 先由固定官方 chat template 渲染;
|
||||||
|
- EOS 为官方序列;
|
||||||
|
- BOS / `x` / 句点在 tokenization 后只改历史边界上的一个 input ID;
|
||||||
|
- 同一 system 水平内等长、同目标位置、同 generation prompt;
|
||||||
|
- system off/on 因 system 文本而不同长,batch 内使用 `PAD=EOS` 左填充且
|
||||||
|
`attention_mask=0`,让八格的 generation prefix 结束于同一列;
|
||||||
|
- 八个条件进入同一个 generation batch。
|
||||||
|
|
||||||
|
三种反事实不是官方合法 chat。它们只识别这个固定 pre-target ID 对生成的影响。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. source 身份与前轮如何衔接
|
||||||
|
|
||||||
|
source ID、selection rank 和每域顺序直接取自 base-routing 的 special-token 正式
|
||||||
|
实验,但行为探针使用**完整 source 文本**:
|
||||||
|
|
||||||
|
- route 实验固定目标前 23 tokens,利于等长统计;
|
||||||
|
- generation 需要完整问题,尤其数学与代码不能在半句话上算行为;
|
||||||
|
- 因此“同 source”不等于“同输入长度”,这一差异必须写进结果。
|
||||||
|
|
||||||
|
首个 smoke 只取每域 1 条,共 4 sources / 32 outputs。闸门通过后,正式运行扩到
|
||||||
|
每域 4 条、16 sources / 128 outputs。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. 解码合同
|
||||||
|
|
||||||
|
三道执行闸门固定:
|
||||||
|
|
||||||
|
```text
|
||||||
|
smoke = 1 source/domain × 8 conditions × 32 new tokens
|
||||||
|
formal = 4 sources/domain × 8 conditions × 128 new tokens
|
||||||
|
long repro = 1 source/domain × 8 conditions × 128 new tokens
|
||||||
|
|
||||||
|
do_sample = false
|
||||||
|
decode = greedy
|
||||||
|
use_cache = true
|
||||||
|
EOS / PAD = official tokenizer IDs
|
||||||
|
```
|
||||||
|
|
||||||
|
官方 `generation_config.json` 的 `temperature=.3 / top_p=.95` 只记录,不在首轮使用。
|
||||||
|
理由不是 greedy 更“真实”,而是它能提供确定性复跑闸门。之后的 sampled
|
||||||
|
distribution 必须另起协议、固定 seeds,并报告多次采样。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. 行为输出记什么
|
||||||
|
|
||||||
|
逐条件保存:
|
||||||
|
|
||||||
|
- prompt token hash;
|
||||||
|
- generated token IDs 与 hash;
|
||||||
|
- 解码文本与 hash;
|
||||||
|
- EOS 是否出现;
|
||||||
|
- 输出长度;
|
||||||
|
- 数学最后数值匹配;
|
||||||
|
- Python 输出 AST 是否可解析。
|
||||||
|
|
||||||
|
逐 source 比较:
|
||||||
|
|
||||||
|
- system edge:同一边界下 S0↔S1;
|
||||||
|
- direct edge:同一 system 下 EOS↔BOS/x/句点;
|
||||||
|
- token IDs 是否 exact;
|
||||||
|
- common-prefix tokens;
|
||||||
|
- token Levenshtein distance;
|
||||||
|
- normalized token similarity。
|
||||||
|
|
||||||
|
代码只解析,不执行。首轮也不把 1–4 个样本称为 benchmark accuracy。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. 通过闸门
|
||||||
|
|
||||||
|
smoke 至少需要:
|
||||||
|
|
||||||
|
1. 四个官方 shards 与 tokenizer/model code hash 固定;
|
||||||
|
2. Chat tokenizer 的四个边界 ID 与前轮完全相同;
|
||||||
|
3. 每个 source 在同一 system 水平内四个边界格等长;system off/on 的差长只由
|
||||||
|
system 文本造成,并由显式左填充吸收;
|
||||||
|
4. official 零 ID、counterfactual 恰好一 ID;
|
||||||
|
5. 32 个生成调用全部正常返回,无 OOM / NaN / runtime exception;是否自然命中 EOS
|
||||||
|
另作 completion 字段,不能把达到长度上限写成完整回答;
|
||||||
|
6. device map 与 offload 事实进入 JSON;
|
||||||
|
7. 独立复跑的 generated token IDs 可逐格比较;
|
||||||
|
8. 结果只命名为小规模 deterministic behavior probe。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. 暂时不允许的结论
|
||||||
|
|
||||||
|
- “EOS 让回答更好”;
|
||||||
|
- “special token 普遍优于普通 token”;
|
||||||
|
- “路由 TV 解释了输出差异”;
|
||||||
|
- “Chat SFT 让某个 expert 学会角色语义”;
|
||||||
|
- “4 条样本代表 HumanEval / GSM8K / 中英文能力”;
|
||||||
|
- “CPU offload 的延迟代表生产吞吐”;
|
||||||
|
- “greedy exact / non-exact 等于采样分布稳定 / 不稳定”。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. 执行入口
|
||||||
|
|
||||||
|
```bash
|
||||||
|
PYTHONPATH=/path/to/pinned/python-deps \
|
||||||
|
python -B \
|
||||||
|
experiments/deepseek/v2_lite_chat_special_token_behavior_probe.py \
|
||||||
|
--artifact-dir /path/to/deepseek-v2-lite-chat \
|
||||||
|
--reference-routing-json \
|
||||||
|
src/data/deepseek-v2-lite-routing-special-token-family-control.json \
|
||||||
|
--human-eval /path/to/HumanEval.jsonl.gz \
|
||||||
|
--gsm8k /path/to/gsm8k/test.jsonl \
|
||||||
|
--tnews /path/to/tnews/test.json \
|
||||||
|
--tnews-archive /path/to/tnews_public.zip \
|
||||||
|
--wikitext /path/to/wikitext-validation.parquet \
|
||||||
|
--output /tmp/deepseek-v2-lite-chat-behavior-smoke.json \
|
||||||
|
--per-domain 1 \
|
||||||
|
--max-new-tokens 32 \
|
||||||
|
--gpu-memory 29GiB \
|
||||||
|
--cpu-memory 80GiB
|
||||||
|
```
|
||||||
|
|
||||||
|
smoke 通过后,正式运行把 `--per-domain` / `--max-new-tokens` 改为 `4 / 128`;长序列
|
||||||
|
复跑保持 `1 / 128`,写入独立 JSON。正式与复跑都必须使用相同模型 revision、工件 hashes、
|
||||||
|
数据来源、依赖版本、设备映射上限与 greedy contract。
|
||||||
@@ -0,0 +1,278 @@
|
|||||||
|
import { createHash } from "node:crypto";
|
||||||
|
import { readFileSync, statSync, writeFileSync } from "node:fs";
|
||||||
|
import { resolve } from "node:path";
|
||||||
|
|
||||||
|
const root = resolve(import.meta.dirname, "..");
|
||||||
|
const formalPath = resolve(
|
||||||
|
root,
|
||||||
|
"src/data/deepseek-v2-lite-chat-behavior.json",
|
||||||
|
);
|
||||||
|
const reproPath = resolve(
|
||||||
|
root,
|
||||||
|
"src/data/deepseek-v2-lite-chat-behavior-repro-1pd.json",
|
||||||
|
);
|
||||||
|
const outputPath = resolve(
|
||||||
|
root,
|
||||||
|
"src/data/deepseek-v2-lite-chat-behavior-compact.json",
|
||||||
|
);
|
||||||
|
|
||||||
|
const sha256 = (path) => createHash("sha256")
|
||||||
|
.update(readFileSync(path))
|
||||||
|
.digest("hex");
|
||||||
|
|
||||||
|
const formal = JSON.parse(readFileSync(formalPath, "utf8"));
|
||||||
|
const repro = JSON.parse(readFileSync(reproPath, "utf8"));
|
||||||
|
const formalSha256 = sha256(formalPath);
|
||||||
|
const reproSha256 = sha256(reproPath);
|
||||||
|
|
||||||
|
const formalOutputByKey = new Map(
|
||||||
|
formal.sources.flatMap((source) => source.outputs.map((output) => [
|
||||||
|
`${source.id}\0${output.condition}`,
|
||||||
|
{ source, output },
|
||||||
|
])),
|
||||||
|
);
|
||||||
|
const reproduction = {
|
||||||
|
sources: repro.sources.length,
|
||||||
|
cells: 0,
|
||||||
|
promptHashExact: 0,
|
||||||
|
generatedTokenIdsExact: 0,
|
||||||
|
generatedTextExact: 0,
|
||||||
|
eosStateExact: 0,
|
||||||
|
};
|
||||||
|
for (const source of repro.sources) {
|
||||||
|
for (const output of source.outputs) {
|
||||||
|
const key = `${source.id}\0${output.condition}`;
|
||||||
|
const reference = formalOutputByKey.get(key);
|
||||||
|
if (!reference) throw new Error(`formal output missing: ${key}`);
|
||||||
|
reproduction.cells += 1;
|
||||||
|
reproduction.promptHashExact += (
|
||||||
|
reference.output.prompt_token_ids_sha256
|
||||||
|
=== output.prompt_token_ids_sha256
|
||||||
|
);
|
||||||
|
reproduction.generatedTokenIdsExact += (
|
||||||
|
JSON.stringify(reference.output.generated_token_ids)
|
||||||
|
=== JSON.stringify(output.generated_token_ids)
|
||||||
|
);
|
||||||
|
reproduction.generatedTextExact += (
|
||||||
|
reference.output.text === output.text
|
||||||
|
);
|
||||||
|
reproduction.eosStateExact += (
|
||||||
|
reference.output.hit_eos === output.hit_eos
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (
|
||||||
|
reproduction.cells !== 32
|
||||||
|
|| reproduction.promptHashExact !== reproduction.cells
|
||||||
|
|| reproduction.generatedTokenIdsExact !== reproduction.cells
|
||||||
|
|| reproduction.generatedTextExact !== reproduction.cells
|
||||||
|
|| reproduction.eosStateExact !== reproduction.cells
|
||||||
|
) {
|
||||||
|
throw new Error(
|
||||||
|
`Chat behavior reproduction mismatch: ${JSON.stringify(reproduction)}`,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
const formalModelHashes = Object.fromEntries(
|
||||||
|
Object.entries(formal.model.files).map(([name, value]) => [
|
||||||
|
name,
|
||||||
|
value.sha256,
|
||||||
|
]),
|
||||||
|
);
|
||||||
|
const reproModelHashes = Object.fromEntries(
|
||||||
|
Object.entries(repro.model.files).map(([name, value]) => [
|
||||||
|
name,
|
||||||
|
value.sha256,
|
||||||
|
]),
|
||||||
|
);
|
||||||
|
if (JSON.stringify(formalModelHashes) !== JSON.stringify(reproModelHashes)) {
|
||||||
|
throw new Error("formal and reproduction model-file hashes differ");
|
||||||
|
}
|
||||||
|
|
||||||
|
const conditions = formal.generation_contract.conditions;
|
||||||
|
const edgeOrder = [
|
||||||
|
"system_eos",
|
||||||
|
"system_bos",
|
||||||
|
"system_x",
|
||||||
|
"system_period",
|
||||||
|
"bos_at_s0",
|
||||||
|
"bos_at_s1",
|
||||||
|
"x_at_s0",
|
||||||
|
"x_at_s1",
|
||||||
|
"period_at_s0",
|
||||||
|
"period_at_s1",
|
||||||
|
];
|
||||||
|
const edgeLabels = {
|
||||||
|
system_eos: "System on − off · EOS",
|
||||||
|
system_bos: "System on − off · BOS",
|
||||||
|
system_x: "System on − off · x",
|
||||||
|
system_period: "System on − off · 句点",
|
||||||
|
bos_at_s0: "BOS − EOS · system off",
|
||||||
|
bos_at_s1: "BOS − EOS · system on",
|
||||||
|
x_at_s0: "x − EOS · system off",
|
||||||
|
x_at_s1: "x − EOS · system on",
|
||||||
|
period_at_s0: "句点 − EOS · system off",
|
||||||
|
period_at_s1: "句点 − EOS · system on",
|
||||||
|
};
|
||||||
|
|
||||||
|
const compactOutput = (output) => ({
|
||||||
|
condition: output.condition,
|
||||||
|
factors: output.factors,
|
||||||
|
promptTokens: output.prompt_tokens,
|
||||||
|
leftPaddingTokens: output.left_padding_tokens,
|
||||||
|
promptTokenIdsSha256: output.prompt_token_ids_sha256,
|
||||||
|
generatedTokens: output.generated_tokens,
|
||||||
|
generatedTokenIdsSha256: output.generated_token_ids_sha256,
|
||||||
|
hitEos: output.hit_eos,
|
||||||
|
stoppedAtMaxNewTokens: output.stopped_at_max_new_tokens,
|
||||||
|
text: output.text,
|
||||||
|
textSha256: output.text_sha256,
|
||||||
|
taskScore: output.task_score,
|
||||||
|
});
|
||||||
|
|
||||||
|
const compact = {
|
||||||
|
schemaVersion: 1,
|
||||||
|
source: {
|
||||||
|
formalSha256,
|
||||||
|
reproSha256,
|
||||||
|
formalBytes: statSync(formalPath).size,
|
||||||
|
reproBytes: statSync(reproPath).size,
|
||||||
|
reproduction,
|
||||||
|
},
|
||||||
|
model: {
|
||||||
|
repo: formal.model.repo,
|
||||||
|
revision: formal.model.revision,
|
||||||
|
checkpointIdentity: formal.model.checkpoint_identity,
|
||||||
|
architecture: formal.model.architecture,
|
||||||
|
dtype: formal.model.dtype,
|
||||||
|
checkpointTensorBytes: formal.model.checkpoint_tensor_bytes,
|
||||||
|
shardFileBytesIncludingHeaders: (
|
||||||
|
formal.model.shard_file_bytes_including_headers
|
||||||
|
),
|
||||||
|
allFilesSameRevision: (
|
||||||
|
formal.model.download_revision_contract.all_files_same_revision
|
||||||
|
),
|
||||||
|
fileCount: Object.keys(formal.model.files).length,
|
||||||
|
fileHashes: formalModelHashes,
|
||||||
|
},
|
||||||
|
tokenizer: formal.tokenizer_contract,
|
||||||
|
sources: formal.sources.map((source) => ({
|
||||||
|
id: source.id,
|
||||||
|
domain: source.domain,
|
||||||
|
label: source.label,
|
||||||
|
withinDomainIndex: source.within_domain_index,
|
||||||
|
selectionRank: source.selection_rank,
|
||||||
|
sourceTextSha256: source.source_text_sha256,
|
||||||
|
sourceCharacters: source.source_characters,
|
||||||
|
sourceTokens: source.source_tokens,
|
||||||
|
promptTokensByCondition: source.prompt_tokens_by_condition,
|
||||||
|
batchPromptTokens: source.batch_prompt_tokens_after_left_padding,
|
||||||
|
batchPaddingSide: source.batch_padding_side,
|
||||||
|
outputs: source.outputs.map(compactOutput),
|
||||||
|
})),
|
||||||
|
contract: {
|
||||||
|
source: formal.source_contract,
|
||||||
|
conditions,
|
||||||
|
edgeOrder,
|
||||||
|
edgeLabels,
|
||||||
|
officialSerialization: (
|
||||||
|
formal.generation_contract.official_serialization
|
||||||
|
),
|
||||||
|
decode: formal.generation_contract.decode,
|
||||||
|
doSample: formal.generation_contract.do_sample,
|
||||||
|
maxNewTokens: formal.generation_contract.max_new_tokens,
|
||||||
|
useCache: formal.generation_contract.use_cache,
|
||||||
|
batchPadding: formal.generation_contract.batch_padding,
|
||||||
|
counterfactualBoundary: (
|
||||||
|
formal.generation_contract.counterfactual_boundary
|
||||||
|
),
|
||||||
|
officialGenerationConfigRecordedNotUsed: (
|
||||||
|
formal.generation_contract
|
||||||
|
.official_generation_config_recorded_not_used
|
||||||
|
),
|
||||||
|
},
|
||||||
|
execution: {
|
||||||
|
python: formal.execution.python,
|
||||||
|
torch: formal.execution.torch,
|
||||||
|
transformers: formal.execution.transformers,
|
||||||
|
accelerate: formal.execution.accelerate,
|
||||||
|
safetensors: formal.execution.safetensors,
|
||||||
|
gpu: formal.execution.gpu,
|
||||||
|
gpuMemoryLimit: formal.execution.gpu_memory_limit,
|
||||||
|
cpuMemoryLimit: formal.execution.cpu_memory_limit,
|
||||||
|
inputDevice: formal.execution.input_device,
|
||||||
|
deviceMap: formal.execution.device_map,
|
||||||
|
parameterBytesByRuntimeParameterDevice: (
|
||||||
|
formal.execution.parameter_bytes_by_runtime_parameter_device
|
||||||
|
),
|
||||||
|
parameterBytesByDtype: formal.execution.parameter_bytes_by_dtype,
|
||||||
|
offloadParameterDeviceNote: (
|
||||||
|
formal.execution.offload_parameter_device_note
|
||||||
|
),
|
||||||
|
loadSeconds: formal.execution.load_seconds,
|
||||||
|
generationSeconds: formal.sources.reduce(
|
||||||
|
(sum, source) => sum + source.generation_seconds,
|
||||||
|
0,
|
||||||
|
),
|
||||||
|
loadPeakCudaMemoryAllocatedBytes: (
|
||||||
|
formal.execution.load_peak_cuda_memory_allocated_bytes
|
||||||
|
),
|
||||||
|
peakCudaMemoryAllocatedBytes: (
|
||||||
|
formal.execution.peak_cuda_memory_allocated_bytes
|
||||||
|
),
|
||||||
|
processMaxRssKib: formal.execution.process_max_rss_kib,
|
||||||
|
officialSingleGpuBf16Requirement: (
|
||||||
|
formal.execution.official_single_gpu_bf16_requirement
|
||||||
|
),
|
||||||
|
localSingleGpuCapacityMib: (
|
||||||
|
formal.execution.local_single_gpu_capacity_mib
|
||||||
|
),
|
||||||
|
offloadRequiredByLocalCapacity: (
|
||||||
|
formal.execution.offload_required_by_local_capacity
|
||||||
|
),
|
||||||
|
},
|
||||||
|
summary: {
|
||||||
|
aggregates: Object.fromEntries(
|
||||||
|
edgeOrder.map((edge) => [
|
||||||
|
edge,
|
||||||
|
formal.summary.aggregates[edge],
|
||||||
|
]),
|
||||||
|
),
|
||||||
|
aggregatesByDomain: formal.summary.aggregates_by_domain,
|
||||||
|
outputByCondition: formal.summary.output_by_condition,
|
||||||
|
taskByCondition: formal.summary.task_by_condition,
|
||||||
|
pairwise: Object.fromEntries(
|
||||||
|
edgeOrder.map((edge) => [
|
||||||
|
edge,
|
||||||
|
formal.summary.pairwise[edge],
|
||||||
|
]),
|
||||||
|
),
|
||||||
|
completedOutputs: formal.sources.reduce(
|
||||||
|
(sum, source) => sum + source.outputs.filter(
|
||||||
|
(output) => output.hit_eos,
|
||||||
|
).length,
|
||||||
|
0,
|
||||||
|
),
|
||||||
|
truncatedOutputs: formal.sources.reduce(
|
||||||
|
(sum, source) => sum + source.outputs.filter(
|
||||||
|
(output) => output.stopped_at_max_new_tokens,
|
||||||
|
).length,
|
||||||
|
0,
|
||||||
|
),
|
||||||
|
},
|
||||||
|
claimBoundary: formal.claim_boundary,
|
||||||
|
};
|
||||||
|
|
||||||
|
writeFileSync(
|
||||||
|
outputPath,
|
||||||
|
`${JSON.stringify(compact, null, 2)}\n`,
|
||||||
|
"utf8",
|
||||||
|
);
|
||||||
|
process.stdout.write(
|
||||||
|
`${outputPath}\n`
|
||||||
|
+ `${formalSha256}\n`
|
||||||
|
+ `${statSync(formalPath).size} bytes formal → `
|
||||||
|
+ `${statSync(outputPath).size} bytes compact\n`
|
||||||
|
+ `${reproduction.generatedTokenIdsExact}`
|
||||||
|
+ ` / ${reproduction.cells} long-sequence cells exact\n`,
|
||||||
|
);
|
||||||
@@ -76,6 +76,10 @@ const overview = await evaluate(`(() => ({
|
|||||||
artifactTabs: document.querySelectorAll("[data-artifact-tab]").length,
|
artifactTabs: document.querySelectorAll("[data-artifact-tab]").length,
|
||||||
artifactPanels: document.querySelectorAll("[data-artifact-panel]").length,
|
artifactPanels: document.querySelectorAll("[data-artifact-panel]").length,
|
||||||
artifactLayers: document.querySelectorAll(".layer-evidence > span").length,
|
artifactLayers: document.querySelectorAll(".layer-evidence > span").length,
|
||||||
|
behaviorTabs: document.querySelectorAll("[data-behavior-tab]").length,
|
||||||
|
behaviorPanels: document.querySelectorAll("[data-behavior-panel]").length,
|
||||||
|
behaviorSources: document.querySelectorAll("[data-behavior-source] option").length,
|
||||||
|
behaviorEdges: document.querySelectorAll("[data-behavior-map-edge]").length,
|
||||||
branches: document.querySelectorAll(".branch-grid > a").length,
|
branches: document.querySelectorAll(".branch-grid > a").length,
|
||||||
followups: document.querySelectorAll(".lineage-row.followup").length,
|
followups: document.querySelectorAll(".lineage-row.followup").length,
|
||||||
navLinks: document.querySelectorAll(".top-nav a").length,
|
navLinks: document.querySelectorAll(".top-nav a").length,
|
||||||
@@ -803,6 +807,78 @@ await evaluate(`(() => {
|
|||||||
await pause(180);
|
await pause(180);
|
||||||
await screenshot("/tmp/llm-atlas-deepseek-artifact-desktop.png");
|
await screenshot("/tmp/llm-atlas-deepseek-artifact-desktop.png");
|
||||||
|
|
||||||
|
const behavior = await evaluate(`(() => {
|
||||||
|
const root = document.querySelector("[data-behavior-lab]");
|
||||||
|
const readPair = () => ({
|
||||||
|
panel: root.querySelector("[data-behavior-panel]:not([hidden])").dataset.behaviorPanel,
|
||||||
|
source: root.querySelector("[data-behavior-source-id]").textContent.trim(),
|
||||||
|
domain: root.querySelector("[data-behavior-domain]").textContent.trim(),
|
||||||
|
exact: root.querySelector("[data-behavior-exact]").textContent.trim(),
|
||||||
|
prefix: root.querySelector("[data-behavior-prefix]").textContent.trim(),
|
||||||
|
edit: root.querySelector("[data-behavior-edit]").textContent.trim(),
|
||||||
|
similarity: root.querySelector("[data-behavior-similarity]").textContent.trim(),
|
||||||
|
conditions: [...root.querySelectorAll("[data-output-condition]")].map((node) => node.textContent.trim()),
|
||||||
|
statuses: [...root.querySelectorAll("[data-output-status]")].map((node) => node.textContent.trim()),
|
||||||
|
outputCharacters: [...root.querySelectorAll("[data-output-text]")].map((node) => node.textContent.length),
|
||||||
|
});
|
||||||
|
const initial = readPair();
|
||||||
|
const source = root.querySelector("[data-behavior-source]");
|
||||||
|
const edge = root.querySelector("[data-behavior-edge]");
|
||||||
|
source.value = "gsm8k/test/1069";
|
||||||
|
source.dispatchEvent(new Event("change", { bubbles: true }));
|
||||||
|
edge.value = "x_at_s1";
|
||||||
|
edge.dispatchEvent(new Event("change", { bubbles: true }));
|
||||||
|
const switched = readPair();
|
||||||
|
root.querySelector('[data-behavior-tab="map"]').click();
|
||||||
|
const map = root.querySelector("[data-behavior-map-domain]");
|
||||||
|
map.value = "math";
|
||||||
|
map.dispatchEvent(new Event("change", { bubbles: true }));
|
||||||
|
const mathMap = {
|
||||||
|
panel: root.querySelector("[data-behavior-panel]:not([hidden])").dataset.behaviorPanel,
|
||||||
|
rows: root.querySelectorAll("[data-behavior-map-edge]").length,
|
||||||
|
note: root.querySelector("[data-behavior-map-note]").textContent.trim(),
|
||||||
|
firstExact: root.querySelector("[data-behavior-map-edge] [data-map-exact]").textContent.trim(),
|
||||||
|
firstSimilarity: root.querySelector("[data-behavior-map-edge] [data-map-similarity]").textContent.trim(),
|
||||||
|
};
|
||||||
|
root.querySelector('[data-behavior-tab="execution"]').click();
|
||||||
|
const execution = {
|
||||||
|
panel: root.querySelector("[data-behavior-panel]:not([hidden])").dataset.behaviorPanel,
|
||||||
|
layers: root.querySelectorAll(".layer-device-map > span").length,
|
||||||
|
gpu: root.querySelectorAll(".layer-device-map > span.gpu").length,
|
||||||
|
cpu: root.querySelectorAll(".layer-device-map > span.cpu").length,
|
||||||
|
runtimeCards: root.querySelectorAll(".runtime-grid > article").length,
|
||||||
|
};
|
||||||
|
root.querySelector('[data-behavior-tab="boundary"]').click();
|
||||||
|
const boundary = {
|
||||||
|
panel: root.querySelector("[data-behavior-panel]:not([hidden])").dataset.behaviorPanel,
|
||||||
|
ladder: root.querySelectorAll(".evidence-ladder > article").length,
|
||||||
|
tokenCards: root.querySelectorAll(".token-contract > article").length,
|
||||||
|
reproCards: root.querySelectorAll(".repro-grid > article").length,
|
||||||
|
links: root.querySelectorAll(".artifact-links > a").length,
|
||||||
|
forbidden: root.querySelector(".forbidden-claims").textContent.replaceAll(/\\s+/g, " ").trim(),
|
||||||
|
};
|
||||||
|
const first = root.querySelector('[data-behavior-tab="pair"]');
|
||||||
|
first.focus();
|
||||||
|
first.dispatchEvent(new KeyboardEvent("keydown", { key: "ArrowRight", bubbles: true }));
|
||||||
|
return {
|
||||||
|
initial,
|
||||||
|
switched,
|
||||||
|
mathMap,
|
||||||
|
execution,
|
||||||
|
boundary,
|
||||||
|
keyboardSelected: root.querySelector('[data-behavior-tab][aria-selected="true"]').dataset.behaviorTab,
|
||||||
|
keyboardVisible: root.querySelector("[data-behavior-panel]:not([hidden])").dataset.behaviorPanel,
|
||||||
|
};
|
||||||
|
})()`);
|
||||||
|
await evaluate(`(() => {
|
||||||
|
const root = document.querySelector("[data-behavior-lab]");
|
||||||
|
root.querySelector('[data-behavior-tab="pair"]').click();
|
||||||
|
root.scrollIntoView({ block: "start", behavior: "instant" });
|
||||||
|
window.scrollBy(0, -82);
|
||||||
|
})()`);
|
||||||
|
await pause(180);
|
||||||
|
await screenshot("/tmp/llm-atlas-deepseek-behavior-desktop.png");
|
||||||
|
|
||||||
await navigate("/");
|
await navigate("/");
|
||||||
const home = await evaluate(`(() => ({
|
const home = await evaluate(`(() => ({
|
||||||
releaseCards: document.querySelectorAll(".release-card").length,
|
releaseCards: document.querySelectorAll(".release-card").length,
|
||||||
@@ -835,6 +911,7 @@ await navigate("/deepseek/");
|
|||||||
const mobile = await evaluate(`(() => {
|
const mobile = await evaluate(`(() => {
|
||||||
const root = document.querySelector("[data-deepseek-lab]");
|
const root = document.querySelector("[data-deepseek-lab]");
|
||||||
const artifact = document.querySelector("[data-dsv2-lab]");
|
const artifact = document.querySelector("[data-dsv2-lab]");
|
||||||
|
const behavior = document.querySelector("[data-behavior-lab]");
|
||||||
root.scrollIntoView({ block: "start", behavior: "instant" });
|
root.scrollIntoView({ block: "start", behavior: "instant" });
|
||||||
const toggle = document.querySelector("#menu-toggle");
|
const toggle = document.querySelector("#menu-toggle");
|
||||||
toggle?.click();
|
toggle?.click();
|
||||||
@@ -845,6 +922,10 @@ const mobile = await evaluate(`(() => {
|
|||||||
mobileLinks: document.querySelectorAll("#mobile-nav a").length,
|
mobileLinks: document.querySelectorAll("#mobile-nav a").length,
|
||||||
tabs: root.querySelectorAll("[data-ds-tab]").length,
|
tabs: root.querySelectorAll("[data-ds-tab]").length,
|
||||||
artifactTabs: artifact.querySelectorAll("[data-artifact-tab]").length,
|
artifactTabs: artifact.querySelectorAll("[data-artifact-tab]").length,
|
||||||
|
behaviorTabs: behavior.querySelectorAll("[data-behavior-tab]").length,
|
||||||
|
behaviorSources: behavior.querySelectorAll("[data-behavior-source] option").length,
|
||||||
|
behaviorEdges: behavior.querySelectorAll("[data-behavior-map-edge]").length,
|
||||||
|
behaviorDeviceCells: behavior.querySelectorAll(".layer-device-map > span").length,
|
||||||
artifactHeatCells: artifact.querySelectorAll("[data-route-heatmap] > span").length,
|
artifactHeatCells: artifact.querySelectorAll("[data-route-heatmap] > span").length,
|
||||||
corpusCohorts: artifact.querySelectorAll("[data-corpus-cohort]").length,
|
corpusCohorts: artifact.querySelectorAll("[data-corpus-cohort]").length,
|
||||||
lengthDeltaCards: artifact.querySelectorAll("[data-length-delta-grid] > article").length,
|
lengthDeltaCards: artifact.querySelectorAll("[data-length-delta-grid] > article").length,
|
||||||
@@ -894,7 +975,7 @@ const mobile = await evaluate(`(() => {
|
|||||||
roleBlockDomainCards: artifact.querySelectorAll("[data-role-block-domain-grid] > article").length,
|
roleBlockDomainCards: artifact.querySelectorAll("[data-role-block-domain-grid] > article").length,
|
||||||
roleBlockDepthCells: artifact.querySelectorAll("[data-role-block-depth-map] > div > span").length,
|
roleBlockDepthCells: artifact.querySelectorAll("[data-role-block-depth-map] > div > span").length,
|
||||||
offenders: [...document.querySelectorAll("body *")]
|
offenders: [...document.querySelectorAll("body *")]
|
||||||
.filter((node) => !node.closest(".paper-chain, .advantage-table, .precision-table, .mapping-table, [data-deepseek-lab], [data-dsv2-lab]"))
|
.filter((node) => !node.closest(".paper-chain, .advantage-table, .precision-table, .mapping-table, [data-deepseek-lab], [data-dsv2-lab], [data-behavior-lab]"))
|
||||||
.filter((node) => node.getBoundingClientRect().right > document.documentElement.clientWidth + 1)
|
.filter((node) => node.getBoundingClientRect().right > document.documentElement.clientWidth + 1)
|
||||||
.slice(0, 12)
|
.slice(0, 12)
|
||||||
.map((node) => ({
|
.map((node) => ({
|
||||||
@@ -1018,19 +1099,28 @@ await evaluate(`(() => {
|
|||||||
})()`);
|
})()`);
|
||||||
await pause(120);
|
await pause(120);
|
||||||
await screenshot("/tmp/llm-atlas-deepseek-role-block-results-mobile.png");
|
await screenshot("/tmp/llm-atlas-deepseek-role-block-results-mobile.png");
|
||||||
|
await evaluate(`(() => {
|
||||||
|
const behavior = document.querySelector("[data-behavior-lab]");
|
||||||
|
behavior.querySelector('[data-behavior-tab="pair"]').click();
|
||||||
|
behavior.scrollIntoView({ block: "start", behavior: "instant" });
|
||||||
|
window.scrollBy(0, -70);
|
||||||
|
})()`);
|
||||||
|
await pause(180);
|
||||||
|
await screenshot("/tmp/llm-atlas-deepseek-behavior-mobile.png");
|
||||||
|
|
||||||
const report = { overview, capacity, cache, codesign, rl, artifactRoute, artifactLoad, artifactCache, artifactAbsorb, artifactCorpus, artifactTemplate, artifactHistory, artifactDistance, artifactBoundary, artifactRole, artifactSpecial, artifactRoleBlock, artifactEvidence, home, papers, mobile, exceptions };
|
const report = { overview, capacity, cache, codesign, rl, artifactRoute, artifactLoad, artifactCache, artifactAbsorb, artifactCorpus, artifactTemplate, artifactHistory, artifactDistance, artifactBoundary, artifactRole, artifactSpecial, artifactRoleBlock, artifactEvidence, behavior, home, papers, mobile, exceptions };
|
||||||
console.log(JSON.stringify(report, null, 2));
|
console.log(JSON.stringify(report, null, 2));
|
||||||
|
|
||||||
const numeric = (text) => Number.parseFloat(text.replaceAll(",", ""));
|
const numeric = (text) => Number.parseFloat(text.replaceAll(",", ""));
|
||||||
const failures = [];
|
const failures = [];
|
||||||
if (!overview.title.includes("为什么转向")) failures.push("专题标题异常");
|
if (!overview.title.includes("为什么转向")) failures.push("专题标题异常");
|
||||||
if (overview.sections !== 26 || overview.tocLinks !== 26) failures.push("二十五个编号专题加阅读链的目录结构异常");
|
if (overview.sections !== 27 || overview.tocLinks !== 27) failures.push("二十六个编号专题加阅读链的目录结构异常");
|
||||||
if (overview.ledgers !== 24 || overview.waves !== 10) failures.push("二十四张问题账或十次转向结构异常");
|
if (overview.ledgers !== 24 || overview.waves !== 10) failures.push("二十四张问题账或十次转向结构异常");
|
||||||
if (overview.paperLinks !== 60 || overview.branches !== 5 || overview.followups !== 1) failures.push("论文链、旁支或公开后续标记异常");
|
if (overview.paperLinks !== 60 || overview.branches !== 5 || overview.followups !== 1) failures.push("论文链、旁支或公开后续标记异常");
|
||||||
if (overview.labTabs !== 4 || overview.labPanels !== 4) failures.push("四联实验结构异常");
|
if (overview.labTabs !== 4 || overview.labPanels !== 4) failures.push("四联实验结构异常");
|
||||||
if (overview.artifactTabs !== 13 || overview.artifactPanels !== 13 || overview.artifactLayers !== 27) failures.push("真实权重十三联实验结构异常");
|
if (overview.artifactTabs !== 13 || overview.artifactPanels !== 13 || overview.artifactLayers !== 27) failures.push("真实权重十三联实验结构异常");
|
||||||
if (overview.heroLabs !== "17 个可操作实验") failures.push("DeepSeek 实验总数账异常");
|
if (overview.behaviorTabs !== 4 || overview.behaviorPanels !== 4 || overview.behaviorSources !== 16 || overview.behaviorEdges !== 10) failures.push("Chat 行为实验结构异常");
|
||||||
|
if (overview.heroLabs !== "18 个可操作实验") failures.push("DeepSeek 实验总数账异常");
|
||||||
if (overview.navLinks !== 20 || home.navLinks !== 20 || mobile.mobileLinks !== 20 || overview.activeNav !== "DeepSeek") failures.push("全站导航未同步 DeepSeek");
|
if (overview.navLinks !== 20 || home.navLinks !== 20 || mobile.mobileLinks !== 20 || overview.activeNav !== "DeepSeek") failures.push("全站导航未同步 DeepSeek");
|
||||||
if (overview.documentOverflow > 1 || mobile.documentOverflow > 1) failures.push("桌面或移动端存在文档级横向溢出");
|
if (overview.documentOverflow > 1 || mobile.documentOverflow > 1) failures.push("桌面或移动端存在文档级横向溢出");
|
||||||
if (capacity.initial.panel !== "capacity" || capacity.initial.total !== "32.1× FFN" || capacity.initial.active !== "1.13× FFN") failures.push("V3 稀疏容量初始账异常");
|
if (capacity.initial.panel !== "capacity" || capacity.initial.total !== "32.1× FFN" || capacity.initial.active !== "1.13× FFN") failures.push("V3 稀疏容量初始账异常");
|
||||||
@@ -1117,9 +1207,15 @@ if (artifactRoleBlock.layer1Head.batchValues.join("|") !== "512 / 512|368 / 512|
|
|||||||
if (artifactEvidence.panel !== "evidence" || artifactEvidence.layers !== 27 || artifactEvidence.executed !== 7 || artifactEvidence.split !== 1 || artifactEvidence.unloaded !== 19 || artifactEvidence.exact !== "31 / 31") failures.push("真实工件执行边界或复跑闸门异常");
|
if (artifactEvidence.panel !== "evidence" || artifactEvidence.layers !== 27 || artifactEvidence.executed !== 7 || artifactEvidence.split !== 1 || artifactEvidence.unloaded !== 19 || artifactEvidence.exact !== "31 / 31") failures.push("真实工件执行边界或复跑闸门异常");
|
||||||
if (!artifactEvidence.dependency.includes("Transformers 5.5") || !artifactEvidence.dependency.includes("4.41.2") || !artifactEvidence.boundary.includes("完整 27 层生成")) failures.push("依赖版本或未覆盖边界异常");
|
if (!artifactEvidence.dependency.includes("Transformers 5.5") || !artifactEvidence.dependency.includes("4.41.2") || !artifactEvidence.boundary.includes("完整 27 层生成")) failures.push("依赖版本或未覆盖边界异常");
|
||||||
if (artifactEvidence.keyboardSelected !== "load" || artifactEvidence.keyboardVisible !== "load") failures.push("真实工件实验键盘 tab 导航异常");
|
if (artifactEvidence.keyboardSelected !== "load" || artifactEvidence.keyboardVisible !== "load") failures.push("真实工件实验键盘 tab 导航异常");
|
||||||
|
if (behavior.initial.panel !== "pair" || behavior.initial.source !== "wikitext2/raw-validation/0443" || behavior.initial.conditions.join("|") !== "S0 · EOS|S1 · EOS" || behavior.initial.exact !== "DIVERGED" || behavior.initial.prefix !== "63" || behavior.initial.edit !== "33" || behavior.initial.similarity !== "74.2%") failures.push("Chat 行为逐格输出初值异常");
|
||||||
|
if (behavior.switched.domain !== "Grade-school math" || behavior.switched.conditions.join("|") !== "S1 · EOS|S1 · x" || behavior.switched.exact !== "DIVERGED" || behavior.switched.outputCharacters.some((value) => value < 100)) failures.push("Chat 行为 source / contrast 切换异常");
|
||||||
|
if (behavior.mathMap.panel !== "map" || behavior.mathMap.rows !== 10 || !behavior.mathMap.note.includes("GSM8K") || behavior.mathMap.firstExact !== "2 / 4" || behavior.mathMap.firstSimilarity !== "84.0%") failures.push("Chat 行为分域分叉地图异常");
|
||||||
|
if (behavior.execution.panel !== "execution" || behavior.execution.layers !== 29 || behavior.execution.gpu !== 25 || behavior.execution.cpu !== 4 || behavior.execution.runtimeCards !== 4) failures.push("Chat BF16 GPU / CPU offload 设备图异常");
|
||||||
|
if (behavior.boundary.panel !== "boundary" || behavior.boundary.ladder !== 3 || behavior.boundary.tokenCards !== 4 || behavior.boundary.reproCards !== 4 || behavior.boundary.links !== 3 || !behavior.boundary.forbidden.includes("route TV")) failures.push("Chat 行为证据阶梯或限制异常");
|
||||||
|
if (behavior.keyboardSelected !== "map" || behavior.keyboardVisible !== "map") failures.push("Chat 行为实验键盘 tab 导航异常");
|
||||||
if (home.releaseCards !== 17 || !home.firstRelease.includes("47 页不再压成摘要") || home.firstHref !== "/k3/" || home.paperCount !== "486") failures.push("首页 DeepSeek 首发入口或论文数异常");
|
if (home.releaseCards !== 17 || !home.firstRelease.includes("47 页不再压成摘要") || home.firstHref !== "/k3/" || home.paperCount !== "486") failures.push("首页 DeepSeek 首发入口或论文数异常");
|
||||||
if (papers.total !== 486 || !papers.hasFilter || papers.visible < 20 || !papers.hasCoder || !papers.hasEngram) failures.push("论文库 DeepSeek 聚光异常");
|
if (papers.total !== 486 || !papers.hasFilter || papers.visible < 20 || !papers.hasCoder || !papers.hasEngram) failures.push("论文库 DeepSeek 聚光异常");
|
||||||
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4 || mobile.artifactTabs !== 13 || mobile.artifactHeatCells !== 64 || mobile.corpusCohorts !== 3 || mobile.lengthDeltaCards !== 4 || mobile.templateLayers !== 6 || mobile.templateScopes !== 2 || mobile.templateModes !== 2 || mobile.templateDomainCards !== 4 || mobile.templateDepthCells !== 24 || mobile.historyLayers !== 6 || mobile.historyScopes !== 2 || mobile.historyModes !== 2 || mobile.historyEffects !== 3 || mobile.historyDomainCards !== 4 || mobile.historyDepthCells !== 24 || mobile.distanceLayers !== 6 || mobile.distanceScopes !== 2 || mobile.distanceModes !== 2 || mobile.distanceContrasts !== 2 || mobile.distanceDomainCards !== 4 || mobile.distanceDepthCells !== 24 || mobile.boundaryLayers !== 6 || mobile.boundaryScopes !== 2 || mobile.boundaryModes !== 2 || mobile.boundaryContrasts !== 3 || mobile.boundaryTokenCards !== 4 || mobile.boundaryDomainCards !== 4 || mobile.boundaryDepthCells !== 24 || mobile.roleLayers !== 6 || mobile.roleScopes !== 2 || mobile.roleModes !== 2 || mobile.roleContrasts !== 3 || mobile.roleLevelCards !== 4 || mobile.roleDomainCards !== 4 || mobile.roleDepthCells !== 24 || mobile.specialLayers !== 6 || mobile.specialScopes !== 2 || mobile.specialModes !== 2 || mobile.specialContrasts !== 4 || mobile.specialTokenCards !== 4 || mobile.specialDomainCards !== 4 || mobile.specialDepthCells !== 24 || mobile.roleBlockLayers !== 6 || mobile.roleBlockScopes !== 2 || mobile.roleBlockModes !== 2 || mobile.roleBlockEffects !== 3 || mobile.roleBlockMatrixCards !== 4 || mobile.roleBlockDomainCards !== 4 || mobile.roleBlockDepthCells !== 24) failures.push("移动端导航或实验异常");
|
if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 4 || mobile.artifactTabs !== 13 || mobile.behaviorTabs !== 4 || mobile.behaviorSources !== 16 || mobile.behaviorEdges !== 10 || mobile.behaviorDeviceCells !== 29 || mobile.artifactHeatCells !== 64 || mobile.corpusCohorts !== 3 || mobile.lengthDeltaCards !== 4 || mobile.templateLayers !== 6 || mobile.templateScopes !== 2 || mobile.templateModes !== 2 || mobile.templateDomainCards !== 4 || mobile.templateDepthCells !== 24 || mobile.historyLayers !== 6 || mobile.historyScopes !== 2 || mobile.historyModes !== 2 || mobile.historyEffects !== 3 || mobile.historyDomainCards !== 4 || mobile.historyDepthCells !== 24 || mobile.distanceLayers !== 6 || mobile.distanceScopes !== 2 || mobile.distanceModes !== 2 || mobile.distanceContrasts !== 2 || mobile.distanceDomainCards !== 4 || mobile.distanceDepthCells !== 24 || mobile.boundaryLayers !== 6 || mobile.boundaryScopes !== 2 || mobile.boundaryModes !== 2 || mobile.boundaryContrasts !== 3 || mobile.boundaryTokenCards !== 4 || mobile.boundaryDomainCards !== 4 || mobile.boundaryDepthCells !== 24 || mobile.roleLayers !== 6 || mobile.roleScopes !== 2 || mobile.roleModes !== 2 || mobile.roleContrasts !== 3 || mobile.roleLevelCards !== 4 || mobile.roleDomainCards !== 4 || mobile.roleDepthCells !== 24 || mobile.specialLayers !== 6 || mobile.specialScopes !== 2 || mobile.specialModes !== 2 || mobile.specialContrasts !== 4 || mobile.specialTokenCards !== 4 || mobile.specialDomainCards !== 4 || mobile.specialDepthCells !== 24 || mobile.roleBlockLayers !== 6 || mobile.roleBlockScopes !== 2 || mobile.roleBlockModes !== 2 || mobile.roleBlockEffects !== 3 || mobile.roleBlockMatrixCards !== 4 || mobile.roleBlockDomainCards !== 4 || mobile.roleBlockDepthCells !== 24) failures.push("移动端导航或实验异常");
|
||||||
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`);
|
||||||
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`);
|
||||||
|
|
||||||
|
|||||||
@@ -1329,7 +1329,7 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
|
|||||||
<article>
|
<article>
|
||||||
<span>CHECKPOINT BOUNDARY</span>
|
<span>CHECKPOINT BOUNDARY</span>
|
||||||
<b>BASE ≠ CHAT / SFT</b>
|
<b>BASE ≠ CHAT / SFT</b>
|
||||||
<p>不能把较小 TV 命名为“理解回合结束”;仍需 V2-Lite-Chat 与行为生成对照。</p>
|
<p>不能把较小 TV 命名为“理解回合结束”;下方 Round 04 已另用 V2-Lite-Chat 做实际生成对照。</p>
|
||||||
</article>
|
</article>
|
||||||
</div>
|
</div>
|
||||||
|
|
||||||
@@ -1345,7 +1345,7 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
|
|||||||
<b>SINGLE-ID ROUTING CAUSALITY, NOT TURN-SEMANTIC OR CAPABILITY PROOF</b>
|
<b>SINGLE-ID ROUTING CAUSALITY, NOT TURN-SEMANTIC OR CAPABILITY PROOF</b>
|
||||||
<p>
|
<p>
|
||||||
EOS 条件下后续目标路由对 system 开关更稳定,但本实验既未生成答案,也未覆盖 Chat 权重;
|
EOS 条件下后续目标路由对 system 开关更稳定,但本实验既未生成答案,也未覆盖 Chat 权重;
|
||||||
更小 TV 不等于更正确。下一步要拆 `User:` 角色标记、special-token 家族与行为指标。
|
更小 TV 不等于更正确。角色标记、special-token 家族与 Chat 行为已由后续实验逐层拆开。
|
||||||
</p>
|
</p>
|
||||||
</div>
|
</div>
|
||||||
</section>
|
</section>
|
||||||
@@ -1496,7 +1496,7 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
|
|||||||
<b>ONE ROLE-HEAD ID, NOT COMPLETE ROLE SEMANTICS</b>
|
<b>ONE ROLE-HEAD ID, NOT COMPLETE ROLE SEMANTICS</b>
|
||||||
<p>
|
<p>
|
||||||
前置词头会改变后续路由,但没有统一调制 system;后置负对照严格为零。
|
前置词头会改变后续路由,但没有统一调制 system;后置负对照严格为零。
|
||||||
special-token family 与完整两-token 角色块已在后续页签闭环;下一步转向 Chat 权重与行为指标。
|
special-token family 与完整两-token 角色块已在后续页签闭环;Chat 权重与行为指标见下方 Round 04。
|
||||||
</p>
|
</p>
|
||||||
</div>
|
</div>
|
||||||
</section>
|
</section>
|
||||||
@@ -1841,7 +1841,7 @@ const shardFraction = trace.provenance.shard_1_bytes / trace.provenance.checkpoi
|
|||||||
|
|
||||||
<div class="artifact-boundary">
|
<div class="artifact-boundary">
|
||||||
<b>U / STILL OPEN</b>
|
<b>U / STILL OPEN</b>
|
||||||
<p>完整 27 层生成、受支持硬件上的 FlashMLA 优化 kernel、生产服务、训练负载、FP8/pipeline 与 R1-like 训练 trace 仍未覆盖。</p>
|
<p>Base checkpoint 的完整 27 层路由 trace、受支持硬件上的 FlashMLA 优化 kernel、生产服务、训练负载、FP8/pipeline 与 R1-like 训练 trace 仍未覆盖;Chat 的完整 27 层生成已在 Round 04 以 GPU+CPU offload 单独执行。</p>
|
||||||
</div>
|
</div>
|
||||||
</section>
|
</section>
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,655 @@
|
|||||||
|
---
|
||||||
|
import rawBehavior from "@/data/deepseek-v2-lite-chat-behavior-compact.json";
|
||||||
|
|
||||||
|
const behavior = rawBehavior as any;
|
||||||
|
const json = JSON.stringify(behavior).replaceAll("<", "\\u003c");
|
||||||
|
const gib = (value: number) => value / 1024 ** 3;
|
||||||
|
const gpuParameterBytes = (
|
||||||
|
behavior.execution.parameterBytesByRuntimeParameterDevice["cuda:0"]
|
||||||
|
);
|
||||||
|
const cpuOffloadBytes = (
|
||||||
|
behavior.execution.parameterBytesByRuntimeParameterDevice.meta
|
||||||
|
);
|
||||||
|
const peakGiB = gib(behavior.execution.peakCudaMemoryAllocatedBytes);
|
||||||
|
const checkpointGiB = gib(behavior.model.checkpointTensorBytes);
|
||||||
|
const gpuParameterGiB = gib(gpuParameterBytes);
|
||||||
|
const cpuOffloadGiB = gib(cpuOffloadBytes);
|
||||||
|
const rssGiB = behavior.execution.processMaxRssKib / 1024 ** 2;
|
||||||
|
const edgeOrder = behavior.contract.edgeOrder as string[];
|
||||||
|
const edgeLabels = behavior.contract.edgeLabels as Record<string, string>;
|
||||||
|
const firstSource = behavior.sources[0];
|
||||||
|
const conditionLabel: Record<string, string> = {
|
||||||
|
s0_eos: "S0 · EOS",
|
||||||
|
s1_eos: "S1 · EOS",
|
||||||
|
s0_bos: "S0 · BOS",
|
||||||
|
s1_bos: "S1 · BOS",
|
||||||
|
s0_x: "S0 · x",
|
||||||
|
s1_x: "S1 · x",
|
||||||
|
s0_period: "S0 · 句点",
|
||||||
|
s1_period: "S1 · 句点",
|
||||||
|
};
|
||||||
|
const deviceLayers = Array.from({ length: 27 }, (_, layer) => ({
|
||||||
|
layer,
|
||||||
|
device: behavior.execution.deviceMap[`model.layers.${layer}`],
|
||||||
|
}));
|
||||||
|
---
|
||||||
|
|
||||||
|
<figure class="behavior-lab" data-behavior-lab>
|
||||||
|
<header class="behavior-head">
|
||||||
|
<div>
|
||||||
|
<p>ROUND 04 / CHAT BEHAVIOR · OFFICIAL BF16</p>
|
||||||
|
<h3>路由变了以后,模型最后真的会说出不同答案吗?</h3>
|
||||||
|
</div>
|
||||||
|
<p>
|
||||||
|
同一组 source 从 Base checkpoint 的路由显微镜进入
|
||||||
|
<code>DeepSeek-V2-Lite-Chat</code>:system off/on × EOS/BOS/x/句点,
|
||||||
|
每条 source 的八格在同一 batch 内做 greedy generation。这里观察的是最终输出,
|
||||||
|
不再用 route TV 代替行为。
|
||||||
|
</p>
|
||||||
|
</header>
|
||||||
|
|
||||||
|
<div class="behavior-ledger">
|
||||||
|
<article><span>CHECKPOINT</span><b>SFT CHAT · BF16</b><p>不是 Base,也不是 R1 / RL</p></article>
|
||||||
|
<article><span>FORMAL GRID</span><b>16 × 8 = 128</b><p>四域各 4 条完整 source</p></article>
|
||||||
|
<article class="complete"><span>EOS COMPLETE</span><b>{behavior.summary.completedOutputs} / 128</b><p>在 128-token 上限内结束</p></article>
|
||||||
|
<article class="warning"><span>TRUNCATED</span><b>{behavior.summary.truncatedOutputs} / 128</b><p>能力分数不得横向解释</p></article>
|
||||||
|
<article><span>LONG RERUN</span><b>32 / 32 EXACT</b><p>每域首条逐 token 复跑</p></article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="behavior-tabs" role="tablist" aria-label="选择 Chat 行为证据视图">
|
||||||
|
<button type="button" role="tab" data-behavior-tab="pair" aria-selected="true">
|
||||||
|
<span>01</span><b>逐格读输出</b><small>source × contrast</small>
|
||||||
|
</button>
|
||||||
|
<button type="button" role="tab" data-behavior-tab="map" aria-selected="false" tabindex="-1">
|
||||||
|
<span>02</span><b>分叉地图</b><small>10 edges × 4 domains</small>
|
||||||
|
</button>
|
||||||
|
<button type="button" role="tab" data-behavior-tab="execution" aria-selected="false" tabindex="-1">
|
||||||
|
<span>03</span><b>完整权重怎样装下</b><small>GPU + CPU offload</small>
|
||||||
|
</button>
|
||||||
|
<button type="button" role="tab" data-behavior-tab="boundary" aria-selected="false" tabindex="-1">
|
||||||
|
<span>04</span><b>证据边界</b><small>route ≠ output ≠ ability</small>
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<section class="behavior-panel" data-behavior-panel="pair">
|
||||||
|
<div class="panel-lead">
|
||||||
|
<div><span>X / GENERATED OUTPUTS</span><h4>固定一条 source,再沿一条边比较两格</h4></div>
|
||||||
|
<p>Exact 表示整段 generated token IDs 完全相同;similarity 是 token Levenshtein 相似度,不是语义得分。</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="pair-controls">
|
||||||
|
<label>
|
||||||
|
<span>source · 16</span>
|
||||||
|
<select data-behavior-source aria-label="选择生成 source">
|
||||||
|
{behavior.sources.map((source: any) => (
|
||||||
|
<option value={source.id}>
|
||||||
|
{source.label} · {source.withinDomainIndex + 1} · {source.id}
|
||||||
|
</option>
|
||||||
|
))}
|
||||||
|
</select>
|
||||||
|
</label>
|
||||||
|
<label>
|
||||||
|
<span>contrast · 10</span>
|
||||||
|
<select data-behavior-edge aria-label="选择输出比较边">
|
||||||
|
{edgeOrder.map((edge) => (
|
||||||
|
<option value={edge}>{edgeLabels[edge]}</option>
|
||||||
|
))}
|
||||||
|
</select>
|
||||||
|
</label>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="source-contract">
|
||||||
|
<div><span>SOURCE</span><b data-behavior-source-id>{firstSource.id}</b></div>
|
||||||
|
<div><span>DOMAIN</span><b data-behavior-domain>{firstSource.label}</b></div>
|
||||||
|
<div><span>FULL INPUT</span><b data-behavior-source-shape>{firstSource.sourceCharacters} chars · {firstSource.sourceTokens} tokens</b></div>
|
||||||
|
<div><span>TEXT SHA-256</span><code data-behavior-source-hash>{firstSource.sourceTextSha256.slice(0, 16)}…</code></div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="pair-metrics">
|
||||||
|
<article><span>TOKEN EXACT</span><b data-behavior-exact>—</b></article>
|
||||||
|
<article><span>COMMON PREFIX</span><b data-behavior-prefix>—</b><small>tokens</small></article>
|
||||||
|
<article><span>EDIT DISTANCE</span><b data-behavior-edit>—</b><small>tokens</small></article>
|
||||||
|
<article><span>NORMALIZED SIMILARITY</span><b data-behavior-similarity>—</b></article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="output-pair">
|
||||||
|
{(["left", "right"] as const).map((side) => (
|
||||||
|
<article data-output-card={side}>
|
||||||
|
<header>
|
||||||
|
<div><span data-output-side={side}>{side === "left" ? "LEFT" : "RIGHT"}</span><b data-output-condition={side}>{conditionLabel[side === "left" ? "s0_eos" : "s1_eos"]}</b></div>
|
||||||
|
<em data-output-status={side}>—</em>
|
||||||
|
</header>
|
||||||
|
<dl>
|
||||||
|
<div><dt>PROMPT</dt><dd data-output-prompt={side}>—</dd></div>
|
||||||
|
<div><dt>GENERATED</dt><dd data-output-tokens={side}>—</dd></div>
|
||||||
|
<div><dt>TOKEN HASH</dt><dd><code data-output-hash={side}>—</code></dd></div>
|
||||||
|
</dl>
|
||||||
|
<p data-output-text={side}></p>
|
||||||
|
</article>
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="pair-reading">
|
||||||
|
<span>怎么读</span>
|
||||||
|
<p data-behavior-reading>
|
||||||
|
一处历史边界 ID 或一条 system message 可以让 greedy 轨迹分叉;分叉只说明这条固定输入、固定 checkpoint、固定解码路径发生变化。
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="behavior-panel" data-behavior-panel="map" hidden>
|
||||||
|
<div class="panel-lead">
|
||||||
|
<div><span>X / DESCRIPTIVE MAP</span><h4>十条边,不要压成一句“有影响 / 没影响”</h4></div>
|
||||||
|
<p>每域只有 4 条 source;下表是完整 token exact 数与平均编辑相似度,适合定位分叉,不是总体效应估计。</p>
|
||||||
|
</div>
|
||||||
|
<div class="map-control">
|
||||||
|
<label>
|
||||||
|
<span>聚合范围</span>
|
||||||
|
<select data-behavior-map-domain aria-label="选择分叉地图聚合范围">
|
||||||
|
<option value="all">四域合计 · n=16</option>
|
||||||
|
<option value="english">英文百科 · n=4</option>
|
||||||
|
<option value="chinese">中文新闻 · n=4</option>
|
||||||
|
<option value="code">Python 代码 · n=4</option>
|
||||||
|
<option value="math">小学数学 · n=4</option>
|
||||||
|
</select>
|
||||||
|
</label>
|
||||||
|
<p data-behavior-map-note>四域合计;不同域的输出长度与完成率不等,均值只作导航。</p>
|
||||||
|
</div>
|
||||||
|
<div class="edge-map" data-behavior-edge-map>
|
||||||
|
<header><b>CONTRAST</b><b>EXACT</b><b>MEAN SIMILARITY</b><b>EDIT</b></header>
|
||||||
|
{edgeOrder.map((edge) => (
|
||||||
|
<button type="button" data-behavior-map-edge={edge}>
|
||||||
|
<span>{edgeLabels[edge]}</span>
|
||||||
|
<b data-map-exact>—</b>
|
||||||
|
<i><em data-map-meter></em></i>
|
||||||
|
<strong data-map-similarity>—</strong>
|
||||||
|
<small data-map-edit>—</small>
|
||||||
|
</button>
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
<div class="map-callout">
|
||||||
|
<article><span>OFFICIAL EDGE</span><b>System · EOS</b><p>16 条里只有 2 条完整输出 exact;平均 token similarity 48.3%。</p></article>
|
||||||
|
<article><span>DIRECT COUNTERFACTUAL</span><b>x · system on</b><p>0 / 16 exact;平均 similarity 29.8%。这是固定 ID 对照,不是“x 更坏”。</p></article>
|
||||||
|
<article><span>DO NOT RANK</span><b>97 / 128 truncated</b><p>不同格的数学 exact / code parse 不满足公平能力比较的完成合同。</p></article>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="behavior-panel" data-behavior-panel="execution" hidden>
|
||||||
|
<div class="panel-lead">
|
||||||
|
<div><span>X / FULL CHECKPOINT</span><h4>31.4GB BF16 权重怎样在 32GB 卡上完成生成</h4></div>
|
||||||
|
<p>官方 model card 给出单卡 BF16 需要 40GB;本机可见 32,607 MiB,因此明确使用 Accelerate GPU+CPU offload。</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="checkpoint-ledger">
|
||||||
|
<article><span>INDEX TENSOR BYTES</span><b>{behavior.model.checkpointTensorBytes.toLocaleString()}</b><p>{checkpointGiB.toFixed(3)} GiB · 全 BF16</p></article>
|
||||||
|
<article><span>SHARD FILE BYTES</span><b>{behavior.model.shardFileBytesIncludingHeaders.toLocaleString()}</b><p>含 safetensors headers</p></article>
|
||||||
|
<article><span>PINNED REVISION</span><code>{behavior.model.revision.slice(0, 12)}…</code><p>12 / 12 文件同 revision</p></article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="offload-flow">
|
||||||
|
<article class="gpu">
|
||||||
|
<header><span>CUDA:0 · RTX 5090</span><b>{gpuParameterGiB.toFixed(2)} GiB parameters</b></header>
|
||||||
|
<div class="memory-track"><i style={`--fill:${Math.min(100, peakGiB / (behavior.execution.localSingleGpuCapacityMib / 1024) * 100).toFixed(2)}%`}></i></div>
|
||||||
|
<p>峰值 CUDA allocation <b>{peakGiB.toFixed(2)} GiB</b>;29GiB 是参数放置上限,KV/cache 会继续占显存。</p>
|
||||||
|
</article>
|
||||||
|
<div class="offload-arrow"><span>Accelerate</span><b>↔</b><small>device_map=auto</small></div>
|
||||||
|
<article class="cpu">
|
||||||
|
<header><span>CPU OFFLOAD</span><b>{cpuOffloadGiB.toFixed(2)} GiB tensors</b></header>
|
||||||
|
<div class="memory-track"><i style={`--fill:${Math.min(100, rssGiB / 80 * 100).toFixed(2)}%`}></i></div>
|
||||||
|
<p>layers 25–26、final norm 与 LM head;进程 max RSS <b>{rssGiB.toFixed(2)} GiB</b>。</p>
|
||||||
|
</article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="layer-device-map" aria-label="27 层运行设备图">
|
||||||
|
{deviceLayers.map(({ layer, device }) => (
|
||||||
|
<span class={device === "cpu" ? "cpu" : "gpu"} title={`layer ${layer} → ${device === "cpu" ? "CPU" : "CUDA:0"}`}>
|
||||||
|
<b>L{layer}</b><small>{device === "cpu" ? "CPU" : "GPU"}</small>
|
||||||
|
</span>
|
||||||
|
))}
|
||||||
|
<span class="cpu tail"><b>NORM</b><small>CPU</small></span>
|
||||||
|
<span class="cpu tail"><b>HEAD</b><small>CPU</small></span>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="runtime-grid">
|
||||||
|
<article><span>LOAD</span><b>{behavior.execution.loadSeconds.toFixed(2)} s</b><p>4 shards → device map</p></article>
|
||||||
|
<article><span>GENERATION</span><b>{behavior.execution.generationSeconds.toFixed(1)} s</b><p>16 source batches · 128 outputs</p></article>
|
||||||
|
<article><span>STACK</span><b>torch {behavior.execution.torch}</b><p>Transformers {behavior.execution.transformers} · Accelerate {behavior.execution.accelerate}</p></article>
|
||||||
|
<article class="warning"><span>NOT THROUGHPUT</span><b>CPU-offloaded eager</b><p>延迟不代表生产 kernel / serving</p></article>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="behavior-panel" data-behavior-panel="boundary" hidden>
|
||||||
|
<div class="panel-lead">
|
||||||
|
<div><span>EVIDENCE LADDER</span><h4>同一个问题,至少要跨过三层证据</h4></div>
|
||||||
|
<p>上一轮能看到专家路径,这一轮能看到生成文本;两者仍不能自动给出基准能力与机制因果。</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="evidence-ladder">
|
||||||
|
<article>
|
||||||
|
<span>01 / BASE ROUTING</span><b>11,289,744 routes</b>
|
||||||
|
<p>回答“历史协议改变后,前六个 MoE 层的目标 token 路由怎样变”。</p>
|
||||||
|
<em>不能回答最终生成了什么</em>
|
||||||
|
</article>
|
||||||
|
<i>→</i>
|
||||||
|
<article class="active">
|
||||||
|
<span>02 / CHAT GENERATION</span><b>128 greedy outputs</b>
|
||||||
|
<p>回答“固定 SFT Chat checkpoint 下,八格输出是否 exact、在哪里分叉”。</p>
|
||||||
|
<em>不能回答普遍能力或采样分布</em>
|
||||||
|
</article>
|
||||||
|
<i>→</i>
|
||||||
|
<article>
|
||||||
|
<span>03 / TASK EVALUATION</span><b>not yet identified</b>
|
||||||
|
<p>需要完整结束、足够样本、可执行 evaluator、采样复跑与预注册统计。</p>
|
||||||
|
<em>97 格截断,所以这一层未过闸</em>
|
||||||
|
</article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="token-contract">
|
||||||
|
<article class="official"><span>OFFICIAL</span><b>EOS · 100001</b><p>官方合法 Chat 序列;同时是 PAD alias。</p></article>
|
||||||
|
<article><span>COUNTERFACTUAL</span><b>BOS · 100000</b><p>另一个 special ID;不是合法历史结束。</p></article>
|
||||||
|
<article><span>COUNTERFACTUAL</span><b>x · 87</b><p>普通单 token 内容控制。</p></article>
|
||||||
|
<article><span>COUNTERFACTUAL</span><b>. · 13</b><p>普通单 token 标点控制。</p></article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="repro-grid">
|
||||||
|
<article class="pass"><span>MODEL FILES</span><b>12 / 12 PINNED</b><p>revision metadata + SHA-256</p></article>
|
||||||
|
<article class="pass"><span>LONG RERUN</span><b>32 / 32 EXACT</b><p>prompt、IDs、text、EOS 全一致</p></article>
|
||||||
|
<article class="pass"><span>SMOKE PREFIX</span><b>32 / 32 EXACT</b><p>32-token smoke = 正式前缀</p></article>
|
||||||
|
<article class="limit"><span>COMPLETION</span><b>31 / 128 EOS</b><p>截断率是结果的一部分</p></article>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="forbidden-claims">
|
||||||
|
<b>这一轮仍然不能写</b>
|
||||||
|
<p>“EOS 让答案更好” · “ordinary token 更差” · “route TV 解释了文本差异” · “4 条/域代表 benchmark” · “greedy exact 等于采样稳定” · “CPU offload 延迟等于生产吞吐”。</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="artifact-links">
|
||||||
|
<a href="https://huggingface.co/deepseek-ai/DeepSeek-V2-Lite-Chat" rel="noreferrer">官方模型与 model card ↗</a>
|
||||||
|
<a href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/src/data/deepseek-v2-lite-chat-behavior.json" rel="noreferrer">542KB 正式原始 JSON ↗</a>
|
||||||
|
<a href="https://git.k1412.top/wuyang/llm-atlas/src/branch/main/research/DEEPSEEK_V2_LITE_CHAT_BEHAVIOR_AUDIT.md" rel="noreferrer">完整审计与非结论 ↗</a>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<figcaption>
|
||||||
|
<span>X / OFFICIAL BF16 CHAT · DESCRIPTIVE BEHAVIOR PROBE</span>
|
||||||
|
固定 revision <code>{behavior.model.revision}</code>;正式原始 JSON
|
||||||
|
<code>{behavior.source.formalSha256.slice(0, 16)}…</code>。32 格长序列复跑逐 token exact;
|
||||||
|
其余 96 格未做完整 128-token 独立复跑。
|
||||||
|
</figcaption>
|
||||||
|
|
||||||
|
<script is:inline type="application/json" data-behavior-data set:html={json}></script>
|
||||||
|
</figure>
|
||||||
|
|
||||||
|
<script>
|
||||||
|
const roots = document.querySelectorAll<HTMLElement>("[data-behavior-lab]");
|
||||||
|
const labels: Record<string, string> = {
|
||||||
|
s0_eos: "S0 · EOS",
|
||||||
|
s1_eos: "S1 · EOS",
|
||||||
|
s0_bos: "S0 · BOS",
|
||||||
|
s1_bos: "S1 · BOS",
|
||||||
|
s0_x: "S0 · x",
|
||||||
|
s1_x: "S1 · x",
|
||||||
|
s0_period: "S0 · 句点",
|
||||||
|
s1_period: "S1 · 句点",
|
||||||
|
};
|
||||||
|
const domainNotes: Record<string, string> = {
|
||||||
|
all: "四域合计;不同域的输出长度与完成率不等,均值只作导航。",
|
||||||
|
english: "英文百科 · 4 条 source;不是开放域 QA benchmark。",
|
||||||
|
chinese: "中文新闻 · 4 条 source;不是 TNEWS 分类准确率。",
|
||||||
|
code: "HumanEval prompt · 4 条;代码只做 AST parse,未执行。",
|
||||||
|
math: "GSM8K question · 4 条;大量输出截断,不报告准确率。",
|
||||||
|
};
|
||||||
|
|
||||||
|
roots.forEach((root) => {
|
||||||
|
const payload = root.querySelector<HTMLScriptElement>("[data-behavior-data]");
|
||||||
|
if (!payload) return;
|
||||||
|
const data = JSON.parse(payload.textContent ?? "{}");
|
||||||
|
const one = <T extends Element>(selector: string) => root.querySelector<T>(selector);
|
||||||
|
const all = <T extends Element>(selector: string) => [...root.querySelectorAll<T>(selector)];
|
||||||
|
const set = (selector: string, value: string) => {
|
||||||
|
const node = one<HTMLElement>(selector);
|
||||||
|
if (node) node.textContent = value;
|
||||||
|
};
|
||||||
|
const sourceSelect = one<HTMLSelectElement>("[data-behavior-source]");
|
||||||
|
const edgeSelect = one<HTMLSelectElement>("[data-behavior-edge]");
|
||||||
|
const mapDomain = one<HTMLSelectElement>("[data-behavior-map-domain]");
|
||||||
|
const sourceById = new Map<string, any>(
|
||||||
|
data.sources.map((source: any) => [source.id, source]),
|
||||||
|
);
|
||||||
|
const shortHash = (value: string) => `${value.slice(0, 12)}…`;
|
||||||
|
|
||||||
|
const renderOutput = (side: "left" | "right", output: any) => {
|
||||||
|
set(`[data-output-condition="${side}"]`, labels[output.condition] ?? output.condition);
|
||||||
|
set(
|
||||||
|
`[data-output-status="${side}"]`,
|
||||||
|
output.hitEos ? "EOS COMPLETE" : "MAX 128 · TRUNCATED",
|
||||||
|
);
|
||||||
|
const status = one<HTMLElement>(`[data-output-status="${side}"]`);
|
||||||
|
status?.classList.toggle("complete", output.hitEos);
|
||||||
|
status?.classList.toggle("truncated", !output.hitEos);
|
||||||
|
set(
|
||||||
|
`[data-output-prompt="${side}"]`,
|
||||||
|
`${output.promptTokens} tokens · left pad ${output.leftPaddingTokens}`,
|
||||||
|
);
|
||||||
|
set(
|
||||||
|
`[data-output-tokens="${side}"]`,
|
||||||
|
`${output.generatedTokens} tokens`,
|
||||||
|
);
|
||||||
|
set(`[data-output-hash="${side}"]`, shortHash(output.generatedTokenIdsSha256));
|
||||||
|
set(`[data-output-text="${side}"]`, output.text || "(empty decoded text)");
|
||||||
|
};
|
||||||
|
|
||||||
|
const renderPair = () => {
|
||||||
|
const sourceId = sourceSelect?.value;
|
||||||
|
if (!sourceId) return;
|
||||||
|
const source = sourceById.get(sourceId);
|
||||||
|
const edge = edgeSelect?.value;
|
||||||
|
if (!source || !edge) return;
|
||||||
|
const pair = data.summary.pairwise[edge].find(
|
||||||
|
(row: any) => row.source_id === source.id,
|
||||||
|
);
|
||||||
|
if (!pair) return;
|
||||||
|
const outputByCondition = new Map(
|
||||||
|
source.outputs.map((output: any) => [output.condition, output]),
|
||||||
|
);
|
||||||
|
const left = outputByCondition.get(pair.left);
|
||||||
|
const right = outputByCondition.get(pair.right);
|
||||||
|
set("[data-behavior-source-id]", source.id);
|
||||||
|
set("[data-behavior-domain]", source.label);
|
||||||
|
set(
|
||||||
|
"[data-behavior-source-shape]",
|
||||||
|
`${source.sourceCharacters.toLocaleString()} chars · ${source.sourceTokens.toLocaleString()} tokens`,
|
||||||
|
);
|
||||||
|
set("[data-behavior-source-hash]", shortHash(source.sourceTextSha256));
|
||||||
|
set("[data-behavior-exact]", pair.token_ids_exact ? "EXACT" : "DIVERGED");
|
||||||
|
one("[data-behavior-exact]")?.classList.toggle("exact", pair.token_ids_exact);
|
||||||
|
set("[data-behavior-prefix]", pair.common_prefix_tokens.toLocaleString());
|
||||||
|
set("[data-behavior-edit]", pair.token_edit_distance.toLocaleString());
|
||||||
|
set(
|
||||||
|
"[data-behavior-similarity]",
|
||||||
|
`${(pair.normalized_token_similarity * 100).toFixed(1)}%`,
|
||||||
|
);
|
||||||
|
renderOutput("left", left);
|
||||||
|
renderOutput("right", right);
|
||||||
|
set(
|
||||||
|
"[data-behavior-reading]",
|
||||||
|
pair.token_ids_exact
|
||||||
|
? `这条边在当前 source 上没有改变完整 greedy token 序列;这不是“因素无效”,只是一条 exact 观测。`
|
||||||
|
: `两格在共同前缀 ${pair.common_prefix_tokens} token 后发生分叉,编辑距离 ${pair.token_edit_distance};它证明这条固定 greedy 轨迹改变,不证明哪一格更正确。`,
|
||||||
|
);
|
||||||
|
};
|
||||||
|
|
||||||
|
const renderMap = () => {
|
||||||
|
const domain = mapDomain?.value ?? "all";
|
||||||
|
const rows = domain === "all"
|
||||||
|
? data.summary.aggregates
|
||||||
|
: data.summary.aggregatesByDomain[domain];
|
||||||
|
set("[data-behavior-map-note]", domainNotes[domain]);
|
||||||
|
all<HTMLElement>("[data-behavior-map-edge]").forEach((row) => {
|
||||||
|
const edge = row.dataset.behaviorMapEdge;
|
||||||
|
if (!edge) return;
|
||||||
|
const metric = rows[edge];
|
||||||
|
const percent = metric.mean_normalized_token_similarity * 100;
|
||||||
|
const exact = row.querySelector<HTMLElement>("[data-map-exact]");
|
||||||
|
const meter = row.querySelector<HTMLElement>("[data-map-meter]");
|
||||||
|
const similarity = row.querySelector<HTMLElement>("[data-map-similarity]");
|
||||||
|
const edit = row.querySelector<HTMLElement>("[data-map-edit]");
|
||||||
|
if (exact) exact.textContent = `${metric.token_ids_exact} / ${metric.sources}`;
|
||||||
|
if (meter) meter.style.width = `${Math.max(1.5, percent)}%`;
|
||||||
|
if (similarity) similarity.textContent = `${percent.toFixed(1)}%`;
|
||||||
|
if (edit) edit.textContent = `edit ${metric.mean_token_edit_distance.toFixed(1)}`;
|
||||||
|
});
|
||||||
|
};
|
||||||
|
|
||||||
|
sourceSelect?.addEventListener("change", renderPair);
|
||||||
|
edgeSelect?.addEventListener("change", renderPair);
|
||||||
|
mapDomain?.addEventListener("change", renderMap);
|
||||||
|
all<HTMLButtonElement>("[data-behavior-map-edge]").forEach((button) => {
|
||||||
|
button.addEventListener("click", () => {
|
||||||
|
if (edgeSelect) edgeSelect.value = button.dataset.behaviorMapEdge ?? "system_eos";
|
||||||
|
one<HTMLButtonElement>('[data-behavior-tab="pair"]')?.click();
|
||||||
|
renderPair();
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
const tabs = all<HTMLButtonElement>("[data-behavior-tab]");
|
||||||
|
const panels = all<HTMLElement>("[data-behavior-panel]");
|
||||||
|
const selectTab = (tab: HTMLButtonElement) => {
|
||||||
|
tabs.forEach((candidate) => {
|
||||||
|
const active = candidate === tab;
|
||||||
|
candidate.setAttribute("aria-selected", String(active));
|
||||||
|
candidate.tabIndex = active ? 0 : -1;
|
||||||
|
});
|
||||||
|
panels.forEach((panel) => {
|
||||||
|
panel.hidden = panel.dataset.behaviorPanel !== tab.dataset.behaviorTab;
|
||||||
|
});
|
||||||
|
};
|
||||||
|
tabs.forEach((tab, index) => {
|
||||||
|
tab.addEventListener("click", () => selectTab(tab));
|
||||||
|
tab.addEventListener("keydown", (event) => {
|
||||||
|
if (!["ArrowLeft", "ArrowRight", "Home", "End"].includes(event.key)) return;
|
||||||
|
event.preventDefault();
|
||||||
|
let next = index;
|
||||||
|
if (event.key === "ArrowRight") next = (index + 1) % tabs.length;
|
||||||
|
if (event.key === "ArrowLeft") next = (index - 1 + tabs.length) % tabs.length;
|
||||||
|
if (event.key === "Home") next = 0;
|
||||||
|
if (event.key === "End") next = tabs.length - 1;
|
||||||
|
tabs[next].focus();
|
||||||
|
selectTab(tabs[next]);
|
||||||
|
});
|
||||||
|
});
|
||||||
|
renderPair();
|
||||||
|
renderMap();
|
||||||
|
});
|
||||||
|
</script>
|
||||||
|
|
||||||
|
<style>
|
||||||
|
.behavior-lab {
|
||||||
|
margin: 2.2rem 0 0;
|
||||||
|
overflow: hidden;
|
||||||
|
border: 1px solid rgba(30, 38, 43, .16);
|
||||||
|
background: #f7f4ec;
|
||||||
|
box-shadow: 0 28px 70px rgba(22, 35, 43, .1);
|
||||||
|
}
|
||||||
|
.behavior-head {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: 1fr 1fr;
|
||||||
|
gap: 2.4rem;
|
||||||
|
align-items: end;
|
||||||
|
padding: 2rem;
|
||||||
|
color: #f8f4e9;
|
||||||
|
background:
|
||||||
|
radial-gradient(circle at 82% 22%, rgba(85, 169, 159, .24), transparent 28%),
|
||||||
|
linear-gradient(135deg, #162a32, #254751);
|
||||||
|
}
|
||||||
|
.behavior-head p { margin: 0; color: rgba(255,255,255,.72); font-size: .77rem; line-height: 1.7; }
|
||||||
|
.behavior-head > div > p { color: #82c3b9; font: 750 .61rem/1.2 var(--font-mono); letter-spacing: .1em; }
|
||||||
|
.behavior-head h3 { max-width: 620px; margin: .75rem 0 0; color: white; font: 760 clamp(1.55rem, 3vw, 2.45rem)/1.14 var(--font-display); }
|
||||||
|
.behavior-head code { color: white; font-size: .7rem; }
|
||||||
|
.behavior-ledger { display: grid; grid-template-columns: repeat(5, 1fr); border-bottom: 1px solid rgba(30,38,43,.14); }
|
||||||
|
.behavior-ledger article { min-width: 0; padding: 1rem 1.15rem; border-right: 1px solid rgba(30,38,43,.12); background: #ece8dd; }
|
||||||
|
.behavior-ledger article:last-child { border-right: 0; }
|
||||||
|
.behavior-ledger article.complete { background: rgba(57, 126, 109, .12); }
|
||||||
|
.behavior-ledger article.warning { background: rgba(190, 100, 51, .14); }
|
||||||
|
.behavior-ledger span,
|
||||||
|
.checkpoint-ledger span,
|
||||||
|
.runtime-grid span,
|
||||||
|
.map-callout span,
|
||||||
|
.repro-grid span,
|
||||||
|
.token-contract span { display: block; color: #60706e; font: 720 .54rem/1.2 var(--font-mono); letter-spacing: .08em; }
|
||||||
|
.behavior-ledger b { display: block; margin-top: .45rem; color: #182b33; font: 760 .82rem/1.25 var(--font-mono); }
|
||||||
|
.behavior-ledger p { margin: .35rem 0 0; color: #68716f; font-size: .62rem; line-height: 1.4; }
|
||||||
|
.behavior-tabs { display: grid; grid-template-columns: repeat(4, 1fr); border-bottom: 1px solid rgba(30,38,43,.14); }
|
||||||
|
.behavior-tabs button { display: grid; grid-template-columns: auto 1fr; grid-template-rows: auto auto; column-gap: .75rem; min-width: 0; padding: .9rem 1rem; border: 0; border-right: 1px solid rgba(30,38,43,.14); color: #25363c; text-align: left; background: #fbf8f1; cursor: pointer; }
|
||||||
|
.behavior-tabs button:last-child { border-right: 0; }
|
||||||
|
.behavior-tabs button[aria-selected="true"] { color: white; background: #b25d36; }
|
||||||
|
.behavior-tabs span { grid-row: 1 / 3; opacity: .72; font: 720 .55rem/1.2 var(--font-mono); }
|
||||||
|
.behavior-tabs b { min-width: 0; font: 720 .76rem/1.25 var(--font-display); }
|
||||||
|
.behavior-tabs small { opacity: .68; font: .54rem/1.3 var(--font-mono); overflow-wrap: anywhere; }
|
||||||
|
.behavior-panel { padding: 1.55rem; }
|
||||||
|
.behavior-panel[hidden] { display: none; }
|
||||||
|
.panel-lead { display: grid; grid-template-columns: 1fr 1fr; gap: 2rem; align-items: end; margin-bottom: 1.25rem; }
|
||||||
|
.panel-lead span { color: #a75231; font: 750 .56rem/1.2 var(--font-mono); letter-spacing: .09em; }
|
||||||
|
.panel-lead h4 { margin: .35rem 0 0; color: #1c3037; font: 750 1.25rem/1.2 var(--font-display); }
|
||||||
|
.panel-lead p { margin: 0; color: #65716f; font-size: .7rem; line-height: 1.65; }
|
||||||
|
.pair-controls { display: grid; grid-template-columns: 1.2fr 1fr; gap: 1px; background: rgba(30,38,43,.14); border: 1px solid rgba(30,38,43,.14); }
|
||||||
|
.pair-controls label,
|
||||||
|
.map-control label { min-width: 0; padding: .8rem; background: #ede9df; }
|
||||||
|
.pair-controls label > span,
|
||||||
|
.map-control label > span { display: block; margin-bottom: .4rem; color: #65716f; font: 720 .54rem/1.2 var(--font-mono); letter-spacing: .08em; }
|
||||||
|
.pair-controls select,
|
||||||
|
.map-control select { width: 100%; min-width: 0; border: 1px solid rgba(30,38,43,.2); padding: .58rem; color: #1f3339; background: #fffdf8; font: 650 .66rem/1.3 var(--font-mono); }
|
||||||
|
.source-contract { display: grid; grid-template-columns: 1.25fr .8fr 1fr 1fr; margin-top: 1px; background: rgba(30,38,43,.12); gap: 1px; }
|
||||||
|
.source-contract > div { min-width: 0; padding: .7rem .8rem; background: #faf7f0; }
|
||||||
|
.source-contract span { display: block; color: #78817e; font: 680 .5rem/1.2 var(--font-mono); }
|
||||||
|
.source-contract b,
|
||||||
|
.source-contract code { display: block; margin-top: .3rem; color: #23343a; font-size: .62rem; overflow-wrap: anywhere; }
|
||||||
|
.pair-metrics { display: grid; grid-template-columns: repeat(4, 1fr); gap: 1px; margin-top: 1.15rem; background: rgba(30,38,43,.14); border: 1px solid rgba(30,38,43,.14); }
|
||||||
|
.pair-metrics article { padding: .8rem; background: #e8e4da; }
|
||||||
|
.pair-metrics span { display: block; color: #65716f; font: 680 .52rem/1.2 var(--font-mono); }
|
||||||
|
.pair-metrics b { display: inline-block; margin-top: .4rem; color: #b05130; font: 790 1.15rem/1 var(--font-display); }
|
||||||
|
.pair-metrics b.exact { color: #277563; }
|
||||||
|
.pair-metrics small { margin-left: .3rem; color: #777; font-size: .55rem; }
|
||||||
|
.output-pair { display: grid; grid-template-columns: 1fr 1fr; gap: 1px; margin-top: 1px; background: rgba(30,38,43,.15); border: 1px solid rgba(30,38,43,.15); }
|
||||||
|
.output-pair > article { min-width: 0; background: #fffdf8; }
|
||||||
|
.output-pair header { display: flex; justify-content: space-between; gap: 1rem; align-items: center; padding: .8rem 1rem; border-bottom: 1px solid rgba(30,38,43,.12); background: #efebe1; }
|
||||||
|
.output-pair header span { display: block; color: #78817e; font: 690 .49rem/1 var(--font-mono); }
|
||||||
|
.output-pair header b { display: block; margin-top: .25rem; color: #21363d; font: 760 .78rem/1.1 var(--font-mono); }
|
||||||
|
.output-pair header em { padding: .35rem .5rem; color: #a15031; background: rgba(177,87,49,.1); font: 720 .5rem/1 var(--font-mono); }
|
||||||
|
.output-pair header em.complete { color: #236a58; background: rgba(45,126,101,.12); }
|
||||||
|
.output-pair dl { display: grid; grid-template-columns: 1fr 1fr 1.15fr; margin: 0; border-bottom: 1px solid rgba(30,38,43,.1); }
|
||||||
|
.output-pair dl div { min-width: 0; padding: .6rem .7rem; border-right: 1px solid rgba(30,38,43,.09); }
|
||||||
|
.output-pair dl div:last-child { border-right: 0; }
|
||||||
|
.output-pair dt { color: #818884; font: 680 .46rem/1 var(--font-mono); }
|
||||||
|
.output-pair dd { margin: .28rem 0 0; color: #35464b; font-size: .56rem; overflow-wrap: anywhere; }
|
||||||
|
.output-pair > article > p { min-height: 15rem; max-height: 23rem; margin: 0; padding: 1rem; overflow: auto; color: #27383d; white-space: pre-wrap; font-size: .72rem; line-height: 1.65; }
|
||||||
|
.pair-reading,
|
||||||
|
.forbidden-claims { display: grid; grid-template-columns: 9rem 1fr; gap: 1rem; margin-top: 1rem; padding: .9rem 1rem; color: white; background: #253e46; }
|
||||||
|
.pair-reading span { color: #79c0b3; font: 750 .57rem/1.3 var(--font-mono); }
|
||||||
|
.pair-reading p,
|
||||||
|
.forbidden-claims p { margin: 0; color: rgba(255,255,255,.78); font-size: .68rem; line-height: 1.55; }
|
||||||
|
.map-control { display: grid; grid-template-columns: minmax(15rem, .7fr) 1.3fr; align-items: stretch; border: 1px solid rgba(30,38,43,.14); background: #ede9df; }
|
||||||
|
.map-control p { display: flex; align-items: center; margin: 0; padding: .8rem 1rem; color: #68736f; font-size: .66rem; line-height: 1.5; background: #faf7ef; }
|
||||||
|
.edge-map { margin-top: 1rem; border: 1px solid rgba(30,38,43,.15); }
|
||||||
|
.edge-map > header,
|
||||||
|
.edge-map > button { display: grid; grid-template-columns: 1.45fr .45fr 1fr .4fr; gap: .8rem; align-items: center; width: 100%; min-width: 0; padding: .65rem .8rem; border: 0; border-bottom: 1px solid rgba(30,38,43,.1); text-align: left; }
|
||||||
|
.edge-map > header { color: #dee9e6; background: #253e46; font: 690 .48rem/1.2 var(--font-mono); }
|
||||||
|
.edge-map > button { color: #2b3d42; background: #fbf8f1; cursor: pointer; }
|
||||||
|
.edge-map > button:hover { background: #f0ebe0; }
|
||||||
|
.edge-map > button:last-child { border-bottom: 0; }
|
||||||
|
.edge-map button > span { font: 670 .63rem/1.25 var(--font-mono); }
|
||||||
|
.edge-map button > b,
|
||||||
|
.edge-map button > strong,
|
||||||
|
.edge-map button > small { font: 720 .59rem/1 var(--font-mono); }
|
||||||
|
.edge-map button > i { height: .48rem; overflow: hidden; background: #ded9cd; }
|
||||||
|
.edge-map button > i > em { display: block; height: 100%; background: linear-gradient(90deg, #b45e37, #3d8275); }
|
||||||
|
.map-callout,
|
||||||
|
.checkpoint-ledger,
|
||||||
|
.runtime-grid,
|
||||||
|
.repro-grid,
|
||||||
|
.token-contract { display: grid; grid-template-columns: repeat(3, 1fr); gap: 1px; margin-top: 1rem; background: rgba(30,38,43,.13); border: 1px solid rgba(30,38,43,.13); }
|
||||||
|
.map-callout article,
|
||||||
|
.checkpoint-ledger article,
|
||||||
|
.runtime-grid article,
|
||||||
|
.repro-grid article,
|
||||||
|
.token-contract article { min-width: 0; padding: .9rem; background: #eeeae0; }
|
||||||
|
.map-callout b,
|
||||||
|
.checkpoint-ledger b,
|
||||||
|
.runtime-grid b,
|
||||||
|
.repro-grid b,
|
||||||
|
.token-contract b,
|
||||||
|
.checkpoint-ledger code { display: block; margin-top: .45rem; color: #23363c; font: 750 .73rem/1.25 var(--font-mono); overflow-wrap: anywhere; }
|
||||||
|
.map-callout p,
|
||||||
|
.checkpoint-ledger p,
|
||||||
|
.runtime-grid p,
|
||||||
|
.repro-grid p,
|
||||||
|
.token-contract p { margin: .45rem 0 0; color: #68736f; font-size: .61rem; line-height: 1.5; }
|
||||||
|
.offload-flow { display: grid; grid-template-columns: 1fr auto 1fr; gap: 1rem; align-items: center; margin-top: 1rem; }
|
||||||
|
.offload-flow > article { padding: 1rem; border: 1px solid rgba(30,38,43,.14); }
|
||||||
|
.offload-flow .gpu { color: white; background: #275d59; }
|
||||||
|
.offload-flow .cpu { color: white; background: #704a3c; }
|
||||||
|
.offload-flow header { display: flex; justify-content: space-between; gap: 1rem; font: 690 .61rem/1.2 var(--font-mono); }
|
||||||
|
.offload-flow p { margin: .7rem 0 0; color: rgba(255,255,255,.75); font-size: .63rem; line-height: 1.5; }
|
||||||
|
.memory-track { height: .65rem; margin-top: .8rem; background: rgba(255,255,255,.18); }
|
||||||
|
.memory-track i { display: block; width: var(--fill); height: 100%; background: #d7a16e; }
|
||||||
|
.offload-arrow { text-align: center; color: #6b716e; }
|
||||||
|
.offload-arrow span,
|
||||||
|
.offload-arrow small { display: block; font: .5rem/1.2 var(--font-mono); }
|
||||||
|
.offload-arrow b { display: block; color: #b35b36; font-size: 1.6rem; }
|
||||||
|
.layer-device-map { display: grid; grid-template-columns: repeat(15, 1fr); gap: 2px; margin-top: 1rem; }
|
||||||
|
.layer-device-map span { display: grid; place-items: center; min-height: 3rem; padding: .3rem .1rem; color: white; background: #34786e; }
|
||||||
|
.layer-device-map span.cpu { background: #875744; }
|
||||||
|
.layer-device-map b { font: 720 .55rem/1 var(--font-mono); }
|
||||||
|
.layer-device-map small { margin-top: .2rem; opacity: .7; font: .43rem/1 var(--font-mono); }
|
||||||
|
.runtime-grid { grid-template-columns: repeat(4, 1fr); }
|
||||||
|
.runtime-grid article.warning { background: rgba(177,86,48,.14); }
|
||||||
|
.evidence-ladder { display: grid; grid-template-columns: 1fr auto 1fr auto 1fr; gap: .7rem; align-items: center; }
|
||||||
|
.evidence-ladder article { min-height: 10rem; padding: 1rem; border: 1px solid rgba(30,38,43,.15); background: #eeeae0; }
|
||||||
|
.evidence-ladder article.active { color: white; background: #24464d; }
|
||||||
|
.evidence-ladder > i { color: #b15b37; font-size: 1.5rem; }
|
||||||
|
.evidence-ladder span { color: #a35434; font: 730 .54rem/1.2 var(--font-mono); }
|
||||||
|
.evidence-ladder article.active span { color: #7cc2b6; }
|
||||||
|
.evidence-ladder b { display: block; margin-top: .5rem; font: 760 .82rem/1.2 var(--font-mono); }
|
||||||
|
.evidence-ladder p { margin: .7rem 0; color: #626e6b; font-size: .65rem; line-height: 1.55; }
|
||||||
|
.evidence-ladder article.active p { color: rgba(255,255,255,.76); }
|
||||||
|
.evidence-ladder em { color: #a85534; font: 680 .57rem/1.4 var(--font-mono); }
|
||||||
|
.evidence-ladder article.active em { color: #e2aa7c; }
|
||||||
|
.token-contract { grid-template-columns: repeat(4, 1fr); }
|
||||||
|
.token-contract article.official { background: rgba(53,126,106,.13); }
|
||||||
|
.repro-grid { grid-template-columns: repeat(4, 1fr); }
|
||||||
|
.repro-grid article.pass { background: rgba(53,126,106,.12); }
|
||||||
|
.repro-grid article.limit { background: rgba(180,91,52,.14); }
|
||||||
|
.forbidden-claims { background: #402e2b; }
|
||||||
|
.forbidden-claims b { color: #e0a178; font: 740 .61rem/1.3 var(--font-mono); }
|
||||||
|
.artifact-links { display: flex; flex-wrap: wrap; gap: .5rem; margin-top: 1rem; }
|
||||||
|
.artifact-links a { padding: .55rem .7rem; color: #2d625b; border: 1px solid rgba(45,98,91,.25); background: rgba(45,98,91,.06); font: 680 .57rem/1.2 var(--font-mono); }
|
||||||
|
.behavior-lab figcaption { padding: .9rem 1.55rem; color: #75807c; border-top: 1px solid rgba(30,38,43,.13); background: #e8e4d9; font-size: .59rem; line-height: 1.5; }
|
||||||
|
.behavior-lab figcaption span { color: #a65031; font-weight: 750; }
|
||||||
|
.behavior-lab figcaption code { font-size: .55rem; overflow-wrap: anywhere; }
|
||||||
|
@media (max-width: 980px) {
|
||||||
|
.behavior-head,
|
||||||
|
.panel-lead { grid-template-columns: 1fr; }
|
||||||
|
.behavior-ledger { grid-template-columns: repeat(2, 1fr); }
|
||||||
|
.behavior-ledger article { border-bottom: 1px solid rgba(30,38,43,.12); }
|
||||||
|
.behavior-tabs { grid-template-columns: repeat(2, 1fr); }
|
||||||
|
.behavior-tabs button:nth-child(2) { border-right: 0; }
|
||||||
|
.behavior-tabs button:nth-child(-n+2) { border-bottom: 1px solid rgba(30,38,43,.14); }
|
||||||
|
.source-contract { grid-template-columns: repeat(2, 1fr); }
|
||||||
|
.output-pair { grid-template-columns: 1fr; }
|
||||||
|
.map-callout,
|
||||||
|
.checkpoint-ledger { grid-template-columns: 1fr; }
|
||||||
|
.layer-device-map { grid-template-columns: repeat(10, 1fr); }
|
||||||
|
.runtime-grid,
|
||||||
|
.token-contract,
|
||||||
|
.repro-grid { grid-template-columns: repeat(2, 1fr); }
|
||||||
|
.evidence-ladder { grid-template-columns: 1fr; }
|
||||||
|
.evidence-ladder > i { transform: rotate(90deg); text-align: center; }
|
||||||
|
}
|
||||||
|
@media (max-width: 620px) {
|
||||||
|
.behavior-head,
|
||||||
|
.behavior-panel { padding: 1rem; }
|
||||||
|
.behavior-ledger { grid-template-columns: 1fr; }
|
||||||
|
.behavior-ledger article { border-right: 0; }
|
||||||
|
.behavior-tabs { display: flex; overflow-x: auto; }
|
||||||
|
.behavior-tabs button { flex: 1 0 10.5rem; border-bottom: 0 !important; }
|
||||||
|
.pair-controls,
|
||||||
|
.source-contract,
|
||||||
|
.pair-metrics,
|
||||||
|
.map-control { grid-template-columns: 1fr; }
|
||||||
|
.pair-metrics { grid-template-columns: repeat(2, 1fr); }
|
||||||
|
.output-pair dl { grid-template-columns: 1fr; }
|
||||||
|
.output-pair dl div { border-right: 0; border-bottom: 1px solid rgba(30,38,43,.09); }
|
||||||
|
.output-pair > article > p { min-height: 10rem; max-height: 19rem; }
|
||||||
|
.pair-reading,
|
||||||
|
.forbidden-claims { grid-template-columns: 1fr; }
|
||||||
|
.edge-map { overflow-x: auto; }
|
||||||
|
.edge-map > header,
|
||||||
|
.edge-map > button { min-width: 35rem; }
|
||||||
|
.offload-flow { grid-template-columns: 1fr; }
|
||||||
|
.offload-arrow b { transform: rotate(90deg); }
|
||||||
|
.layer-device-map { grid-template-columns: repeat(5, 1fr); }
|
||||||
|
.runtime-grid,
|
||||||
|
.token-contract,
|
||||||
|
.repro-grid { grid-template-columns: 1fr; }
|
||||||
|
}
|
||||||
|
</style>
|
||||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -3,6 +3,7 @@ import BaseLayout from "@/layouts/BaseLayout.astro";
|
|||||||
import DeepSeekLineage from "@/components/DeepSeekLineage.astro";
|
import DeepSeekLineage from "@/components/DeepSeekLineage.astro";
|
||||||
import DeepSeekLab from "@/components/DeepSeekLab.astro";
|
import DeepSeekLab from "@/components/DeepSeekLab.astro";
|
||||||
import DeepSeekArtifactLab from "@/components/DeepSeekArtifactLab.astro";
|
import DeepSeekArtifactLab from "@/components/DeepSeekArtifactLab.astro";
|
||||||
|
import DeepSeekBehaviorLab from "@/components/DeepSeekBehaviorLab.astro";
|
||||||
import { deepseekBranches, deepseekLedgers, deepseekPaperChain, deepseekWaves } from "@/data/deepseek";
|
import { deepseekBranches, deepseekLedgers, deepseekPaperChain, deepseekWaves } from "@/data/deepseek";
|
||||||
|
|
||||||
const toc = [
|
const toc = [
|
||||||
@@ -29,21 +30,22 @@ const toc = [
|
|||||||
["20", "k3", "与 K3 的继承边界"],
|
["20", "k3", "与 K3 的继承边界"],
|
||||||
["21", "lab", "四联交互实验"],
|
["21", "lab", "四联交互实验"],
|
||||||
["22", "artifact", "真实权重执行"],
|
["22", "artifact", "真实权重执行"],
|
||||||
["23", "branches", "别漏掉旁支"],
|
["23", "behavior", "Chat:最终生成行为"],
|
||||||
["24", "audit", "事实、推导与教学模型"],
|
["24", "branches", "别漏掉旁支"],
|
||||||
|
["25", "audit", "事实、推导与教学模型"],
|
||||||
["↳", "papers", "六十节点阅读链"],
|
["↳", "papers", "六十节点阅读链"],
|
||||||
];
|
];
|
||||||
---
|
---
|
||||||
|
|
||||||
<BaseLayout
|
<BaseLayout
|
||||||
title="DeepSeek 技术谱系与真实权重深读:从 Dense、MoE、MLA 到 R1 与 V4"
|
title="DeepSeek 技术谱系与真实权重深读:从 Dense、MoE、MLA 到 R1 与 V4"
|
||||||
description="用二十四张问题账、十次技术转向、十七个交互实验、真实 V2-Lite 权重、公开语料路由区间、官方模板、消息历史、等长 filler、特殊词元家族与完整角色块控制、吸收式缓存 trace 和六十个一手节点,完整理解 DeepSeek 的 MoE、MLA、FP8、DualPipe、GRPO、R1、V3.2 与 V4。"
|
description="用二十四张问题账、十次技术转向、十八个交互实验、真实 V2-Lite Base / Chat 权重、公开语料路由区间、官方模板、消息历史、等长 filler、特殊词元家族、完整角色块与最终生成行为控制、吸收式缓存 trace 和六十个一手节点,完整理解 DeepSeek 的 MoE、MLA、FP8、DualPipe、GRPO、R1、V3.2 与 V4。"
|
||||||
section="deepseek"
|
section="deepseek"
|
||||||
>
|
>
|
||||||
<header class="page-hero deepseek-hero">
|
<header class="page-hero deepseek-hero">
|
||||||
<div class="page-hero-inner">
|
<div class="page-hero-inner">
|
||||||
<div>
|
<div>
|
||||||
<p class="eyebrow"><span>SPOTLIGHT / DEEPSEEK · ROUND 03</span> ALGORITHM × SYSTEM × REAL WEIGHTS</p>
|
<p class="eyebrow"><span>SPOTLIGHT / DEEPSEEK · ROUND 04</span> ROUTING × CHAT OUTPUT × REAL WEIGHTS</p>
|
||||||
<h1>不要背模型名<br />要看懂每次为什么转向</h1>
|
<h1>不要背模型名<br />要看懂每次为什么转向</h1>
|
||||||
<p class="lead">
|
<p class="lead">
|
||||||
这不是七篇报告的摘要,而是一套可追问、可计算、可反驳的技术谱系:
|
这不是七篇报告的摘要,而是一套可追问、可计算、可反驳的技术谱系:
|
||||||
@@ -55,7 +57,7 @@ const toc = [
|
|||||||
<div><dt>SPAN</dt><dd>2024.01 → 2026.06</dd></div>
|
<div><dt>SPAN</dt><dd>2024.01 → 2026.06</dd></div>
|
||||||
<div><dt>LEDGERS</dt><dd>24 张问题账</dd></div>
|
<div><dt>LEDGERS</dt><dd>24 张问题账</dd></div>
|
||||||
<div><dt>LINEAGE</dt><dd>10 次技术转向</dd></div>
|
<div><dt>LINEAGE</dt><dd>10 次技术转向</dd></div>
|
||||||
<div><dt>LABS</dt><dd>17 个可操作实验</dd></div>
|
<div><dt>LABS</dt><dd>18 个可操作实验</dd></div>
|
||||||
<div><dt>EVIDENCE</dt><dd>60 个一手 / 官方节点</dd></div>
|
<div><dt>EVIDENCE</dt><dd>60 个一手 / 官方节点</dd></div>
|
||||||
<div><dt>STATUS</dt><dd>三轮 · 真实权重执行</dd></div>
|
<div><dt>STATUS</dt><dd>三轮 · 真实权重执行</dd></div>
|
||||||
</dl>
|
</dl>
|
||||||
@@ -781,8 +783,20 @@ const toc = [
|
|||||||
<DeepSeekArtifactLab />
|
<DeepSeekArtifactLab />
|
||||||
</section>
|
</section>
|
||||||
|
|
||||||
|
<section class="article-section" id="behavior">
|
||||||
|
<p class="eyebrow"><span>23</span> ROUTING IS NOT THE ANSWER</p>
|
||||||
|
<h2>第七层之后不再只看 expert:加载完整 Chat 权重,实际生成 128 个输出</h2>
|
||||||
|
<p class="lede">
|
||||||
|
Base checkpoint 的路由实验回答“输入协议怎样改变专家路径”,却不能告诉我们模型最终说了什么。
|
||||||
|
这一轮固定官方 <code>DeepSeek-V2-Lite-Chat</code> revision 与 31,412,968,448 bytes BF16 参数,
|
||||||
|
把同一批 source 带进完整 27 层 generation。由于本机 32GB 显存低于官方 40GB 单卡边界,
|
||||||
|
layers 25–26、final norm 与 LM head 明确落到 CPU;offload 事实、截断率和复跑覆盖都直接展示。
|
||||||
|
</p>
|
||||||
|
<DeepSeekBehaviorLab />
|
||||||
|
</section>
|
||||||
|
|
||||||
<section class="article-section" id="branches">
|
<section class="article-section" id="branches">
|
||||||
<p class="eyebrow"><span>23</span> THE MAIN LINE IS NOT THE WHOLE TREE</p>
|
<p class="eyebrow"><span>24</span> THE MAIN LINE IS NOT THE WHOLE TREE</p>
|
||||||
<h2>如果只读 V2 → V3 → R1 → V4,会漏掉五条反过来影响主线的旁支</h2>
|
<h2>如果只读 V2 → V3 → R1 → V4,会漏掉五条反过来影响主线的旁支</h2>
|
||||||
<div class="branch-grid">
|
<div class="branch-grid">
|
||||||
{deepseekBranches.map(([name, line, text, url]) => (
|
{deepseekBranches.map(([name, line, text, url]) => (
|
||||||
@@ -802,7 +816,7 @@ const toc = [
|
|||||||
</section>
|
</section>
|
||||||
|
|
||||||
<section class="article-section" id="audit">
|
<section class="article-section" id="audit">
|
||||||
<p class="eyebrow"><span>24</span> EVIDENCE AUDIT</p>
|
<p class="eyebrow"><span>25</span> EVIDENCE AUDIT</p>
|
||||||
<h2>同一张页面里有三种知识,它们的语气必须不同</h2>
|
<h2>同一张页面里有三种知识,它们的语气必须不同</h2>
|
||||||
<div class="audit-grid">
|
<div class="audit-grid">
|
||||||
<article class="reported">
|
<article class="reported">
|
||||||
|
|||||||
@@ -145,17 +145,18 @@ const paths = [
|
|||||||
</a>
|
</a>
|
||||||
<a class="release-card deepseek-release" href="/deepseek/">
|
<a class="release-card deepseek-release" href="/deepseek/">
|
||||||
<div>
|
<div>
|
||||||
<p class="eyebrow"><span>NEW / DEEPSEEK ROUND 03</span> LINEAGE · REAL WEIGHTS · ROUTES · CACHE</p>
|
<p class="eyebrow"><span>NEW / DEEPSEEK ROUND 04</span> LINEAGE · REAL WEIGHTS · ROUTES · GENERATION</p>
|
||||||
<h2>从 Dense 到百万上下文:每次创新都在偿还上一代最贵的一张账</h2>
|
<h2>从 Dense 到百万上下文:每次创新都在偿还上一代最贵的一张账</h2>
|
||||||
<p>
|
<p>
|
||||||
用二十四张问题账和十次技术转向走完 Dense→V4,再固定官方 V2-Lite 权重执行 7/27 层:
|
用二十四张问题账和十次技术转向走完 Dense→V4,再把 Base 路由证据接到官方
|
||||||
逐 token 检查 3,240 次专家选择,并把 latent 状态与 HF eager cache 的实现差距摆在同一张账上。
|
V2-Lite-Chat 的 31.4 GB 完整 BF16 权重:16 个公开来源、8 种边界条件生成
|
||||||
|
128 个输出,并把 31 个自然 EOS 与 97 个长度截断分开解释。
|
||||||
</p>
|
</p>
|
||||||
</div>
|
</div>
|
||||||
<dl>
|
<dl>
|
||||||
<div><dt>LINEAGE</dt><dd>1991 → 2026 · 10 次转向</dd></div>
|
<div><dt>LINEAGE</dt><dd>1991 → 2026 · 10 次转向</dd></div>
|
||||||
<div><dt>NODES</dt><dd>60 个一手 / 官方节点</dd></div>
|
<div><dt>NODES</dt><dd>60 个一手 / 官方节点</dd></div>
|
||||||
<div><dt>LAB</dt><dd>4 公式实验 · 4 真实工件实验</dd></div>
|
<div><dt>LAB</dt><dd>4 公式 · 13 Base 工件 · 1 Chat 行为</dd></div>
|
||||||
</dl>
|
</dl>
|
||||||
<span class="release-arrow" aria-hidden="true">进入 DeepSeek 完整技术谱系 →</span>
|
<span class="release-arrow" aria-hidden="true">进入 DeepSeek 完整技术谱系 →</span>
|
||||||
</a>
|
</a>
|
||||||
|
|||||||
@@ -15,7 +15,7 @@ const workstreams = [
|
|||||||
{ label: "表示、位置与残差高速公路", value: 81, next: "加入真实 hidden-state / norm traces、长上下文位置外推复现与更多深层稳定性消融" },
|
{ label: "表示、位置与残差高速公路", value: 81, next: "加入真实 hidden-state / norm traces、长上下文位置外推复现与更多深层稳定性消融" },
|
||||||
{ label: "Scaling Laws", value: 74, next: "加入真实拟合复现、置信区间与更多模型族对照" },
|
{ label: "Scaling Laws", value: 74, next: "加入真实拟合复现、置信区间与更多模型族对照" },
|
||||||
{ label: "数据工程与预训练配方", value: 73, next: "逐图精读 FineWeb / DCLM,加入真实去重与 mixture traces" },
|
{ label: "数据工程与预训练配方", value: 73, next: "逐图精读 FineWeb / DCLM,加入真实去重与 mixture traces" },
|
||||||
{ label: "DeepSeek 专题", value: 98, next: "V2-Lite-Chat 生成/行为对照、完整 27 层与固定 batch content,再推进 SM90 FlashMLA、FP8/pipeline 与 R1-like RL" },
|
{ label: "DeepSeek 专题", value: 98, next: "completion-aware 生成评测、完整 27 层与固定 batch content,再推进 SM90 FlashMLA、FP8/pipeline 与 R1-like RL" },
|
||||||
{ label: "指令微调与人类偏好", value: 75, next: "加入真实偏好分歧样本、RM 长度偏置与 PPO/DPO 小模型复现" },
|
{ label: "指令微调与人类偏好", value: 75, next: "加入真实偏好分歧样本、RM 长度偏置与 PPO/DPO 小模型复现" },
|
||||||
{ label: "推理与测试时扩展", value: 76, next: "真实模型采样曲线、PRM 案例与逐篇图表精读" },
|
{ label: "推理与测试时扩展", value: 76, next: "真实模型采样曲线、PRM 案例与逐篇图表精读" },
|
||||||
{ label: "工具使用与长程 Agent", value: 74, next: "补真实环境 traces、cross-harness 对照、Agent RL 训练曲线与安全案例" },
|
{ label: "工具使用与长程 Agent", value: 74, next: "补真实环境 traces、cross-harness 对照、Agent RL 训练曲线与安全案例" },
|
||||||
@@ -97,12 +97,12 @@ const workstreams = [
|
|||||||
<article><span>✓</span><h3>K3 报告已结构化拆解</h3><p>47 页报告目录、151 条参考来源和架构/后训练/系统主线已经提取。</p></article>
|
<article><span>✓</span><h3>K3 报告已结构化拆解</h3><p>47 页报告目录、151 条参考来源和架构/后训练/系统主线已经提取。</p></article>
|
||||||
<article><span>✓</span><h3>17 专题知识图</h3><p>从语言模型基础到评测安全,包含先修依赖和三条贯穿案例。</p></article>
|
<article><span>✓</span><h3>17 专题知识图</h3><p>从语言模型基础到评测安全,包含先修依赖和三条贯穿案例。</p></article>
|
||||||
<article><span>✓</span><h3>编辑式网站系统</h3><p>响应式导航、章节模板、侧栏、进度、论文链和证据提示组件。</p></article>
|
<article><span>✓</span><h3>编辑式网站系统</h3><p>响应式导航、章节模板、侧栏、进度、论文链和证据提示组件。</p></article>
|
||||||
<article><span>✓</span><h3>八十四个原创交互视图</h3><p>K3 三轴图、八联报告实验与四联开放工件实验,DeepSeek 四联公式实验与十三联真实权重实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
|
<article><span>✓</span><h3>八十五个原创交互视图</h3><p>K3 三轴图、八联报告实验与四联开放工件实验,DeepSeek 四联公式实验、十三联 Base 工件实验与一联完整 Chat 行为实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
|
||||||
<article><span>✓</span><h3>十七篇首版长文</h3><p>K3、语言模型前史、Transformer、表示/位置/残差、DeepSeek、Scaling、数据工程、长上下文、MoE、后训练、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全专题。</p></article>
|
<article><span>✓</span><h3>十七篇首版长文</h3><p>K3、语言模型前史、Transformer、表示/位置/残差、DeepSeek、Scaling、数据工程、长上下文、MoE、后训练、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全专题。</p></article>
|
||||||
<article><span>✓</span><h3>语言模型前史深度专题</h3><p>八张独立问题账、33 个正式节点、20 段长文与概率—向量—记忆—对齐四联实验。</p></article>
|
<article><span>✓</span><h3>语言模型前史深度专题</h3><p>八张独立问题账、33 个正式节点、20 段长文与概率—向量—记忆—对齐四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>Transformer 深度专题</h3><p>十张独立问题账、40 个正式节点、21 段正文与 QKV—Mask—多头位置—Block 成本四联实验。</p></article>
|
<article><span>✓</span><h3>Transformer 深度专题</h3><p>十张独立问题账、40 个正式节点、21 段正文与 QKV—Mask—多头位置—Block 成本四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>表示、位置与残差高速公路深度专题</h3><p>二十张问题账、66 个一手节点、DeepSeek/Kimi 双谱系,以及 Token—位置—Norm—Residual/FFN 四联实验。</p></article>
|
<article><span>✓</span><h3>表示、位置与残差高速公路深度专题</h3><p>二十张问题账、66 个一手节点、DeepSeek/Kimi 双谱系,以及 Token—位置—Norm—Residual/FFN 四联实验。</p></article>
|
||||||
<article><span>✓</span><h3>DeepSeek 三轮真实权重里程碑</h3><p>在二十四张问题账、十次转向与四联公式实验上,新增 V2-Lite 7/27 层连续 forward、官方 V3 absorb,以及长度、模板、消息历史、边界、角色词头、完整 special inventory 与两-token 角色块控制;累计 11,289,744 次真实路由。最新两组八格各自 byte-exact:BOS 不复现 EOS,四 ID 的 2-vs-2 只作描述;head、delimiter 与 interaction 均无统一 system 调制方向。</p></article>
|
<article><span>✓</span><h3>DeepSeek 四轮真实权重里程碑</h3><p>在 Base 路由与缓存实证上,新增官方 V2-Lite-Chat 的 31.4 GB 完整 BF16 生成:16 个公开来源 × 8 条件得到 128 个输出,31 个自然 EOS、97 个长度截断;独立复跑的 32 / 32 token 序列 exact。逐来源双输出、十边分歧、GPU/CPU offload 与证据边界共同组成第十八个实验,不把生成差异越界写成能力。</p></article>
|
||||||
<article><span>✓</span><h3>Kimi K3 技术报告二轮深读</h3><p>三十二张问题账、Figure 1–16 / Table 1–5 审计、100 节点阅读链,以及 Delta—Decay—AttnRes—LatentMoE—SiTU—QB—MOPD—Cache 八联实验。</p></article>
|
<article><span>✓</span><h3>Kimi K3 技术报告二轮深读</h3><p>三十二张问题账、Figure 1–16 / Table 1–5 审计、100 节点阅读链,以及 Delta—Decay—AttnRes—LatentMoE—SiTU—QB—MOPD—Cache 八联实验。</p></article>
|
||||||
<article><span>✓</span><h3>Kimi K3 三轮开放工件里程碑</h3><p>固定官方 revisions,审计 96 个 shards、497,220 个 tensor entries 与真实 KDA / MLA / MoE / MoonViT shapes;四联实验分开显示层型、tensor anatomy、参数范围和复现边界。</p></article>
|
<article><span>✓</span><h3>Kimi K3 三轮开放工件里程碑</h3><p>固定官方 revisions,审计 96 个 shards、497,220 个 tensor entries 与真实 KDA / MLA / MoE / MoonViT shapes;四联实验分开显示层型、tensor anatomy、参数范围和复现边界。</p></article>
|
||||||
<article><span>✓</span><h3>FlashKDA RTX 5090 执行闸门</h3><p>隔离 CUDA 13.0 / glibc 2.39 编译 sm_120a wheel;6/6 官方参考逐元素相等,并完成 fixed / varlen、三种 state mode 的 1,800 个 CUDA Event samples。</p></article>
|
<article><span>✓</span><h3>FlashKDA RTX 5090 执行闸门</h3><p>隔离 CUDA 13.0 / glibc 2.39 编译 sm_120a wheel;6/6 官方参考逐元素相等,并完成 fixed / varlen、三种 state mode 的 1,800 个 CUDA Event samples。</p></article>
|
||||||
@@ -134,7 +134,7 @@ const workstreams = [
|
|||||||
<div class="queue-table">
|
<div class="queue-table">
|
||||||
<div class="head"><b>优先级</b><b>专题</b><b>本轮交付</b><b>完成闸门</b></div>
|
<div class="head"><b>优先级</b><b>专题</b><b>本轮交付</b><b>完成闸门</b></div>
|
||||||
<div><span>P0</span><strong>K3 三轮</strong><p>开放权重 traces → FlashKDA / AttnRes / MoE 真实行为 → Figure 1–16 数值重绘与独立复现</p><em>运行证据 + 逐图复现</em></div>
|
<div><span>P0</span><strong>K3 三轮</strong><p>开放权重 traces → FlashKDA / AttnRes / MoE 真实行为 → Figure 1–16 数值重绘与独立复现</p><em>运行证据 + 逐图复现</em></div>
|
||||||
<div><span>P0</span><strong>DeepSeek 三轮</strong><p>V2-Lite-Chat 生成/行为对照 → 完整 27 层与固定 batch content → SM90 FlashMLA / FP8 / pipeline traces → R1-like RL 小模型复现</p><em>运行证据 + 独立复现</em></div>
|
<div><span>P0</span><strong>DeepSeek 四轮</strong><p>completion-aware 行为评测 → 完整 27 层与固定 batch content → SM90 FlashMLA / FP8 / pipeline traces → R1-like RL 小模型复现</p><em>运行证据 + 独立复现</em></div>
|
||||||
<div><span>P0</span><strong>Transformer 二轮</strong><p>多头电路逐图 → Pre/Post-LN 真实 traces → Flash/KV 配置与 kernel 对照</p><em>逐图笔记 + 实测边界</em></div>
|
<div><span>P0</span><strong>Transformer 二轮</strong><p>多头电路逐图 → Pre/Post-LN 真实 traces → Flash/KV 配置与 kernel 对照</p><em>逐图笔记 + 实测边界</em></div>
|
||||||
<div><span>P0</span><strong>表示、位置与残差二轮</strong><p>真实 hidden-state / norm traces → 长上下文位置外推 → mHC / AttnRes 深层稳定性消融</p><em>可复现实验 + 逐图笔记</em></div>
|
<div><span>P0</span><strong>表示、位置与残差二轮</strong><p>真实 hidden-state / norm traces → 长上下文位置外推 → mHC / AttnRes 深层稳定性消融</p><em>可复现实验 + 逐图笔记</em></div>
|
||||||
<div><span>P0</span><strong>语言模型前史二轮</strong><p>Kneser–Ney / LSTM / Bahdanau 逐图 → 真实小语料复现 → tokenizer 公平性</p><em>可复现实验 + 逐图笔记</em></div>
|
<div><span>P0</span><strong>语言模型前史二轮</strong><p>Kneser–Ney / LSTM / Bahdanau 逐图 → 真实小语料复现 → tokenizer 公平性</p><em>可复现实验 + 逐图笔记</em></div>
|
||||||
@@ -226,6 +226,9 @@ const workstreams = [
|
|||||||
<div><time>2026-07-29</time><b>完整 special inventory 与普通对照分身份</b><p>BOS/EOS 穷尽固定 tokenizer 的两个 special IDs;x/句点只是两个选定普通对照,2-vs-2 只描述四个 ID。</p></div>
|
<div><time>2026-07-29</time><b>完整 special inventory 与普通对照分身份</b><p>BOS/EOS 穷尽固定 tokenizer 的两个 special IDs;x/句点只是两个选定普通对照,2-vs-2 只描述四个 ID。</p></div>
|
||||||
<div><time>2026-07-29</time><b>两-token 角色块按因子分解</b><p>`User:` / `Assistant:` 都是 head + delimiter 两个普通 IDs;head、delimiter、interaction 与直接边分别记账。</p></div>
|
<div><time>2026-07-29</time><b>两-token 角色块按因子分解</b><p>`User:` / `Assistant:` 都是 head + delimiter 两个普通 IDs;head、delimiter、interaction 与直接边分别记账。</p></div>
|
||||||
<div><time>2026-07-29</time><b>下游目标与完整输入分因果身份</b><p>目标内容只看编辑位置后的精确对齐路由;完整输入包含被编辑 token 自身,只作稳健性账。</p></div>
|
<div><time>2026-07-29</time><b>下游目标与完整输入分因果身份</b><p>目标内容只看编辑位置后的精确对齐路由;完整输入包含被编辑 token 自身,只作稳健性账。</p></div>
|
||||||
|
<div><time>2026-07-29</time><b>Chat、生成行为与能力永久分层</b><p>完整官方 Chat 权重可以支撑真实生成;成对输出分歧只证明干预传播,没有 evaluator 与足够 completion 就不升级成能力判断。</p></div>
|
||||||
|
<div><time>2026-07-29</time><b>completion 是生成实验的首要审计字段</b><p>31 / 128 自然 EOS 与 97 / 128 长度截断同时显示;截断答案不冒充完整回答。</p></div>
|
||||||
|
<div><time>2026-07-29</time><b>offload 拓扑进入复现合同</b><p>31.4 GB BF16 权重按 GPU 25 层、CPU 2 层 + norm / lm_head 执行;设备切分不被写成模型结构。</p></div>
|
||||||
</div>
|
</div>
|
||||||
</section>
|
</section>
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user