feat: add DeepSeek Chat behavior evidence

This commit is contained in:
wuyang
2026-07-29 22:47:03 +08:00
parent b615224798
commit 96443d4dfc
18 changed files with 35709 additions and 34 deletions
@@ -0,0 +1,434 @@
# DeepSeek-V2-Lite-Chat 最终生成行为审计
> 状态:Round 04 正式结果
>
> 正式原始 JSON:`src/data/deepseek-v2-lite-chat-behavior.json`
>
> 长序列复跑 JSON:`src/data/deepseek-v2-lite-chat-behavior-repro-1pd.json`
>
> 执行脚本:`experiments/deepseek/v2_lite_chat_special_token_behavior_probe.py`
>
> 预注册协议:`research/DEEPSEEK_V2_LITE_CHAT_BEHAVIOR_PROTOCOL.md`
## 0. 一句话结论
固定官方 SFT Chat checkpoint、greedy decoding 与同 source 八格 batch 后,一处历史边界
ID 或一条 system message 经常足以让最终 token 轨迹分叉;但本轮 128 个输出中只有 31 个
在 128-token 上限内遇到 EOS,因此它识别的是**小样本确定性输出敏感性**,不是任务能力、
采样分布稳定性,也不是“某种 token 更好”。
---
## 1. 为什么上一轮还不能回答这个问题
上一组 special-token family 实验加载的是 `DeepSeek-V2-Lite` Base checkpoint,只观察
前六个 MoE 层的 hidden states 与 top-6 expert routes。它能回答:
```text
同一个 pre-target input ID 被替换以后,
后续目标 token 的专家路径怎样改变?
```
但它不能回答:
```text
完整模型最后生成了什么?
两格输出是否逐 token 相同?
任务结果有没有改变?
```
本轮因此换到官方
[`deepseek-ai/DeepSeek-V2-Lite-Chat`](https://huggingface.co/deepseek-ai/DeepSeek-V2-Lite-Chat)
SFT checkpoint,并真正调用 `generate()`。DeepSeek 官方仓库把 Lite-Chat 明确列作
16B total / 2.4B active / 32K context 的 **SFT** 模型;它与较大的 V2-Chat (RL) 不是
同一个 checkpoint 身份。
来源:
[DeepSeek-V2 官方仓库](https://github.com/deepseek-ai/DeepSeek-V2)、
[DeepSeek-V2 报告](https://arxiv.org/abs/2405.04434)。
这也意味着:
- Base routing 与 Chat generation 是两种证据层;
- 不能把二者的 effect size 直接相减;
- 不能把 Chat 输出分叉倒推成某个 expert 的因果中介;
- 不能把 SFT checkpoint 写成 R1、RL 或当前线上 DeepSeek。
---
## 2. 固定了什么
### 2.1 模型工件
| 对象 | 固定值 |
|---|---:|
| 模型 | `deepseek-ai/DeepSeek-V2-Lite-Chat` |
| revision | `85864749cd611b4353ce1decdb286193298f64c7` |
| checkpoint 身份 | SFT Chat |
| dtype | BF16 |
| safetensors shards | 4 |
| index tensor bytes | `31,412,968,448` |
| shard file bytes(含 headers) | `31,413,626,576` |
| 文件数 | 12 |
| revision metadata | 12 / 12 同一 revision |
运行时参数 bytes 按 dtype 汇总也是 `31,412,968,448`,全部为 `torch.bfloat16`,与
index 的 `metadata.total_size` 精确相等。
正式原始结果:
```text
542,559 bytes
SHA-256 54496955d0dd20a3116e6b758a55f56e2444c43dd99bfef94a2cb7b140e6196e
```
长序列复跑:
```text
153,819 bytes
SHA-256 b74c31606c0fc70a8e52cc6c3d6135c4e916d698b4d5dae51863124d025b0c72
```
### 2.2 tokenizer
Chat 与 Base 的 `tokenizer.json`、`tokenizer_config.json` 逐字节相同:
```text
tokenizer.json
41f3bf64213da8c012d8bd0871a58a1fdf70463e8f08f110ddbb1082f529f669
tokenizer_config.json
31181eaf79394ea26728d95ecb54fe7c8413e6f56085dbabc8b0818134380ec8
```
固定边界 ID:
| level | ID | 类别 | 官方合法历史边界 |
|---|---:|---|---|
| EOS | `100001` | special;同时是 PAD alias | 是 |
| BOS | `100000` | special | 否 |
| `x` | `87` | ordinary content | 否 |
| `.` | `13` | ordinary punctuation | 否 |
EOS 与 BOS 穷尽这个固定 tokenizer 的 special-token inventory。BOS / `x` / 句点三格
都不是官方合法 Chat 序列,它们只是在 official rendering 后替换一个 pre-target
input ID 的反事实控制。
---
## 3. 八格设计
```text
system off/on × EOS / BOS / x / period
```
同一 system 水平内:
- 四格 token 数相同;
- target content、generation prompt 与 target position 相同;
- EOS 格相对官方序列改 0 个 ID;
- BOS / `x` / period 各恰好改 1 个 ID。
system on 比 system off 多 16 个 token。八格同 batch 时使用:
```text
PAD = EOS
left padding
attention_mask = 0 on padding
generation prefix ends at the same batch column
```
这次协议修订发生在成功加载模型之前。最初“八格全部同长度”的检查忽略了 system 文本本身
会增加长度;失败后把合同改成“同一 system 水平内等长 + batch 左填充”。这不是结果后改
统计口径,而是修复一个无法构造 tensor 的前置序列错误。
---
## 4. source 合同
正式运行:
```text
4 domains × 4 sources/domain = 16 sources
16 sources × 8 conditions = 128 outputs
```
四域:
- English encyclopedia / WikiText-2 raw validation;
- Chinese news / CLUE TNEWS public test;
- Python code / HumanEval;
- Grade-school math / GSM8K test。
source ID、selection rank 与域内顺序直接复用 Base special-token 正式实验。每条 source
的文本 SHA-256 重新核验;但本轮使用**完整 source 文本**,而路由实验使用固定目标前
23 tokens。这是为了让数学题和代码 prompt 不在半句话上生成。
因此:
```text
同 source 身份 = 是
同 checkpoint = 否(Base → SFT Chat)
同输入文本长度 = 否(23-token prefix → full source)
同证据对象 = 否(routes → generated outputs)
```
每域 4 条是预定 source 顺序的前四条,不是 benchmark 的代表性抽样,也没有能力估计所需
的样本量。
---
## 5. 解码合同
正式运行固定:
```text
decode = greedy
do_sample = false
temperature = unset
top_p = unset
max_new_tokens = 128
use_cache = true
```
官方 `generation_config.json` 另行记录:
```text
do_sample = true
temperature = 0.3
top_p = 0.95
```
它没有用于本轮。选择 greedy 的目的,是先建立逐 token 复跑闸门,不是宣称 greedy 比
官方 sampling 更真实。
---
## 6. 完整 BF16 权重怎样装进本机
官方 Lite model card 写明 BF16 inference 需要一张 40GB GPU。本机 RTX 5090 可见
32,607 MiB,因此不能声称“单卡完整 BF16”。实际运行使用:
```text
Accelerate device_map = auto
max_memory[CUDA:0] = 29 GiB
max_memory[CPU] = 80 GiB
attention = official eager path
```
Hugging Face 的
[Big Model Inference 文档](https://huggingface.co/docs/accelerate/main/concept_guides/big_model_inference)
说明,`hf_device_map` 是查看自动放置的正式入口;CPU-offloaded weights 会通过 hook 在
需要时进入执行设备。正式 device map:
```text
CUDA:0 embed_tokens + layers 0–24
CPU layers 25–26 + final norm + lm_head
```
| 执行账 | 观测 |
|---|---:|
| CUDA 参数 tensor bytes | `28,654,142,464` |
| offload hooks 暴露为 meta 的 tensor bytes | `2,758,825,984` |
| load peak CUDA allocation | `29,786,618,368` bytes |
| formal peak CUDA allocation | `31,578,740,224` bytes ≈ `29.41 GiB` |
| process max RSS | `18,665,132 KiB` ≈ `17.80 GiB` |
| model load | `8.58 s` |
| 16 个 source batch generation | `369.67 s` |
`29GiB max_memory` 约束参数放置,不是包含 KV cache、activation 和 CUDA runtime 的硬峰值;
因此实际 generation peak 高于 29GiB 不构成合同矛盾。
Accelerate-offloaded modules 在 forward 间可把 `parameter.device` 暴露为 `meta`;这里用
`hf_device_map` 判断 CPU placement,不能把 `meta` 字面解释为“参数不存在”。
这些延迟只描述本机 eager + CPU offload 审计,不代表 SGLang、vLLM、FlashMLA 或生产
serving 吞吐。DeepSeek 官方仓库也明确区分 Hugging Face 开源实现与优化 serving 路径。
---
## 7. 正式结果:输出是否 exact
十条预注册比较边:
| edge | exact / 16 | mean common prefix | mean edit distance | mean token similarity |
|---|---:|---:|---:|---:|
| system · EOS | 2 / 16 | 32.875 | 64.625 | 48.26% |
| system · BOS | 1 / 16 | 22.125 | 74.000 | 40.95% |
| system · x | 0 / 16 | 4.875 | 95.188 | 24.15% |
| system · period | 0 / 16 | 4.500 | 94.813 | 24.82% |
| BOS − EOS · S0 | 2 / 16 | 24.563 | 75.125 | 39.89% |
| BOS − EOS · S1 | 1 / 16 | 17.500 | 77.813 | 37.49% |
| x − EOS · S0 | 0 / 16 | 12.563 | 79.188 | 37.18% |
| x − EOS · S1 | 0 / 16 | 9.813 | 85.500 | 29.82% |
| period − EOS · S0 | 1 / 16 | 22.250 | 80.375 | 35.88% |
| period − EOS · S1 | 0 / 16 | 3.688 | 85.750 | 30.30% |
`token similarity = 1 - Levenshtein distance / max(output lengths)`。
能写的最窄结论:
> 在这 16 条固定 source、固定 SFT Chat checkpoint 与 greedy contract 下,system
> message 和一个历史边界 ID 都可能改变最终生成 token 序列;完整 exact 是少数,而不是
> 默认状态。
不能写:
- system / EOS “普遍重要”;
- x / 句点“更差”;
- special token “更稳定”;
- edit distance 越大,能力变化越大;
- 输出分叉由前六层 route TV 解释。
四个 ID 既不是随机从 token 类别总体抽样,16 条 source 也不是总体样本,所以不存在可靠
的 “special vs ordinary 类别效应”。
---
## 8. 分域结果最醒目的差异
### 8.1 official system edge(S0 EOS ↔ S1 EOS)
| 域 | exact / 4 | mean similarity |
|---|---:|---:|
| English encyclopedia | 0 / 4 | 51.37% |
| Chinese news | 0 / 4 | 28.57% |
| Python code | 0 / 4 | 29.10% |
| Grade-school math | 2 / 4 | 83.98% |
数学格在这个极小 cohort 上更常 exact,但不能据此说“数学更不受 system 影响”。它可能来自
source 内容、固定 greedy path、输出长度或 checkpoint 特定模式;每域只有 4 条。
### 8.2 x / period 下的 system edge
system × `x` 和 system × period 都是 0 / 16 exact,平均 similarity 分别为 24.15% /
24.82%。这说明两个固定 ordinary counterfactual 下,system on/off 的 greedy 轨迹在本
cohort 中更常早分叉;它仍不证明 ordinary token 类别导致更强 system effect。
---
## 9. 截断率是主结果,不是脚注
按 condition 的 EOS 完成数:
| condition | hit EOS / 16 | max-128 truncated / 16 |
|---|---:|---:|
| S0 EOS | 2 | 14 |
| S1 EOS | 3 | 13 |
| S0 BOS | 3 | 13 |
| S1 BOS | 3 | 13 |
| S0 x | 2 | 14 |
| S1 x | 8 | 8 |
| S0 period | 2 | 14 |
| S1 period | 8 | 8 |
| **合计** | **31 / 128** | **97 / 128** |
因此本轮保存了 math final-number 与 Python AST parse 作为诊断字段,却不把它们做成能力
比较。比如某格的数学 final number 没出现,可能只是输出仍在推导;某格 AST parse 失败,
也可能只是 code fence 在第 128 token 被截断。
在公平能力表之前,下一协议至少需要:
1. 足够高的 completion contract,或明确定义 stop / answer extractor;
2. 代码实际执行与 sandbox;
3. 每个任务足够样本;
4. 对所有格使用相同完成预算与失败规则;
5. 报告 truncation、refusal、format failure;
6. 与 greedy 分开的多 seed sampling protocol。
---
## 10. 复跑闸门
### 10.1 128-token 长序列复跑
复跑每域首条,共 4 source / 32 格。与正式 16-source 结果按
`source_id × condition` 对齐:
```text
prompt token hash exact 32 / 32
generated token IDs exact 32 / 32
decoded text exact 32 / 32
EOS state exact 32 / 32
```
其余 96 格没有做完整 128-token 独立复跑,网站和文档都不写成“128 / 128 rerun”。
### 10.2 smoke 与正式前缀
4 source / 32 格的 32-token smoke 做了独立复跑,32 / 32 generated token IDs 与文本
exact;正式 128-token 结果的前 32 tokens 也与 smoke 32 / 32 exact。
这支持“固定 greedy contract 可复跑”,不支持 sampling distribution 稳定。
---
## 11. 执行过程中暴露并记录的问题
### 11.1 system 长度
最初错误地检查“八格全等长”。system on 本来就多 16 tokens;改为同一 system 水平内
四个 boundary 格等长,并用显式左填充组成八格 batch。
### 11.2 Python 环境污染
早期命令把 `/usr/lib/python3/dist-packages` 混入 pyenv Python 3.10,触发 Pillow
`_imaging` 二进制不匹配。当前虚拟环境已经包含 pandas / pyarrow / Pillow,正式运行移除
系统路径。
### 11.3 `flash_attn` 动态依赖检查
Transformers 4.41 的 remote-code import scanner 把官方源码中受
`is_flash_attn_2_available()` 保护的 import 仍判成硬依赖。本机没有 flash-attn,正式运行
沿用前轮审计过的 loader,把官方 `configuration_deepseek.py` /
`modeling_deepseek.py` 作为本地 package 直接加载,并走官方 eager attention。
官方模型源码没有打补丁:
```text
modeling_deepseek.py
7d8e5221095286eea991137760893fd7ba52727c0b4ebf48ec09e8bc56b45b9c
```
### 11.4 official sampling config 的 warning
第一次 smoke 明确传 `do_sample=False`,但 model 自带 `.3 / .95` 导致 warning。正式脚本
同时把 `temperature` / `top_p` unset,消除“记录但未使用”的歧义;生成前缀复跑仍逐格
exact。
---
## 12. 这轮最重要的非结论
1. **不是 benchmark。** 每域 4 条,且 97 / 128 截断。
2. **不是 token 类别总体效应。** 两个 special、两个人为选择的 ordinary ID。
3. **不是合法 Chat 对比。** BOS / `x` / period 是单 ID 反事实。
4. **不是 route → output 中介分析。** Base / Chat checkpoint 与输入长度都不同。
5. **不是 sampling robustness。** 只测 greedy。
6. **不是单卡 BF16。** 明确使用 CPU offload。
7. **不是生产性能。** eager + CPU offload 延迟只供执行审计。
8. **不是 RL checkpoint。** Lite-Chat 是 SFT。
---
## 13. 下一闸门
优先级:
1. 把正式行为实验扩成 completion-aware protocol,先解决 97 / 128 截断;
2. 对数学做完整 answer extraction,对代码做 sandbox execution;
3. 独立设计多 seed sampling distribution 对照;
4. 在 Base 与 Chat 上用同一完整输入抽取 27 层 hidden / routes;
5. 预注册 route features 与 output divergence 的关联分析,并明确它仍不是中介因果;
6. 再推进支持硬件上的 FlashMLA、FP8 / pipeline traces 与 R1-like 小模型训练。
本轮真正完成的是从:
```text
“路由看起来变了”
```
前进到:
```text
“完整 SFT Chat 权重实际生成后,哪些固定格子逐 token 相同,
哪些分叉,以及证据为什么仍然不能越界。”
```