# DeepSeek-V2-Lite-Chat 最终生成行为审计 > 状态:Round 04 正式结果 > > 正式原始 JSON:`src/data/deepseek-v2-lite-chat-behavior.json` > > 长序列复跑 JSON:`src/data/deepseek-v2-lite-chat-behavior-repro-1pd.json` > > 执行脚本:`experiments/deepseek/v2_lite_chat_special_token_behavior_probe.py` > > 预注册协议:`research/DEEPSEEK_V2_LITE_CHAT_BEHAVIOR_PROTOCOL.md` ## 0. 一句话结论 固定官方 SFT Chat checkpoint、greedy decoding 与同 source 八格 batch 后,一处历史边界 ID 或一条 system message 经常足以让最终 token 轨迹分叉;但本轮 128 个输出中只有 31 个 在 128-token 上限内遇到 EOS,因此它识别的是**小样本确定性输出敏感性**,不是任务能力、 采样分布稳定性,也不是“某种 token 更好”。 --- ## 1. 为什么上一轮还不能回答这个问题 上一组 special-token family 实验加载的是 `DeepSeek-V2-Lite` Base checkpoint,只观察 前六个 MoE 层的 hidden states 与 top-6 expert routes。它能回答: ```text 同一个 pre-target input ID 被替换以后, 后续目标 token 的专家路径怎样改变? ``` 但它不能回答: ```text 完整模型最后生成了什么? 两格输出是否逐 token 相同? 任务结果有没有改变? ``` 本轮因此换到官方 [`deepseek-ai/DeepSeek-V2-Lite-Chat`](https://huggingface.co/deepseek-ai/DeepSeek-V2-Lite-Chat) SFT checkpoint,并真正调用 `generate()`。DeepSeek 官方仓库把 Lite-Chat 明确列作 16B total / 2.4B active / 32K context 的 **SFT** 模型;它与较大的 V2-Chat (RL) 不是 同一个 checkpoint 身份。 来源: [DeepSeek-V2 官方仓库](https://github.com/deepseek-ai/DeepSeek-V2)、 [DeepSeek-V2 报告](https://arxiv.org/abs/2405.04434)。 这也意味着: - Base routing 与 Chat generation 是两种证据层; - 不能把二者的 effect size 直接相减; - 不能把 Chat 输出分叉倒推成某个 expert 的因果中介; - 不能把 SFT checkpoint 写成 R1、RL 或当前线上 DeepSeek。 --- ## 2. 固定了什么 ### 2.1 模型工件 | 对象 | 固定值 | |---|---:| | 模型 | `deepseek-ai/DeepSeek-V2-Lite-Chat` | | revision | `85864749cd611b4353ce1decdb286193298f64c7` | | checkpoint 身份 | SFT Chat | | dtype | BF16 | | safetensors shards | 4 | | index tensor bytes | `31,412,968,448` | | shard file bytes(含 headers) | `31,413,626,576` | | 文件数 | 12 | | revision metadata | 12 / 12 同一 revision | 运行时参数 bytes 按 dtype 汇总也是 `31,412,968,448`,全部为 `torch.bfloat16`,与 index 的 `metadata.total_size` 精确相等。 正式原始结果: ```text 542,559 bytes SHA-256 54496955d0dd20a3116e6b758a55f56e2444c43dd99bfef94a2cb7b140e6196e ``` 长序列复跑: ```text 153,819 bytes SHA-256 b74c31606c0fc70a8e52cc6c3d6135c4e916d698b4d5dae51863124d025b0c72 ``` ### 2.2 tokenizer Chat 与 Base 的 `tokenizer.json`、`tokenizer_config.json` 逐字节相同: ```text tokenizer.json 41f3bf64213da8c012d8bd0871a58a1fdf70463e8f08f110ddbb1082f529f669 tokenizer_config.json 31181eaf79394ea26728d95ecb54fe7c8413e6f56085dbabc8b0818134380ec8 ``` 固定边界 ID: | level | ID | 类别 | 官方合法历史边界 | |---|---:|---|---| | EOS | `100001` | special;同时是 PAD alias | 是 | | BOS | `100000` | special | 否 | | `x` | `87` | ordinary content | 否 | | `.` | `13` | ordinary punctuation | 否 | EOS 与 BOS 穷尽这个固定 tokenizer 的 special-token inventory。BOS / `x` / 句点三格 都不是官方合法 Chat 序列,它们只是在 official rendering 后替换一个 pre-target input ID 的反事实控制。 --- ## 3. 八格设计 ```text system off/on × EOS / BOS / x / period ``` 同一 system 水平内: - 四格 token 数相同; - target content、generation prompt 与 target position 相同; - EOS 格相对官方序列改 0 个 ID; - BOS / `x` / period 各恰好改 1 个 ID。 system on 比 system off 多 16 个 token。八格同 batch 时使用: ```text PAD = EOS left padding attention_mask = 0 on padding generation prefix ends at the same batch column ``` 这次协议修订发生在成功加载模型之前。最初“八格全部同长度”的检查忽略了 system 文本本身 会增加长度;失败后把合同改成“同一 system 水平内等长 + batch 左填充”。这不是结果后改 统计口径,而是修复一个无法构造 tensor 的前置序列错误。 --- ## 4. source 合同 正式运行: ```text 4 domains × 4 sources/domain = 16 sources 16 sources × 8 conditions = 128 outputs ``` 四域: - English encyclopedia / WikiText-2 raw validation; - Chinese news / CLUE TNEWS public test; - Python code / HumanEval; - Grade-school math / GSM8K test。 source ID、selection rank 与域内顺序直接复用 Base special-token 正式实验。每条 source 的文本 SHA-256 重新核验;但本轮使用**完整 source 文本**,而路由实验使用固定目标前 23 tokens。这是为了让数学题和代码 prompt 不在半句话上生成。 因此: ```text 同 source 身份 = 是 同 checkpoint = 否(Base → SFT Chat) 同输入文本长度 = 否(23-token prefix → full source) 同证据对象 = 否(routes → generated outputs) ``` 每域 4 条是预定 source 顺序的前四条,不是 benchmark 的代表性抽样,也没有能力估计所需 的样本量。 --- ## 5. 解码合同 正式运行固定: ```text decode = greedy do_sample = false temperature = unset top_p = unset max_new_tokens = 128 use_cache = true ``` 官方 `generation_config.json` 另行记录: ```text do_sample = true temperature = 0.3 top_p = 0.95 ``` 它没有用于本轮。选择 greedy 的目的,是先建立逐 token 复跑闸门,不是宣称 greedy 比 官方 sampling 更真实。 --- ## 6. 完整 BF16 权重怎样装进本机 官方 Lite model card 写明 BF16 inference 需要一张 40GB GPU。本机 RTX 5090 可见 32,607 MiB,因此不能声称“单卡完整 BF16”。实际运行使用: ```text Accelerate device_map = auto max_memory[CUDA:0] = 29 GiB max_memory[CPU] = 80 GiB attention = official eager path ``` Hugging Face 的 [Big Model Inference 文档](https://huggingface.co/docs/accelerate/main/concept_guides/big_model_inference) 说明,`hf_device_map` 是查看自动放置的正式入口;CPU-offloaded weights 会通过 hook 在 需要时进入执行设备。正式 device map: ```text CUDA:0 embed_tokens + layers 0–24 CPU layers 25–26 + final norm + lm_head ``` | 执行账 | 观测 | |---|---:| | CUDA 参数 tensor bytes | `28,654,142,464` | | offload hooks 暴露为 meta 的 tensor bytes | `2,758,825,984` | | load peak CUDA allocation | `29,786,618,368` bytes | | formal peak CUDA allocation | `31,578,740,224` bytes ≈ `29.41 GiB` | | process max RSS | `18,665,132 KiB` ≈ `17.80 GiB` | | model load | `8.58 s` | | 16 个 source batch generation | `369.67 s` | `29GiB max_memory` 约束参数放置,不是包含 KV cache、activation 和 CUDA runtime 的硬峰值; 因此实际 generation peak 高于 29GiB 不构成合同矛盾。 Accelerate-offloaded modules 在 forward 间可把 `parameter.device` 暴露为 `meta`;这里用 `hf_device_map` 判断 CPU placement,不能把 `meta` 字面解释为“参数不存在”。 这些延迟只描述本机 eager + CPU offload 审计,不代表 SGLang、vLLM、FlashMLA 或生产 serving 吞吐。DeepSeek 官方仓库也明确区分 Hugging Face 开源实现与优化 serving 路径。 --- ## 7. 正式结果:输出是否 exact 十条预注册比较边: | edge | exact / 16 | mean common prefix | mean edit distance | mean token similarity | |---|---:|---:|---:|---:| | system · EOS | 2 / 16 | 32.875 | 64.625 | 48.26% | | system · BOS | 1 / 16 | 22.125 | 74.000 | 40.95% | | system · x | 0 / 16 | 4.875 | 95.188 | 24.15% | | system · period | 0 / 16 | 4.500 | 94.813 | 24.82% | | BOS − EOS · S0 | 2 / 16 | 24.563 | 75.125 | 39.89% | | BOS − EOS · S1 | 1 / 16 | 17.500 | 77.813 | 37.49% | | x − EOS · S0 | 0 / 16 | 12.563 | 79.188 | 37.18% | | x − EOS · S1 | 0 / 16 | 9.813 | 85.500 | 29.82% | | period − EOS · S0 | 1 / 16 | 22.250 | 80.375 | 35.88% | | period − EOS · S1 | 0 / 16 | 3.688 | 85.750 | 30.30% | `token similarity = 1 - Levenshtein distance / max(output lengths)`。 能写的最窄结论: > 在这 16 条固定 source、固定 SFT Chat checkpoint 与 greedy contract 下,system > message 和一个历史边界 ID 都可能改变最终生成 token 序列;完整 exact 是少数,而不是 > 默认状态。 不能写: - system / EOS “普遍重要”; - x / 句点“更差”; - special token “更稳定”; - edit distance 越大,能力变化越大; - 输出分叉由前六层 route TV 解释。 四个 ID 既不是随机从 token 类别总体抽样,16 条 source 也不是总体样本,所以不存在可靠 的 “special vs ordinary 类别效应”。 --- ## 8. 分域结果最醒目的差异 ### 8.1 official system edge(S0 EOS ↔ S1 EOS) | 域 | exact / 4 | mean similarity | |---|---:|---:| | English encyclopedia | 0 / 4 | 51.37% | | Chinese news | 0 / 4 | 28.57% | | Python code | 0 / 4 | 29.10% | | Grade-school math | 2 / 4 | 83.98% | 数学格在这个极小 cohort 上更常 exact,但不能据此说“数学更不受 system 影响”。它可能来自 source 内容、固定 greedy path、输出长度或 checkpoint 特定模式;每域只有 4 条。 ### 8.2 x / period 下的 system edge system × `x` 和 system × period 都是 0 / 16 exact,平均 similarity 分别为 24.15% / 24.82%。这说明两个固定 ordinary counterfactual 下,system on/off 的 greedy 轨迹在本 cohort 中更常早分叉;它仍不证明 ordinary token 类别导致更强 system effect。 --- ## 9. 截断率是主结果,不是脚注 按 condition 的 EOS 完成数: | condition | hit EOS / 16 | max-128 truncated / 16 | |---|---:|---:| | S0 EOS | 2 | 14 | | S1 EOS | 3 | 13 | | S0 BOS | 3 | 13 | | S1 BOS | 3 | 13 | | S0 x | 2 | 14 | | S1 x | 8 | 8 | | S0 period | 2 | 14 | | S1 period | 8 | 8 | | **合计** | **31 / 128** | **97 / 128** | 因此本轮保存了 math final-number 与 Python AST parse 作为诊断字段,却不把它们做成能力 比较。比如某格的数学 final number 没出现,可能只是输出仍在推导;某格 AST parse 失败, 也可能只是 code fence 在第 128 token 被截断。 在公平能力表之前,下一协议至少需要: 1. 足够高的 completion contract,或明确定义 stop / answer extractor; 2. 代码实际执行与 sandbox; 3. 每个任务足够样本; 4. 对所有格使用相同完成预算与失败规则; 5. 报告 truncation、refusal、format failure; 6. 与 greedy 分开的多 seed sampling protocol。 --- ## 10. 复跑闸门 ### 10.1 128-token 长序列复跑 复跑每域首条,共 4 source / 32 格。与正式 16-source 结果按 `source_id × condition` 对齐: ```text prompt token hash exact 32 / 32 generated token IDs exact 32 / 32 decoded text exact 32 / 32 EOS state exact 32 / 32 ``` 其余 96 格没有做完整 128-token 独立复跑,网站和文档都不写成“128 / 128 rerun”。 ### 10.2 smoke 与正式前缀 4 source / 32 格的 32-token smoke 做了独立复跑,32 / 32 generated token IDs 与文本 exact;正式 128-token 结果的前 32 tokens 也与 smoke 32 / 32 exact。 这支持“固定 greedy contract 可复跑”,不支持 sampling distribution 稳定。 --- ## 11. 执行过程中暴露并记录的问题 ### 11.1 system 长度 最初错误地检查“八格全等长”。system on 本来就多 16 tokens;改为同一 system 水平内 四个 boundary 格等长,并用显式左填充组成八格 batch。 ### 11.2 Python 环境污染 早期命令把 `/usr/lib/python3/dist-packages` 混入 pyenv Python 3.10,触发 Pillow `_imaging` 二进制不匹配。当前虚拟环境已经包含 pandas / pyarrow / Pillow,正式运行移除 系统路径。 ### 11.3 `flash_attn` 动态依赖检查 Transformers 4.41 的 remote-code import scanner 把官方源码中受 `is_flash_attn_2_available()` 保护的 import 仍判成硬依赖。本机没有 flash-attn,正式运行 沿用前轮审计过的 loader,把官方 `configuration_deepseek.py` / `modeling_deepseek.py` 作为本地 package 直接加载,并走官方 eager attention。 官方模型源码没有打补丁: ```text modeling_deepseek.py 7d8e5221095286eea991137760893fd7ba52727c0b4ebf48ec09e8bc56b45b9c ``` ### 11.4 official sampling config 的 warning 第一次 smoke 明确传 `do_sample=False`,但 model 自带 `.3 / .95` 导致 warning。正式脚本 同时把 `temperature` / `top_p` unset,消除“记录但未使用”的歧义;生成前缀复跑仍逐格 exact。 --- ## 12. 这轮最重要的非结论 1. **不是 benchmark。** 每域 4 条,且 97 / 128 截断。 2. **不是 token 类别总体效应。** 两个 special、两个人为选择的 ordinary ID。 3. **不是合法 Chat 对比。** BOS / `x` / period 是单 ID 反事实。 4. **不是 route → output 中介分析。** Base / Chat checkpoint 与输入长度都不同。 5. **不是 sampling robustness。** 只测 greedy。 6. **不是单卡 BF16。** 明确使用 CPU offload。 7. **不是生产性能。** eager + CPU offload 延迟只供执行审计。 8. **不是 RL checkpoint。** Lite-Chat 是 SFT。 --- ## 13. 下一闸门 优先级: 1. 把正式行为实验扩成 completion-aware protocol,先解决 97 / 128 截断; 2. 对数学做完整 answer extraction,对代码做 sandbox execution; 3. 独立设计多 seed sampling distribution 对照; 4. 在 Base 与 Chat 上用同一完整输入抽取 27 层 hidden / routes; 5. 预注册 route features 与 output divergence 的关联分析,并明确它仍不是中介因果; 6. 再推进支持硬件上的 FlashMLA、FP8 / pipeline traces 与 R1-like 小模型训练。 本轮真正完成的是从: ```text “路由看起来变了” ``` 前进到: ```text “完整 SFT Chat 权重实际生成后,哪些固定格子逐 token 相同, 哪些分叉,以及证据为什么仍然不能越界。” ```