13 KiB
DeepSeek-V2-Lite-Chat 多种子采样稳健性审计
执行日期:2026-07-30
协议:
llm-atlas-deepseek-chat-sampling-v1模型:
deepseek-ai/DeepSeek-V2-Lite-Chatrevision:
85864749cd611b4353ce1decdb286193298f64c7预注册:
research/DEEPSEEK_V2_LITE_CHAT_SAMPLING_PROTOCOL.md
0. 一句话结论
在固定官方 Chat checkpoint、固定四条 source、固定八格 prompt batch 与官方
temperature=0.3 / top_p=0.95 下:
256 sampled outputs
├─ 251 natural EOS
├─ 5 budget truncated(全部来自 English / system off)
├─ 242 个不同的完整 token trajectory hash
├─ Math:62 / 64 strict exact
├─ Code:63 / 64 official tests pass
└─ 新进程 R0/R1:64 / 64 全合同字段 exact
这说明在这份执行合同下:
- sampling 确实产生了跨 seed 轨迹分叉;
- 同 seed 可以在新进程精确复现;
- 一条流畅、自然结束、可解析或可执行的回答仍然可能答错;
- 四条 source 与八个 seed 仍远不足以估计 benchmark 能力或完整生成分布。
1. 为什么在 greedy 之后还要做这一轮
Round 05 的 512-token 实验固定 do_sample=false。它验证的是每一步都取最高概率 token
时的单条轨迹:
同一 source + 同一 checkpoint + 不同输入边界
→ greedy token 轨迹会怎样分叉
greedy 不能回答:
- 换 seed 后,同一条件是否只会复述同一条轨迹;
- greedy 轨迹是否会出现在有限的 nucleus samples 里;
- 两个条件的八样本集合是否有 exact 重合;
- 完成与任务结果是否会随 seed 改变。
本轮只补这个证据缺口,不回头把 4 条 source 写成能力评测。
2. 冻结合同
2.1 四条 source
| domain | source ID | 角色 |
|---|---|---|
| English | wikitext2/raw-validation/0443 |
长 continuation,最容易触及 512 上限 |
| Chinese | tnews/test/4855 |
中文分类式输入 |
| Code | HumanEval/31 |
is_prime,可用官方 tests 执行 |
| Math | gsm8k/test/1069 |
可抽取数值 gold |
它们是前序固定 source 排序中每域第 1 条,不是看到结果后重选。
2.2 八格输入
system off/on × EOS / BOS / x / period
固定行顺序:
s0_eos, s1_eos, s0_bos, s1_bos,
s0_x, s1_x, s0_period, s1_period
32 / 32 个 prompt token hash 与 Round 05 greedy 运行 exact。只有 EOS 是官方聊天序列;
BOS、x 与句点都是单 ID 反事实,不能称为官方有效 chat。
2.3 sampling 参数
checkpoint 自带的 generation_config.json 固定:
do_sample = true
temperature = 0.3
top_p = 0.95
本轮另外显式传入:
top_k = 0
max_new_tokens = 512
use_cache = true
top_k=0 很重要:否则 Transformers 的通用默认 top-k 可能在不显眼处继续截断候选集合。
2.4 八个 seed
八个 base seed 全部由协议字符串做 SHA-256 派生:
| replicate | base seed |
|---|---|
| R0 | 19,683,830 |
| R1 | 1,560,062,173 |
| R2 | 3,978,401,375 |
| R3 | 1,280,933,274 |
| R4 | 1,467,459,869 |
| R5 | 1,297,359,489 |
| R6 | 2,722,953,988 |
| R7 | 3,330,978,061 |
每个 source 再由 base seed 与 source ID 派生独立 run seed。每次 batch 前执行
torch.manual_seed() 与 torch.cuda.manual_seed_all(),并记录 CPU/CUDA RNG 的前后
state hash。
2.5 batch-seed aligned,不是 common random numbers
Transformers 4.41.2 对 batch 调一次 torch.multinomial。八行共享 run seed 与调用
时序,但不同 row 消费不同 RNG 子流。因此:
- 可以比较相同 replicate label 的两格;
- 只能称为 batch-seed aligned;
- 不能当作逐行共享同一随机数的 paired causal design;
- 本轮不报告 paired p-value 或置信区间。
3. 执行闸门
3.1 16-token smoke
4 sources × R0/R1 × 8 conditions = 64 short outputs
4 sources × R0 replay × 8 conditions = 32 replay outputs
结果:
| gate | 结果 |
|---|---|
| prompt hash exact | 32 / 32 |
| R0 同进程重放全合同 exact | 32 / 32 |
| R0 vs R1 可比格 | 32 |
| R0 vs R1 分叉 | 18 |
| OOM / NaN / exception | 0 |
16-token smoke 的 64 条输出全部触顶,符合技术检查的预期;它们没有进入正式统计。
3.2 正式网格
4 sources × 8 seeds × 8 conditions = 256 outputs
结构审计:
| 检查 | 结果 |
|---|---|
| sources | 4 |
| runs / source | 8 |
| outputs / source | 64 |
| outputs / condition | 32 |
| total outputs | 256 |
| prompt hash exact | 32 / 32 |
| missing / duplicate grid cells | 0 |
生成耗时 2,251.225s,CUDA peak allocated 30,868,238,336 bytes。这是 CPU-offloaded
eager 机制审计延迟,不是服务吞吐。
4. stopping 与轨迹多样性
4.1 总账
| 指标 | 结果 |
|---|---|
| natural EOS | 251 / 256 |
| budget truncated | 5 / 256 |
| 全局 unique trajectory hashes | 242 / 256 |
| R0/R1 同格 trajectory 不同 | 31 / 32 |
| greedy trajectory 出现在八样本集合 | 9 / 32 source×condition sets |
242/256 unique 只说明完整 token ID 序列的 hash 不同,不说明有 242 种语义或思路。
4.2 每个 condition
| condition | EOS / 32 | mean tokens | 四个 source-set 的 unique 总数 / 32 | 8/8 unique sets / 4 | greedy included / 4 |
|---|---|---|---|---|---|
s0_eos |
29 | 290.1 | 31 | 3 | 1 |
s1_eos |
32 | 241.4 | 32 | 4 | 2 |
s0_bos |
32 | 265.4 | 30 | 3 | 1 |
s1_bos |
32 | 221.2 | 32 | 4 | 1 |
s0_x |
30 | 263.6 | 32 | 4 | 0 |
s1_x |
32 | 146.8 | 30 | 3 | 2 |
s0_period |
32 | 259.5 | 32 | 4 | 0 |
s1_period |
32 | 121.7 | 26 | 3 | 2 |
32 个 source×condition set 中:
28 sets:8 unique trajectories
2 sets:6 unique trajectories
1 set :7 unique trajectories
1 set :2 unique trajectories
只有 HumanEval/31 × s1_period 收缩到 2 条完整轨迹;它不是“低创造力”的证明,只是该
source/condition/参数下八次抽样的 exact-sequence 重复。
4.3 五个截断
五格全部属于 wikitext2/raw-validation/0443:
| replicate | condition | tokens |
|---|---|---|
| R0 | s0_x |
512 |
| R1 | s0_eos |
512 |
| R4 | s0_eos |
512 |
| R5 | s0_x |
512 |
| R6 | s0_eos |
512 |
所以“全局 251/256 EOS”不能简化成每条 source 都同样容易完成。
5. 十条 edge 的集合比较
每个 source×edge 同时计算:
- 8 个相同 replicate label 对的 token similarity;
- 左集合每条到右集合的最高 similarity;
- 右集合每条到左集合的最高 similarity;
- 两个方向 16 个最近邻值的平均;
- 完整 trajectory hash-set 的 intersection / union。
四条 source 的描述性均值:
| edge | aligned similarity | symmetric nearest similarity | exact intersection / union |
|---|---|---|---|
| system · EOS | .366 | .492 | 1 / 62 |
| system · BOS | .314 | .491 | 0 / 62 |
| system · x | .275 | .392 | 0 / 62 |
| system · period | .253 | .355 | 0 / 58 |
| BOS − EOS · S0 | .444 | .589 | 0 / 61 |
| BOS − EOS · S1 | .430 | .589 | 0 / 64 |
| x − EOS · S0 | .332 | .470 | 0 / 63 |
| x − EOS · S1 | .278 | .395 | 0 / 62 |
| period − EOS · S0 | .404 | .546 | 1 / 62 |
| period − EOS · S1 | .261 | .375 | 0 / 58 |
40 个 source-level edge set 中只有两个出现跨侧 exact trajectory 重合:
gsm8k/test/1069 × system_eos;gsm8k/test/1069 × period_at_s0。
nearest-neighbor 总是高于相同 label 对并不奇怪:它从 8 条右侧样本里主动选择最接近的一条。 这不是无偏分布距离,也不能据此给 boundary effect 排名。
6. 任务账:完成、可评与正确分开
6.1 GSM8K
| 指标 | 结果 |
|---|---|
| sampled outputs | 64 |
| numeric evaluator covered | 64 |
| strict-complete exact | 62 |
answer 300 |
62 |
answer 100 |
2 |
| unique absolute majority | 300 |
| gold | 300 |
两条失败分别位于:
- R4 /
s1_x; - R7 /
s1_eos。
两条都不是抽取器误判。回答明确写出 boxed 100,错误推理是:
1/4 trucks failed
→ 错写成只有 1/4 trucks delivered
→ 5 trucks × 20 tons = 100
正确逻辑应保留 3/4 × 20 = 15 辆车,得到 300。这说明 natural EOS、明确 final
marker 与可抽取数值都不能保证 reasoning 正确。
这里的多数答案只是 1 条 GSM8K source 上的八条件×八 seed 描述,不是标准 self-consistency benchmark。
6.2 HumanEval
| 指标 | 结果 |
|---|---|
| sampled outputs | 64 |
| Python AST parse | 64 |
| sandbox evaluated | 64 |
| official tests pass | 63 |
| assertion failure | 1 |
| unique candidate execution keys | 24 |
| exact duplicate cache hits | 40 |
唯一失败位于 R5 / s0_period。候选代码:
- 正确排除了
n < 2; - 正确特判
n == 2; - 只检查奇数除数;
- 没有先排除大于 2 的偶数。
因此代码流畅、code fence 闭合、AST 合法、可以执行,却对偶数输入返回错误。这个案例直接 展示了四张账:
NATURAL EOS
→ TASK TERMINAL
→ EVALUATOR COVERED
→ ASSERTION FAILED
每个唯一 candidate 的执行环境:
python:3.11-alpine
@sha256:25976e9d34a0fab1f278cae931f34c8303d97bf0c0d7f85b6b4dcf641d7702a4
network none · read-only filesystem · no host mounts
user 65534:65534 · cap-drop ALL · no-new-privileges
256MiB memory/swap · pids 64 · cpus .5 · timeout 5s
cache 只复用完全相同的 candidate hash + task-test hash + harness hash;每格评测行仍保留。
7. 新进程复现
正式运行结束后重新启动 Python、重新加载 checkpoint,只执行 R0/R1:
4 sources × 2 seeds × 8 conditions = 64 outputs
逐格结果:
| field | exact |
|---|---|
| run seed | 64 / 64 |
| prompt hash | 64 / 64 |
| complete generated token IDs | 64 / 64 |
| decoded text | 64 / 64 |
| EOS state | 64 / 64 |
| truncation state | 64 / 64 |
| CPU RNG pre-state hash | 64 / 64 |
| CUDA RNG pre-state hash | 64 / 64 |
| all preregistered fields | 64 / 64 |
这只能写成“R0/R1 的 64 条轨迹在固定环境中独立复现”。R2–R7 没有新进程复跑,不能写 256 / 256。
8. artifact hash 链
| artifact | bytes | SHA-256 |
|---|---|---|
| formal raw | 1,852,892 | 46c7edfce1409e798d6b0f06e905dd3a9d4ab6911acd8e42e2ec9d06d65345af |
| independent eval | 316,800 | 078f486e021ffd0e4af8ad942d001a9f2429933b2fb0fb3523b7d1da5b24c4d8 |
| R0/R1 rerun raw | 537,669 | 72d050eace531bb38fb2592e39dc7f1b9c9003306c92c568dcd4acaaa691f6a0 |
| reproduction comparison | 29,938 | 4d59a775943459604ea2d9976bb3be01ff7b3b08d71eb0ca22c376122ef82d15 |
| website compact | 371,495 | c014dc0ed6c27b19a61e7abb9b84a077b414e8472d40cc69d025693c25e1372e |
compact builder 在写前端 JSON 前强制验证:
- evaluator 声明的 sampling input SHA 与 formal raw 文件一致;
- reproduction 声明的 formal/rerun SHA 与两份 raw 一致;
- reproduction
all_preregistered_fields_exact == cells。
任一失败都会中止构建。
9. 依赖与放置
| object | value |
|---|---|
| PyTorch | 2.11.0+cu128 |
| Transformers | 4.41.2 |
| dtype | BF16 |
| local GPU | RTX 5090 · 32,607 MiB |
| CUDA placement | embedding + layers 0–23 |
| CPU offload | layers 24–26 + final norm + LM head |
| max memory | GPU 28 GiB / CPU 80 GiB |
| allocator | expandable_segments:True |
| formal generation time | 2,251.225 s |
| peak CUDA allocated | 30,868,238,336 bytes |
官方模型卡写出的 40GB 单 GPU BF16 边界高于本机容量,所以这不是 single-GPU BF16 execution。
10. 一手来源
- DeepSeek-V2-Lite-Chat model card
- Pinned official generation_config.json
- Transformers 4.41.2 generation utils
- The Curious Case of Neural Text Degeneration / nucleus sampling
- Self-Consistency Improves Chain of Thought Reasoning
- HumanEval repository
- GSM8K paper
11. 可以说与不能说
可以说:
- 在固定四条 source 与官方 sampling 参数下,R0/R1 有 31/32 同格轨迹分叉;
- 256 条样本中 251 条自然 EOS,242 条完整 token trajectory hash 不同;
- 这 1 条 GSM8K 的 64 个样本中 62 个 strict exact;
- 这 1 条 HumanEval 的 64 个样本中 63 个通过官方 tests;
- 固定环境下 R0/R1 新进程复跑 64/64 全合同 exact。
不能说:
- 八个 seed 已恢复完整生成分布;
- 242 个 hash 表示 242 种语义;
- 4 条 source 是 benchmark;
- 62/64 与 63/64 可以横比标准 accuracy 或 pass@k;
- majority answer 是标准 self-consistency;
- batch-seed aligned 是 common random numbers;
- nearest-neighbor 是无偏分布距离;
- 某个 condition 更“有创造力”;
- BOS、
x、句点是官方有效聊天格式; - output 差异由某层 hidden state 或 router 因果中介;
- fixed-version exact replay 能跨 PyTorch、CUDA 或硬件保证。