feat: audit DeepSeek Chat across sources
This commit is contained in:
+14
-3
@@ -14,7 +14,7 @@
|
||||
| 表示、位置与残差高速公路 | 完成首版 | 81% | 真实 hidden-state / norm traces、长上下文位置外推与深层稳定性消融 |
|
||||
| Scaling Laws | 完成首版 | 74% | 真实拟合复现、置信区间与更多模型族对照 |
|
||||
| 数据工程与预训练配方 | 完成首版 | 73% | FineWeb / DCLM 逐图精读、真实去重误伤与 mixture traces |
|
||||
| DeepSeek 专题 | 六轮实证进行中 | 99% | 扩大 sampling 的 source/task 覆盖,再推进干预式 mediation、SM90 FlashMLA、FP8/pipeline 与 R1-like RL 复现 |
|
||||
| DeepSeek 专题 | 七轮实证进行中 | 99% | 扩大到可做 task-level bootstrap,加入 per-row RNG,再推进干预式 mediation、SM90 FlashMLA、FP8/pipeline 与 R1-like RL |
|
||||
| 指令微调与人类偏好 | 完成首版 | 75% | 真实偏好分歧、RM 长度偏置与 PPO/DPO 小模型复现 |
|
||||
| 推理与测试时扩展 | 完成首版 | 76% | 真实模型采样曲线、PRM 案例与逐篇图表精读 |
|
||||
| 工具使用与长程 Agent | 完成首版 | 74% | 真实环境 traces、cross-harness 对照、Agent RL 曲线与安全案例 |
|
||||
@@ -41,7 +41,7 @@
|
||||
- [x] 完成 486 篇关键论文索引,覆盖 16 个标签专题与 Kimi/DeepSeek 聚光主线。
|
||||
- [x] 完成可检索、可按专题筛选的论文库页面。
|
||||
- [x] 完成 K3、语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全十七篇首版长文。
|
||||
- [x] 完成 K3 三轴架构、八联报告实验与四联开放工件实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 二十联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等八十七个原创交互视图。
|
||||
- [x] 完成 K3 三轴架构、八联报告实验与四联开放工件实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 二十一联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等八十八个原创交互视图。
|
||||
- [x] 完成长上下文首版:五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。
|
||||
- [x] 核验 FlashAttention、DeepSeek-V2/V3.2/V4、Kimi Linear/K3 等六份论文原文,并建立长上下文研究账本。
|
||||
- [x] 核验 Switch、ST-MoE、DeepSeekMoE、Loss-Free、V3、LatentMoE 与 K3 原文,并建立 MoE 研究账本。
|
||||
@@ -256,11 +256,18 @@
|
||||
- [x] 第二十个 DeepSeek 交互实验以四页签讲解八 seed 轨迹显微镜、Math/Code 失败账、aligned/nearest 样本集合比较与随机性×复现合同;完整协议、结果审计、运行脚本与 compact 数据均已记录在项目本地。
|
||||
- [x] Round 06 本地闸门通过:80 个受检文件零诊断,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;DeepSeek 历史全量与 sampling 专项真实 Chrome 回归均通过,桌面/390px 移动端无文档级溢出、无 offender、无运行时异常。
|
||||
- [x] DeepSeek Round 06 以源提交 `580f696`、不可变镜像 `20260729T180228Z-580f696` 发布;OCI digest `sha256:aabc13a4…fa3ed`,复用 NAS `12010→8080`、NPM host 31 / cert 41 与门户 order 180。NAS / VPS / HTTPS / gzip / immutable assets、sampling 专项和 DeepSeek 全量生产 Chrome 回归均通过,保留 `20260729T162625Z-8bb488f` 回滚。
|
||||
- [x] DeepSeek Round 07 在正式输出前冻结跨来源 sampling 协议:16 条 source 来自既有 SHA-256 selection rank,每域 4 条;4 个新 SHA-256 seed、`system off/on × EOS/句点` 四行 batch 与 `.3/.95/top-k 0` 固定。source 是主要覆盖单位,seed 是题内重复;句点是受 Round 06 启发的定向复查,不冒充盲确认。
|
||||
- [x] 16-token smoke 的 64 / 64 prompt hashes 与 Round 05 exact,R0 同进程 replay 64 / 64 全合同 exact,R0/R1 34 / 64 格分叉且无 OOM/NaN/exception;正式 256 条输出为 250 natural EOS、6 条截断、247 个完整 token trajectory hashes,R0/R1 63 / 64 同格分叉。
|
||||
- [x] 跨题 evaluator 分离 stopping、task terminal、coverage 与 correctness:Math 47 / 64、Code 52 / 64;Math 四题为 `14/16 · 16/16 · 8/16 · 9/16`,Code 四题为 `16/16 · 8/16 · 12/16 · 16/16`。修复跨不同 GSM8K gold 聚合 final-answer frequency 的统计身份错误,单题前序总账不变。
|
||||
- [x] source-blocked 分析逐 source 计算 contrast,不报告 p-value 或 population CI。句点在 11 / 16 条 source 上缩短平均长度;English `/0443` 为 −121,但另三条为 +44.5 / +41 / +1.375,导致 domain mean −8.5 与 3 / 4 source 的正方向相反。Code task interaction 的 +1 / −1 也在 domain mean 0 中完全抵消。
|
||||
- [x] 全新进程重跑 16 sources × R0 × 4 conditions;run seed、prompt hash、完整 token IDs、decoded text、EOS、truncation 与 CPU/CUDA RNG pre-state 八项均 64 / 64 exact。formal/eval/rerun/compare/analysis SHA-256 为 `f013132…d7f7c` / `e88b274…b8975` / `143dc9d…02a6` / `ec4a478…a556` / `d0dece3…bc84`。
|
||||
- [x] 第二十一个 DeepSeek 交互实验以四页签讲解 source→seed 层级、Math/Code 4×4 task matrix、source-level 方向与均值反例、64 格八字段复现及五段 hash chain;协议、完整审计、探针、独立 evaluator、analysis 与 compact builder 全部落盘。
|
||||
- [x] Round 07 本地闸门通过:83 个 Astro 文件零诊断,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;Python 探针/评测/分析编译通过,compact builder 重建 SHA-256 不变。跨来源专项与 DeepSeek 全量真实 Chrome 回归通过,4×4 矩阵、English 反例、键盘页签、桌面/390px 移动端无溢出及零运行时异常均有自动断言。
|
||||
|
||||
## 正在进行
|
||||
|
||||
- [ ] K3 三轮下一闸门:获得真实 token hidden states、expert load 与 cache traces,解释或修订 `A_log [128]` 工件冲突,再做 Figure 3/4/5 数值重绘和独立小模型复现。
|
||||
- [ ] DeepSeek 六轮下一闸门:扩大 GSM8K/HumanEval 与语言 source 的 sampling 覆盖,加入干预式 mediation;再推进 SM90 FlashMLA、FP8 / pipeline traces 与 R1-like RL 小模型复现。
|
||||
- [ ] DeepSeek 七轮下一闸门:扩大到可做 task-level bootstrap 的预注册 source 抽样框,加入 per-row RNG/common-random-number 对照与 failure taxonomy;再推进干预式 mediation、SM90 FlashMLA、FP8 / pipeline traces 与 R1-like RL 小模型复现。
|
||||
- [ ] 表示、位置与残差二轮:真实 hidden-state / norm traces、长上下文位置外推复现与 mHC / AttnRes 深层稳定性消融。
|
||||
- [ ] 评测安全二轮:真实 cross-harness / pass@k 复跑、Judge 元评测、动态污染与过拒案例。
|
||||
- [ ] 推理服务二轮:真实 GPU kernel / workload traces、功耗与成本、跨 vLLM / SGLang / TensorRT-LLM 复现。
|
||||
@@ -438,6 +445,10 @@
|
||||
| 2026-07-30 | sampled completion 与任务失败分账 | 256 条中 251 natural EOS、242 unique token trajectories;Math 62/64、Code 63/64,具体错误回到推理与官方 tests |
|
||||
| 2026-07-30 | sampling 的随机性与复现同时过闸 | R0/R1 31/32 同格分叉;新进程复跑 64 格的 seed、prompt、token、text、stop 与 CPU/CUDA RNG pre-state 八项全部 exact |
|
||||
| 2026-07-30 | DeepSeek sampling robustness 里程碑以 `20260729T180228Z-580f696` 发布 | OCI digest `sha256:aabc13a4…fa3ed`;复用 NAS `12010→8080`、NPM 31 / cert 41、门户 order 180;sampling 专项与 DeepSeek 全量生产 Chrome 回归通过,保留 `20260729T162625Z-8bb488f` 回滚 |
|
||||
| 2026-07-30 | Round 07 把 source 冻结为主要覆盖单位 | 16 sources × 4 seeds × 4 conditions;seed 是题内重复,四题方向不升级为 benchmark、p-value 或总体置信区间 |
|
||||
| 2026-07-30 | 跨题矩阵替代单题 micro total | Math 四题 14/16、16/16、8/16、9/16;Code 四题 16/16、8/16、12/16、16/16;跨不同 gold 不再聚合 final-answer frequency |
|
||||
| 2026-07-30 | 域均值必须与 source 方向同时展示 | English 句点 contrast 的均值 −8.5,但 3/4 source 为正;Code task interaction 的 +1 与 −1 在域均值 0 中抵消 |
|
||||
| 2026-07-30 | Round 07 随机分叉与复现同时过闸 | R0/R1 63/64 同格 trajectory 分叉;新进程 R0 的八项预注册字段 64/64 exact,五段 artifact hash chain 闭合 |
|
||||
|
||||
## 未决问题
|
||||
|
||||
|
||||
@@ -19,7 +19,7 @@
|
||||
|
||||
当前里程碑包含 17 专题学习地图、486 篇关键论文索引、Kimi K3 完整导读,
|
||||
语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 技术谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、工具使用与长程 Agent、原生多模态、训练系统、推理服务、数值优化,以及评测与安全深度专题,
|
||||
以及 87 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、
|
||||
以及 88 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、
|
||||
8 个交互实验和 100 个一手/官方节点,完整覆盖架构、预训练、后训练、系统、评测、案例与附录。
|
||||
第三轮已完成开放工件与首个真实 kernel 里程碑:固定官方模型与 FlashKDA revisions,审计 96 个 checkpoint shards、
|
||||
497,220 个 tensor entries、真实 KDA / MLA / MoE / MoonViT shapes 与小范围参数统计,并用 4 个新视图
|
||||
@@ -27,8 +27,8 @@
|
||||
`sm_120a` wheel,在 RTX 5090 上完成 6/6 官方参考 exact-match 和 K3 fixed / varlen 形状计时。详见
|
||||
[K3_ARTIFACT_AUDIT.md](./research/K3_ARTIFACT_AUDIT.md) 与
|
||||
[checkpoint_probe.py](./experiments/k3/checkpoint_probe.py)、[FlashKDA probe](./experiments/k3/flashkda/)。
|
||||
DeepSeek 六轮专题以 24 张问题账、10 次技术转向、
|
||||
20 个交互实验和 60 个一手/官方节点,串起 Dense、MoE、MLA、V3 协同、R1 与 V4;
|
||||
DeepSeek 七轮专题以 24 张问题账、10 次技术转向、
|
||||
21 个交互实验和 60 个一手/官方节点,串起 Dense、MoE、MLA、V3 协同、R1 与 V4;
|
||||
并固定官方 V2-Lite revision,在 RTX 5090 上连续执行 7/27 层,记录 3,240 次真实专家选择、
|
||||
MLA/HF eager cache shapes 与 `31/31` exact 独立复跑;进一步用真实 layer-1 权重执行官方 V3
|
||||
naive/absorb 路径,实际写入 576 元素 latent cache,并以 FP32 将两种结合顺序的最大误差压到
|
||||
@@ -78,6 +78,13 @@ Round 06 冻结 4 条 source、8 个 SHA-256 派生 seed 与 8 个 condition,
|
||||
为 `62 / 64`,单条 HumanEval 的官方 tests pass 为 `63 / 64`,不外推为 benchmark。
|
||||
全新进程复跑 R0/R1 后,run seed、prompt hash、完整 token IDs、文本、停止状态与
|
||||
CPU/CUDA RNG pre-state 八项均为 `64 / 64` exact。
|
||||
Round 07 保持 256 条输出预算不变,改为 16 条预先冻结的 source、4 个 seed 与
|
||||
`system off/on × EOS/句点` 四格;source 是覆盖单位,seed 只是题内重复。正式结果为
|
||||
250 / 256 natural EOS、247 个完整 token trajectory hashes,四道 GSM8K 分别
|
||||
`14/16 · 16/16 · 8/16 · 9/16`,四道 HumanEval 分别
|
||||
`16/16 · 8/16 · 12/16 · 16/16`。句点在 11 / 16 条 source 上缩短平均长度,但
|
||||
English 的首条 source 为 `−121`、其余三条为正,证明单题均值会与多数 source
|
||||
方向相反;新进程 R0 的八项合同字段仍为 `64 / 64` exact。
|
||||
详见
|
||||
[DEEPSEEK_V2_LITE_TRACE.md](./research/DEEPSEEK_V2_LITE_TRACE.md) 与
|
||||
[DEEPSEEK_MLA_ABSORB_AUDIT.md](./research/DEEPSEEK_MLA_ABSORB_AUDIT.md)、
|
||||
@@ -93,7 +100,9 @@ CPU/CUDA RNG pre-state 八项均为 `64 / 64` exact。
|
||||
[DEEPSEEK_V2_LITE_CHAT_BEHAVIOR_AUDIT.md](./research/DEEPSEEK_V2_LITE_CHAT_BEHAVIOR_AUDIT.md) 与
|
||||
[DEEPSEEK_V2_LITE_CHAT_COMPLETION_DEPTH_AUDIT.md](./research/DEEPSEEK_V2_LITE_CHAT_COMPLETION_DEPTH_AUDIT.md)、
|
||||
[DEEPSEEK_V2_LITE_CHAT_SAMPLING_PROTOCOL.md](./research/DEEPSEEK_V2_LITE_CHAT_SAMPLING_PROTOCOL.md) 与
|
||||
[DEEPSEEK_V2_LITE_CHAT_SAMPLING_AUDIT.md](./research/DEEPSEEK_V2_LITE_CHAT_SAMPLING_AUDIT.md)。
|
||||
[DEEPSEEK_V2_LITE_CHAT_SAMPLING_AUDIT.md](./research/DEEPSEEK_V2_LITE_CHAT_SAMPLING_AUDIT.md),以及
|
||||
[DEEPSEEK_V2_LITE_CHAT_CROSS_SOURCE_SAMPLING_PROTOCOL.md](./research/DEEPSEEK_V2_LITE_CHAT_CROSS_SOURCE_SAMPLING_PROTOCOL.md) 与
|
||||
[DEEPSEEK_V2_LITE_CHAT_CROSS_SOURCE_SAMPLING_AUDIT.md](./research/DEEPSEEK_V2_LITE_CHAT_CROSS_SOURCE_SAMPLING_AUDIT.md)。
|
||||
其余专题按进度账本持续扩建。
|
||||
|
||||
## 本地开发
|
||||
|
||||
@@ -754,3 +754,102 @@ See `research/DEEPSEEK_V2_LITE_CHAT_SAMPLING_PROTOCOL.md` and
|
||||
`research/DEEPSEEK_V2_LITE_CHAT_SAMPLING_AUDIT.md` for seed derivation,
|
||||
smoke gates, source-condition diversity tables, edge-set comparisons, task
|
||||
failure cases, exact reproduction scope, primary sources, and non-claims.
|
||||
|
||||
## Source-blocked cross-source Chat sampling
|
||||
|
||||
`v2_lite_chat_cross_source_sampling_probe.py` reuses the audited sampling
|
||||
runtime but changes the coverage contract before any outputs are inspected:
|
||||
|
||||
```text
|
||||
16 preregistered sources
|
||||
× 4 SHA-256-derived seeds
|
||||
× 4 conditions (system off/on × EOS/period)
|
||||
= 256 outputs
|
||||
```
|
||||
|
||||
The four first-ranked sources in each of WikiText-2, TNEWS, HumanEval, and
|
||||
GSM8K come from the frozen routing-corpus selection contract. Source is the
|
||||
primary coverage unit; seed is a within-source repeat. The period condition is
|
||||
a directed Round-06 follow-up, not a blind independent confirmation.
|
||||
|
||||
```bash
|
||||
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
|
||||
PYTHONPATH=/path/to/transformers-4.41.2-deps \
|
||||
python -B \
|
||||
experiments/deepseek/v2_lite_chat_cross_source_sampling_probe.py \
|
||||
--artifact-dir /path/to/deepseek-v2-lite-chat \
|
||||
--reference-routing-json \
|
||||
src/data/deepseek-v2-lite-routing-special-token-family-control.json \
|
||||
--greedy-json src/data/deepseek-v2-lite-chat-completion-512.json \
|
||||
--human-eval /path/to/HumanEval.jsonl.gz \
|
||||
--gsm8k /path/to/gsm8k/test.jsonl \
|
||||
--tnews /path/to/tnews/test.json \
|
||||
--tnews-archive /path/to/tnews_public.zip \
|
||||
--wikitext /path/to/wikitext-validation.parquet \
|
||||
--base-seeds 2101325316 2511573438 1677220094 2412346607 \
|
||||
--max-new-tokens 512 \
|
||||
--gpu-memory 28GiB \
|
||||
--cpu-memory 80GiB \
|
||||
--output \
|
||||
src/data/deepseek-v2-lite-chat-cross-source-sampling.json
|
||||
```
|
||||
|
||||
The independent evaluator uses the same pinned, networkless, read-only
|
||||
HumanEval sandbox and avoids pooling final-answer frequencies across different
|
||||
GSM8K tasks:
|
||||
|
||||
```bash
|
||||
python -B \
|
||||
experiments/deepseek/v2_lite_chat_cross_source_sampling_evaluator.py \
|
||||
--sampling-json \
|
||||
src/data/deepseek-v2-lite-chat-cross-source-sampling.json \
|
||||
--human-eval /path/to/HumanEval.jsonl.gz \
|
||||
--gsm8k /path/to/gsm8k/test.jsonl \
|
||||
--sandbox-image \
|
||||
python:3.11-alpine@sha256:25976e9d34a0fab1f278cae931f34c8303d97bf0c0d7f85b6b4dcf641d7702a4 \
|
||||
--output \
|
||||
src/data/deepseek-v2-lite-chat-cross-source-sampling-eval.json
|
||||
```
|
||||
|
||||
The source-blocked analysis computes each contrast within source, then reports
|
||||
the four source directions in each domain. It deliberately emits no p-values
|
||||
or population confidence intervals:
|
||||
|
||||
```bash
|
||||
python -B \
|
||||
experiments/deepseek/v2_lite_chat_cross_source_sampling_analysis.py \
|
||||
--sampling-json \
|
||||
src/data/deepseek-v2-lite-chat-cross-source-sampling.json \
|
||||
--evaluation-json \
|
||||
src/data/deepseek-v2-lite-chat-cross-source-sampling-eval.json \
|
||||
--reproduction-json \
|
||||
src/data/deepseek-v2-lite-chat-cross-source-sampling-reproduction.json \
|
||||
--output \
|
||||
src/data/deepseek-v2-lite-chat-cross-source-sampling-analysis.json
|
||||
```
|
||||
|
||||
A fresh Python process reruns R0 for all 16 sources. The existing reproduction
|
||||
comparer verifies the eight preregistered fields in all 64
|
||||
source-condition cells.
|
||||
|
||||
Formal results are 250/256 natural EOS and 247/256 unique complete token
|
||||
trajectories. Strict task totals are Math 47/64 and Code 52/64, but the
|
||||
per-task totals range from 8/16 to 16/16 in both domains. Period shortens mean
|
||||
generation length for 11/16 sources. English is the important counterexample:
|
||||
the first source is −121 tokens while the other three are +44.5, +41, and
|
||||
+1.375, so the domain mean is negative even though three of four sources are
|
||||
positive.
|
||||
|
||||
```text
|
||||
formal f013132485f27adce008f03f781bed9982efc0d7939f13faede01c9f6f3d7f7c
|
||||
eval e88b274599fc5951561f9e5e7438fb4d6d25f341ae3bc4a2c4121877a0db8975
|
||||
rerun 143dc9d0f7c914db4781e36b1401cdc9fbc2971a0dca71188bab0f8cedb002a6
|
||||
compare ec4a47894953f5d73bb62211b588632ef7032da0e915218992b7fc3ef9c2a556
|
||||
analysis d0dece388998fee419d34ff33f140695a9fedef6e79799cdf42eb283b047bc84
|
||||
compact d0ab65646c6119bdeafeb451103dc6afebff3624a1e05f13a45a52ad965be1af
|
||||
```
|
||||
|
||||
See `research/DEEPSEEK_V2_LITE_CHAT_CROSS_SOURCE_SAMPLING_PROTOCOL.md` and
|
||||
`research/DEEPSEEK_V2_LITE_CHAT_CROSS_SOURCE_SAMPLING_AUDIT.md` for the
|
||||
preregistered source frame, task matrices, source-direction contrasts,
|
||||
failure identities, hash chain, exact replay scope, and non-claims.
|
||||
|
||||
@@ -0,0 +1,678 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Build source-blocked summaries for the Round 07 sampling grid.
|
||||
|
||||
The primary unit is the preregistered source. Seed-level outputs remain
|
||||
within-source repeats and are never counted as independent benchmark tasks.
|
||||
No p-values or population confidence intervals are produced from four sources
|
||||
per domain.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import hashlib
|
||||
import json
|
||||
from collections import Counter
|
||||
from pathlib import Path
|
||||
from statistics import mean, median
|
||||
from typing import Any, Callable
|
||||
|
||||
|
||||
PROTOCOL_ID = "llm-atlas-deepseek-chat-cross-source-sampling-v1"
|
||||
CONDITIONS = (
|
||||
"s0_eos",
|
||||
"s1_eos",
|
||||
"s0_period",
|
||||
"s1_period",
|
||||
)
|
||||
DOMAINS = ("english", "chinese", "code", "math")
|
||||
TASK_DOMAINS = ("code", "math")
|
||||
METRICS = (
|
||||
"natural_eos_rate",
|
||||
"mean_generated_tokens",
|
||||
"fixed_budget_success_rate",
|
||||
"strict_complete_success_rate",
|
||||
)
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--sampling-json", type=Path, required=True)
|
||||
parser.add_argument("--evaluation-json", type=Path, required=True)
|
||||
parser.add_argument("--reproduction-json", type=Path)
|
||||
parser.add_argument("--output", type=Path, required=True)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def sha256_file(path: Path) -> str:
|
||||
digest = hashlib.sha256()
|
||||
with path.open("rb") as handle:
|
||||
while chunk := handle.read(16 * 1024 * 1024):
|
||||
digest.update(chunk)
|
||||
return digest.hexdigest()
|
||||
|
||||
|
||||
def canonical_hash(value: Any) -> str:
|
||||
return hashlib.sha256(
|
||||
json.dumps(
|
||||
value,
|
||||
ensure_ascii=False,
|
||||
sort_keys=True,
|
||||
separators=(",", ":"),
|
||||
).encode()
|
||||
).hexdigest()
|
||||
|
||||
|
||||
def summarize(values: list[float | int]) -> dict[str, Any]:
|
||||
return {
|
||||
"count": len(values),
|
||||
"mean": mean(values) if values else None,
|
||||
"median": median(values) if values else None,
|
||||
"min": min(values) if values else None,
|
||||
"max": max(values) if values else None,
|
||||
}
|
||||
|
||||
|
||||
def direction_counts(values: list[float]) -> dict[str, int]:
|
||||
epsilon = 1e-12
|
||||
return {
|
||||
"positive": sum(value > epsilon for value in values),
|
||||
"zero": sum(abs(value) <= epsilon for value in values),
|
||||
"negative": sum(value < -epsilon for value in values),
|
||||
}
|
||||
|
||||
|
||||
def row_success(row: dict[str, Any], strict: bool) -> int | None:
|
||||
evaluation = row["task_evaluation"]
|
||||
if row["domain"] == "math":
|
||||
key = (
|
||||
"strict_complete_numeric_exact"
|
||||
if strict
|
||||
else "fixed_budget_numeric_exact"
|
||||
)
|
||||
return int(evaluation[key])
|
||||
if row["domain"] == "code":
|
||||
key = (
|
||||
"strict_complete_tests_pass"
|
||||
if strict
|
||||
else "fixed_budget_tests_pass"
|
||||
)
|
||||
return int(evaluation[key])
|
||||
return None
|
||||
|
||||
|
||||
def cell_summary(rows: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
if not rows:
|
||||
raise ValueError("cell has no rows")
|
||||
fixed = [
|
||||
value
|
||||
for row in rows
|
||||
if (value := row_success(row, strict=False)) is not None
|
||||
]
|
||||
strict = [
|
||||
value
|
||||
for row in rows
|
||||
if (value := row_success(row, strict=True)) is not None
|
||||
]
|
||||
lengths = [row["generated_tokens"] for row in rows]
|
||||
hashes = {
|
||||
row["generated_token_ids_sha256"] for row in rows
|
||||
}
|
||||
result = {
|
||||
"outputs": len(rows),
|
||||
"natural_eos": sum(row["hit_eos"] for row in rows),
|
||||
"natural_eos_rate": mean(
|
||||
int(row["hit_eos"]) for row in rows
|
||||
),
|
||||
"budget_truncated": sum(
|
||||
row["stopped_at_max_new_tokens"] for row in rows
|
||||
),
|
||||
"generated_tokens": summarize(lengths),
|
||||
"generated_token_seed_range": max(lengths) - min(lengths),
|
||||
"unique_trajectory_hashes": len(hashes),
|
||||
"completion_classes": dict(
|
||||
Counter(row["completion_class"] for row in rows)
|
||||
),
|
||||
"fixed_budget_success": sum(fixed) if fixed else None,
|
||||
"fixed_budget_success_rate": mean(fixed) if fixed else None,
|
||||
"strict_complete_success": sum(strict) if strict else None,
|
||||
"strict_complete_success_rate": (
|
||||
mean(strict) if strict else None
|
||||
),
|
||||
"observed_any_fixed_budget_pass": (
|
||||
any(fixed) if fixed else None
|
||||
),
|
||||
"observed_any_strict_complete_pass": (
|
||||
any(strict) if strict else None
|
||||
),
|
||||
"seed_success_range": (
|
||||
max(fixed) - min(fixed) if fixed else None
|
||||
),
|
||||
}
|
||||
if rows[0]["domain"] == "math":
|
||||
answers = Counter(
|
||||
row["task_evaluation"]["predicted_final"]
|
||||
for row in rows
|
||||
if row["task_evaluation"]["predicted_final"] is not None
|
||||
)
|
||||
largest = max(answers.values(), default=0)
|
||||
modes = sorted(
|
||||
answer
|
||||
for answer, count in answers.items()
|
||||
if count == largest
|
||||
)
|
||||
result["math_answers"] = {
|
||||
"frequencies": dict(answers),
|
||||
"modes": modes,
|
||||
"unique_absolute_majority": (
|
||||
modes[0]
|
||||
if len(modes) == 1 and largest > len(rows) / 2
|
||||
else None
|
||||
),
|
||||
"gold": rows[0]["task_evaluation"]["gold_final"],
|
||||
}
|
||||
if rows[0]["domain"] == "code":
|
||||
result["code"] = {
|
||||
"ast_parse": sum(
|
||||
row["task_evaluation"]["python_ast_parse"]
|
||||
for row in rows
|
||||
),
|
||||
"executed": sum(
|
||||
row["task_evaluation"]["execution"]["status"]
|
||||
!= "not_run"
|
||||
for row in rows
|
||||
),
|
||||
"execution_statuses": dict(
|
||||
Counter(
|
||||
row["task_evaluation"]["execution"]["status"]
|
||||
for row in rows
|
||||
)
|
||||
),
|
||||
}
|
||||
return result
|
||||
|
||||
|
||||
def contrast(
|
||||
cells: dict[str, dict[str, Any]],
|
||||
metric: str,
|
||||
) -> dict[str, float] | None:
|
||||
def metric_value(cell: dict[str, Any]) -> float | None:
|
||||
if metric == "mean_generated_tokens":
|
||||
return cell["generated_tokens"]["mean"]
|
||||
return cell[metric]
|
||||
|
||||
values = {
|
||||
condition: metric_value(cells[condition])
|
||||
for condition in CONDITIONS
|
||||
}
|
||||
if any(value is None for value in values.values()):
|
||||
return None
|
||||
system = (
|
||||
(values["s1_eos"] + values["s1_period"]) / 2
|
||||
- (values["s0_eos"] + values["s0_period"]) / 2
|
||||
)
|
||||
boundary = (
|
||||
(values["s0_period"] + values["s1_period"]) / 2
|
||||
- (values["s0_eos"] + values["s1_eos"]) / 2
|
||||
)
|
||||
interaction = (
|
||||
values["s1_period"]
|
||||
- values["s1_eos"]
|
||||
- values["s0_period"]
|
||||
+ values["s0_eos"]
|
||||
)
|
||||
return {
|
||||
"system_main": system,
|
||||
"boundary_main_period_minus_eos": boundary,
|
||||
"interaction": interaction,
|
||||
"period_minus_eos_s0": (
|
||||
values["s0_period"] - values["s0_eos"]
|
||||
),
|
||||
"period_minus_eos_s1": (
|
||||
values["s1_period"] - values["s1_eos"]
|
||||
),
|
||||
"system_on_minus_off_eos": (
|
||||
values["s1_eos"] - values["s0_eos"]
|
||||
),
|
||||
"system_on_minus_off_period": (
|
||||
values["s1_period"] - values["s0_period"]
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def domain_condition_summary(
|
||||
*,
|
||||
domain: str,
|
||||
condition: str,
|
||||
source_ids: list[str],
|
||||
source_cells: dict[str, dict[str, dict[str, Any]]],
|
||||
rows: list[dict[str, Any]],
|
||||
) -> dict[str, Any]:
|
||||
subset = [
|
||||
row
|
||||
for row in rows
|
||||
if row["domain"] == domain
|
||||
and row["condition"] == condition
|
||||
]
|
||||
cells = [source_cells[source_id][condition] for source_id in source_ids]
|
||||
fixed_rates = [
|
||||
cell["fixed_budget_success_rate"]
|
||||
for cell in cells
|
||||
if cell["fixed_budget_success_rate"] is not None
|
||||
]
|
||||
strict_rates = [
|
||||
cell["strict_complete_success_rate"]
|
||||
for cell in cells
|
||||
if cell["strict_complete_success_rate"] is not None
|
||||
]
|
||||
source_mean_lengths = [
|
||||
cell["generated_tokens"]["mean"] for cell in cells
|
||||
]
|
||||
return {
|
||||
"sources": len(source_ids),
|
||||
"outputs": len(subset),
|
||||
"micro": {
|
||||
"natural_eos": sum(row["hit_eos"] for row in subset),
|
||||
"natural_eos_rate": mean(
|
||||
int(row["hit_eos"]) for row in subset
|
||||
),
|
||||
"budget_truncated": sum(
|
||||
row["stopped_at_max_new_tokens"]
|
||||
for row in subset
|
||||
),
|
||||
"mean_generated_tokens": mean(
|
||||
row["generated_tokens"] for row in subset
|
||||
),
|
||||
"fixed_budget_success": (
|
||||
sum(
|
||||
value
|
||||
for row in subset
|
||||
if (
|
||||
value := row_success(row, strict=False)
|
||||
)
|
||||
is not None
|
||||
)
|
||||
if fixed_rates
|
||||
else None
|
||||
),
|
||||
"strict_complete_success": (
|
||||
sum(
|
||||
value
|
||||
for row in subset
|
||||
if (
|
||||
value := row_success(row, strict=True)
|
||||
)
|
||||
is not None
|
||||
)
|
||||
if strict_rates
|
||||
else None
|
||||
),
|
||||
},
|
||||
"source_macro": {
|
||||
"natural_eos_rate": summarize(
|
||||
[cell["natural_eos_rate"] for cell in cells]
|
||||
),
|
||||
"mean_generated_tokens": summarize(source_mean_lengths),
|
||||
"fixed_budget_success_rate": summarize(fixed_rates),
|
||||
"strict_complete_success_rate": summarize(strict_rates),
|
||||
},
|
||||
"variability": {
|
||||
"within_source_seed_length_range": summarize(
|
||||
[
|
||||
cell["generated_token_seed_range"]
|
||||
for cell in cells
|
||||
]
|
||||
),
|
||||
"between_source_mean_length_range": (
|
||||
max(source_mean_lengths) - min(source_mean_lengths)
|
||||
),
|
||||
"within_source_seed_success_range": summarize(
|
||||
[
|
||||
cell["seed_success_range"]
|
||||
for cell in cells
|
||||
if cell["seed_success_range"] is not None
|
||||
]
|
||||
),
|
||||
"between_source_success_rate_range": (
|
||||
max(fixed_rates) - min(fixed_rates)
|
||||
if fixed_rates
|
||||
else None
|
||||
),
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def build_source_cells(
|
||||
rows: list[dict[str, Any]],
|
||||
source_order: list[str],
|
||||
) -> dict[str, dict[str, dict[str, Any]]]:
|
||||
result: dict[str, dict[str, dict[str, Any]]] = {}
|
||||
for source_id in source_order:
|
||||
result[source_id] = {}
|
||||
for condition in CONDITIONS:
|
||||
subset = [
|
||||
row
|
||||
for row in rows
|
||||
if row["source_id"] == source_id
|
||||
and row["condition"] == condition
|
||||
]
|
||||
if len(subset) != 4:
|
||||
raise RuntimeError(
|
||||
f"{source_id}/{condition} has {len(subset)} rows; "
|
||||
"expected 4"
|
||||
)
|
||||
result[source_id][condition] = cell_summary(subset)
|
||||
return result
|
||||
|
||||
|
||||
def load_json(path: Path) -> dict[str, Any]:
|
||||
if not path.is_file():
|
||||
raise FileNotFoundError(path)
|
||||
return json.loads(path.read_text(encoding="utf-8"))
|
||||
|
||||
|
||||
def main() -> None:
|
||||
args = parse_args()
|
||||
sampling = load_json(args.sampling_json)
|
||||
evaluation = load_json(args.evaluation_json)
|
||||
if sampling["protocol_id"] != PROTOCOL_ID:
|
||||
raise RuntimeError("sampling protocol ID differs")
|
||||
if evaluation["protocol_id"] != PROTOCOL_ID:
|
||||
raise RuntimeError("evaluation protocol ID differs")
|
||||
if evaluation["input"]["sampling_sha256"] != sha256_file(
|
||||
args.sampling_json
|
||||
):
|
||||
raise RuntimeError("evaluation does not reference this sampling file")
|
||||
if tuple(
|
||||
sampling["seed_contract"]["condition_row_order"]
|
||||
) != CONDITIONS:
|
||||
raise RuntimeError("condition order differs from preregistration")
|
||||
if len(
|
||||
sampling["seed_contract"]["executed_base_seeds"]
|
||||
) != 4:
|
||||
raise RuntimeError("formal grid must contain four seeds")
|
||||
|
||||
source_metadata = {
|
||||
source["id"]: {
|
||||
"domain": source["domain"],
|
||||
"within_domain_index": source["within_domain_index"],
|
||||
"selection_rank": source["selection_rank"],
|
||||
}
|
||||
for source in sampling["sources"]
|
||||
}
|
||||
source_order = [source["id"] for source in sampling["sources"]]
|
||||
if len(source_order) != 16 or len(set(source_order)) != 16:
|
||||
raise RuntimeError("formal grid must contain 16 unique sources")
|
||||
by_domain = {
|
||||
domain: [
|
||||
source_id
|
||||
for source_id in source_order
|
||||
if source_metadata[source_id]["domain"] == domain
|
||||
]
|
||||
for domain in DOMAINS
|
||||
}
|
||||
if any(len(source_ids) != 4 for source_ids in by_domain.values()):
|
||||
raise RuntimeError("each domain must contain four sources")
|
||||
|
||||
rows = evaluation["rows"]
|
||||
if len(rows) != 256:
|
||||
raise RuntimeError(f"expected 256 evaluated rows; got {len(rows)}")
|
||||
grid_keys = {
|
||||
(
|
||||
row["source_id"],
|
||||
row["base_seed"],
|
||||
row["condition"],
|
||||
)
|
||||
for row in rows
|
||||
}
|
||||
if len(grid_keys) != 256:
|
||||
raise RuntimeError("evaluation grid has duplicate or missing cells")
|
||||
|
||||
source_cells = build_source_cells(rows, source_order)
|
||||
source_contrasts = {
|
||||
source_id: {
|
||||
metric: contrast(source_cells[source_id], metric)
|
||||
for metric in METRICS
|
||||
}
|
||||
for source_id in source_order
|
||||
}
|
||||
domain_conditions = {
|
||||
domain: {
|
||||
condition: domain_condition_summary(
|
||||
domain=domain,
|
||||
condition=condition,
|
||||
source_ids=by_domain[domain],
|
||||
source_cells=source_cells,
|
||||
rows=rows,
|
||||
)
|
||||
for condition in CONDITIONS
|
||||
}
|
||||
for domain in DOMAINS
|
||||
}
|
||||
|
||||
domain_contrasts: dict[str, Any] = {}
|
||||
for domain in DOMAINS:
|
||||
domain_contrasts[domain] = {}
|
||||
for metric in METRICS:
|
||||
available = {
|
||||
source_id: source_contrasts[source_id][metric]
|
||||
for source_id in by_domain[domain]
|
||||
if source_contrasts[source_id][metric] is not None
|
||||
}
|
||||
domain_contrasts[domain][metric] = {}
|
||||
for contrast_name in (
|
||||
"system_main",
|
||||
"boundary_main_period_minus_eos",
|
||||
"interaction",
|
||||
"period_minus_eos_s0",
|
||||
"period_minus_eos_s1",
|
||||
"system_on_minus_off_eos",
|
||||
"system_on_minus_off_period",
|
||||
):
|
||||
values = [
|
||||
value[contrast_name]
|
||||
for value in available.values()
|
||||
]
|
||||
domain_contrasts[domain][metric][contrast_name] = {
|
||||
**summarize(values),
|
||||
"directions": direction_counts(values),
|
||||
"by_source": {
|
||||
source_id: value[contrast_name]
|
||||
for source_id, value in available.items()
|
||||
},
|
||||
}
|
||||
|
||||
task_matrix = {
|
||||
domain: [
|
||||
{
|
||||
"source_id": source_id,
|
||||
"within_domain_index": source_metadata[source_id][
|
||||
"within_domain_index"
|
||||
],
|
||||
"conditions": {
|
||||
condition: {
|
||||
key: source_cells[source_id][condition][key]
|
||||
for key in (
|
||||
"fixed_budget_success",
|
||||
"strict_complete_success",
|
||||
"observed_any_fixed_budget_pass",
|
||||
"observed_any_strict_complete_pass",
|
||||
"natural_eos",
|
||||
"generated_tokens",
|
||||
)
|
||||
}
|
||||
for condition in CONDITIONS
|
||||
},
|
||||
}
|
||||
for source_id in by_domain[domain]
|
||||
]
|
||||
for domain in TASK_DOMAINS
|
||||
}
|
||||
|
||||
reproduction = None
|
||||
if args.reproduction_json is not None:
|
||||
reproduction_payload = load_json(args.reproduction_json)
|
||||
if reproduction_payload["protocol_id"] != PROTOCOL_ID:
|
||||
raise RuntimeError("reproduction protocol ID differs")
|
||||
if reproduction_payload["formal"]["sha256"] != sha256_file(
|
||||
args.sampling_json
|
||||
):
|
||||
raise RuntimeError(
|
||||
"reproduction does not reference this formal file"
|
||||
)
|
||||
reproduction = {
|
||||
"path": str(args.reproduction_json),
|
||||
"sha256": sha256_file(args.reproduction_json),
|
||||
"summary": reproduction_payload["summary"],
|
||||
}
|
||||
|
||||
result = {
|
||||
"schema_version": 1,
|
||||
"protocol_id": PROTOCOL_ID,
|
||||
"inputs": {
|
||||
"sampling": {
|
||||
"path": str(args.sampling_json),
|
||||
"sha256": sha256_file(args.sampling_json),
|
||||
"content_hash": sampling["content_hash"],
|
||||
},
|
||||
"evaluation": {
|
||||
"path": str(args.evaluation_json),
|
||||
"sha256": sha256_file(args.evaluation_json),
|
||||
"content_hash": evaluation["content_hash"],
|
||||
},
|
||||
"reproduction": reproduction,
|
||||
},
|
||||
"contract": {
|
||||
"source_is_primary_coverage_unit": True,
|
||||
"seed_is_within_source_repeat": True,
|
||||
"sources": len(source_order),
|
||||
"sources_per_domain": 4,
|
||||
"seeds_per_source_condition": 4,
|
||||
"conditions": list(CONDITIONS),
|
||||
"outputs": len(rows),
|
||||
"source_order": source_order,
|
||||
"sources_by_domain": by_domain,
|
||||
"no_population_confidence_intervals": True,
|
||||
"no_p_values": True,
|
||||
},
|
||||
"overall": {
|
||||
"outputs": len(rows),
|
||||
"natural_eos": sum(row["hit_eos"] for row in rows),
|
||||
"budget_truncated": sum(
|
||||
row["stopped_at_max_new_tokens"] for row in rows
|
||||
),
|
||||
"unique_trajectory_hashes": len(
|
||||
{
|
||||
row["generated_token_ids_sha256"]
|
||||
for row in rows
|
||||
}
|
||||
),
|
||||
"math": {
|
||||
"outputs": sum(
|
||||
row["domain"] == "math" for row in rows
|
||||
),
|
||||
"fixed_budget_success": sum(
|
||||
row_success(row, strict=False) or 0
|
||||
for row in rows
|
||||
if row["domain"] == "math"
|
||||
),
|
||||
"strict_complete_success": sum(
|
||||
row_success(row, strict=True) or 0
|
||||
for row in rows
|
||||
if row["domain"] == "math"
|
||||
),
|
||||
},
|
||||
"code": {
|
||||
"outputs": sum(
|
||||
row["domain"] == "code" for row in rows
|
||||
),
|
||||
"fixed_budget_success": sum(
|
||||
row_success(row, strict=False) or 0
|
||||
for row in rows
|
||||
if row["domain"] == "code"
|
||||
),
|
||||
"strict_complete_success": sum(
|
||||
row_success(row, strict=True) or 0
|
||||
for row in rows
|
||||
if row["domain"] == "code"
|
||||
),
|
||||
},
|
||||
},
|
||||
"source_metadata": source_metadata,
|
||||
"source_condition_cells": source_cells,
|
||||
"domain_condition_summary": domain_conditions,
|
||||
"source_contrasts": source_contrasts,
|
||||
"domain_contrasts": domain_contrasts,
|
||||
"task_matrix": task_matrix,
|
||||
"prior_direction_check": {
|
||||
domain: {
|
||||
"period_shortens_mean_tokens": {
|
||||
"by_source": {
|
||||
source_id: (
|
||||
source_contrasts[source_id][
|
||||
"mean_generated_tokens"
|
||||
][
|
||||
"boundary_main_period_minus_eos"
|
||||
]
|
||||
< 0
|
||||
)
|
||||
for source_id in by_domain[domain]
|
||||
},
|
||||
},
|
||||
"system_on_raises_natural_eos_rate": {
|
||||
"by_source": {
|
||||
source_id: (
|
||||
source_contrasts[source_id][
|
||||
"natural_eos_rate"
|
||||
]["system_main"]
|
||||
> 0
|
||||
)
|
||||
for source_id in by_domain[domain]
|
||||
},
|
||||
},
|
||||
}
|
||||
for domain in DOMAINS
|
||||
},
|
||||
"claim_boundary": [
|
||||
"Four sources per domain are not full benchmark estimates.",
|
||||
"Seed-level repeats are not independent task observations.",
|
||||
"Source-blocked contrasts are descriptive and have no p-values.",
|
||||
"Observed any-pass is not standard HumanEval pass@4.",
|
||||
"Natural EOS, evaluator coverage, and correctness remain separate.",
|
||||
"The period condition is a counterfactual, not an official-valid chat.",
|
||||
"Round 06 motivated the period contrast, so this is a directed follow-up.",
|
||||
],
|
||||
}
|
||||
result["content_hash"] = canonical_hash(
|
||||
{
|
||||
"protocol_id": result["protocol_id"],
|
||||
"contract": result["contract"],
|
||||
"source_condition_cells": result[
|
||||
"source_condition_cells"
|
||||
],
|
||||
"source_contrasts": result["source_contrasts"],
|
||||
"task_matrix": result["task_matrix"],
|
||||
}
|
||||
)
|
||||
args.output.parent.mkdir(parents=True, exist_ok=True)
|
||||
args.output.write_text(
|
||||
json.dumps(result, ensure_ascii=False, indent=2) + "\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
print(
|
||||
json.dumps(
|
||||
{
|
||||
"output": str(args.output),
|
||||
"sha256": sha256_file(args.output),
|
||||
"bytes": args.output.stat().st_size,
|
||||
"overall": result["overall"],
|
||||
"reproduction": reproduction,
|
||||
},
|
||||
ensure_ascii=False,
|
||||
indent=2,
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,23 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Evaluate the preregistered cross-source sampling grid.
|
||||
|
||||
This keeps the Round 06 sandbox and four-ledger evaluator unchanged while
|
||||
restricting edge summaries to the preregistered EOS/period factorial.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import v2_lite_chat_sampling_evaluator as evaluator
|
||||
|
||||
|
||||
COMPARISONS = (
|
||||
("system_eos", "s0_eos", "s1_eos"),
|
||||
("system_period", "s0_period", "s1_period"),
|
||||
("period_at_s0", "s0_eos", "s0_period"),
|
||||
("period_at_s1", "s1_eos", "s1_period"),
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
evaluator.COMPARISONS = COMPARISONS
|
||||
evaluator.main()
|
||||
@@ -0,0 +1,63 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Run the preregistered cross-source DeepSeek Chat sampling grid.
|
||||
|
||||
The implementation deliberately reuses the audited Round 06 runner while
|
||||
installing a new immutable protocol identity, four SHA-256-derived seeds, and
|
||||
the fixed system x EOS/period four-row batch. Keeping the execution path
|
||||
shared avoids silently forking model loading, RNG capture, generation, and
|
||||
trajectory summaries.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import v2_lite_chat_sampling_probe as sampling
|
||||
|
||||
|
||||
PROTOCOL_ID = "llm-atlas-deepseek-chat-cross-source-sampling-v1"
|
||||
CONDITIONS = (
|
||||
"s0_eos",
|
||||
"s1_eos",
|
||||
"s0_period",
|
||||
"s1_period",
|
||||
)
|
||||
BASE_SEEDS = (
|
||||
2101325316,
|
||||
2511573438,
|
||||
1677220094,
|
||||
2412346607,
|
||||
)
|
||||
COMPARISONS = (
|
||||
("system_eos", "s0_eos", "s1_eos"),
|
||||
("system_period", "s0_period", "s1_period"),
|
||||
("period_at_s0", "s0_eos", "s0_period"),
|
||||
("period_at_s1", "s1_eos", "s1_period"),
|
||||
)
|
||||
|
||||
|
||||
def install_protocol() -> None:
|
||||
special = sampling.special
|
||||
factors = {
|
||||
condition: special.FACTORS[condition]
|
||||
for condition in CONDITIONS
|
||||
}
|
||||
special.BOUNDARY_LEVELS = ("eos", "period")
|
||||
special.CONDITIONS = CONDITIONS
|
||||
special.FACTORS = factors
|
||||
special.SYSTEM_CELLS = {
|
||||
"eos": ("s0_eos", "s1_eos"),
|
||||
"period": ("s0_period", "s1_period"),
|
||||
}
|
||||
special.SYSTEM_EDGE_CONTRASTS = {
|
||||
"period_minus_eos": ("eos", "period"),
|
||||
}
|
||||
special.COMPARISONS = COMPARISONS
|
||||
special.ALIGNMENT_COMPARISONS = COMPARISONS
|
||||
|
||||
sampling.PROTOCOL_ID = PROTOCOL_ID
|
||||
sampling.PREREGISTERED_BASE_SEEDS = BASE_SEEDS
|
||||
sampling.EXPECTED_CONDITIONS = CONDITIONS
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
install_protocol()
|
||||
sampling.main()
|
||||
@@ -153,11 +153,19 @@ def summarize_group(rows: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
code_rows = [
|
||||
row for row in rows if row["domain"] == "code"
|
||||
]
|
||||
answers = Counter(
|
||||
math_source_ids = sorted(
|
||||
{row["source_id"] for row in math_rows}
|
||||
)
|
||||
single_math_source = len(math_source_ids) == 1
|
||||
answers = (
|
||||
Counter(
|
||||
row["task_evaluation"]["predicted_final"]
|
||||
for row in math_rows
|
||||
if row["task_evaluation"]["predicted_final"] is not None
|
||||
)
|
||||
if single_math_source
|
||||
else Counter()
|
||||
)
|
||||
max_count = max(answers.values(), default=0)
|
||||
modes = sorted(
|
||||
answer
|
||||
@@ -166,7 +174,7 @@ def summarize_group(rows: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
)
|
||||
gold = (
|
||||
math_rows[0]["task_evaluation"]["gold_final"]
|
||||
if math_rows
|
||||
if single_math_source
|
||||
else None
|
||||
)
|
||||
return {
|
||||
@@ -185,6 +193,13 @@ def summarize_group(rows: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
),
|
||||
"math": {
|
||||
"outputs": len(math_rows),
|
||||
"sources": len(math_source_ids),
|
||||
"source_ids": math_source_ids,
|
||||
"answer_aggregation_scope": (
|
||||
"single_source"
|
||||
if single_math_source
|
||||
else "disabled_across_distinct_gold_answers"
|
||||
),
|
||||
"evaluator_covered": sum(
|
||||
row["task_evaluation"]["evaluator_covered"]
|
||||
for row in math_rows
|
||||
@@ -206,13 +221,13 @@ def summarize_group(rows: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
"modal_count": max_count,
|
||||
"absolute_majority_exists": (
|
||||
max_count > len(math_rows) / 2
|
||||
if math_rows
|
||||
if single_math_source
|
||||
else None
|
||||
),
|
||||
"unique_absolute_majority": (
|
||||
modes[0]
|
||||
if (
|
||||
math_rows
|
||||
single_math_source
|
||||
and len(modes) == 1
|
||||
and max_count > len(math_rows) / 2
|
||||
)
|
||||
@@ -223,12 +238,15 @@ def summarize_group(rows: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
len(modes) == 1
|
||||
and max_count > len(math_rows) / 2
|
||||
and modes[0] == gold
|
||||
if math_rows
|
||||
if single_math_source
|
||||
else None
|
||||
),
|
||||
},
|
||||
"code": {
|
||||
"outputs": len(code_rows),
|
||||
"sources": len(
|
||||
{row["source_id"] for row in code_rows}
|
||||
),
|
||||
"ast_parse": sum(
|
||||
row["task_evaluation"]["python_ast_parse"]
|
||||
for row in code_rows
|
||||
@@ -516,7 +534,11 @@ def main() -> None:
|
||||
"by_source_edge": edge_summary(rows),
|
||||
},
|
||||
"claim_boundary": [
|
||||
"Four sources and eight seeds are not benchmark estimates.",
|
||||
(
|
||||
f"{len(sampling['sources'])} sources and "
|
||||
f"{len(sampling['seed_contract']['executed_base_seeds'])} "
|
||||
"seeds are not full benchmark estimates."
|
||||
),
|
||||
"Completion-conditioned metrics are selection-biased diagnostics.",
|
||||
"A passing HumanEval test is functional evidence, not code-safety evidence.",
|
||||
"A modal sampled math answer is not standard self-consistency.",
|
||||
|
||||
@@ -965,13 +965,17 @@ def main() -> None:
|
||||
"use_cache": True,
|
||||
"official_generation_config": official_generation,
|
||||
"official_parameters_explicitly_passed_to_generate": True,
|
||||
"all_eight_conditions_same_source_batch": True,
|
||||
"all_conditions_same_source_batch": True,
|
||||
"batch_conditions": len(EXPECTED_CONDITIONS),
|
||||
"all_eight_conditions_same_source_batch": (
|
||||
len(EXPECTED_CONDITIONS) == 8
|
||||
),
|
||||
"batch_padding": (
|
||||
"left padding with PAD=EOS and attention_mask=0; "
|
||||
"generation prefix ends at the same batch column"
|
||||
),
|
||||
"counterfactual_boundary": (
|
||||
"BOS/x/period cells edit one pre-target token ID after "
|
||||
"Non-EOS cells edit one pre-target token ID after "
|
||||
"official rendering and are not valid official chats"
|
||||
),
|
||||
},
|
||||
@@ -1018,12 +1022,22 @@ def main() -> None:
|
||||
"sources": generated_sources,
|
||||
"summary": summarize(generated_sources, greedy_rows),
|
||||
"claim_boundary": [
|
||||
"This is a four-source, eight-seed mechanism probe, not a benchmark.",
|
||||
"Eight trajectories do not estimate the full generation distribution.",
|
||||
(
|
||||
f"This is a {len(source_rows)}-source, "
|
||||
f"{len(args.base_seeds)}-seed mechanism probe, "
|
||||
"not a benchmark."
|
||||
),
|
||||
(
|
||||
f"{len(args.base_seeds)} trajectories per cell do not "
|
||||
"estimate the full generation distribution."
|
||||
),
|
||||
"Batch-seed-aligned rows are not common-random-number pairs.",
|
||||
"Counterfactual BOS/x/period token sequences are not official-valid chats.",
|
||||
"Counterfactual non-EOS token sequences are not official-valid chats.",
|
||||
"Unique token sequences are not semantic-diversity measurements.",
|
||||
"One GSM8K and one HumanEval source do not estimate task ability.",
|
||||
(
|
||||
f"{args.per_domain} GSM8K and {args.per_domain} HumanEval "
|
||||
"sources do not estimate full benchmark ability."
|
||||
),
|
||||
"CPU-offloaded eager latency is not serving throughput.",
|
||||
"Sampling differences do not identify a hidden-state or router mediator.",
|
||||
],
|
||||
|
||||
@@ -16,6 +16,7 @@
|
||||
"build:data:deepseek-chat-behavior": "node scripts/build-deepseek-chat-behavior-compact.mjs",
|
||||
"build:data:deepseek-chat-completion-depth": "node scripts/build-deepseek-chat-completion-depth-compact.mjs",
|
||||
"build:data:deepseek-chat-sampling": "node scripts/build-deepseek-chat-sampling-compact.mjs",
|
||||
"build:data:deepseek-chat-cross-source-sampling": "node scripts/build-deepseek-chat-cross-source-sampling-compact.mjs",
|
||||
"check:site": "node scripts/check-site.mjs",
|
||||
"check:moe-browser": "node scripts/check-moe-browser.mjs",
|
||||
"check:reasoning-browser": "node scripts/check-reasoning-browser.mjs",
|
||||
@@ -33,6 +34,7 @@
|
||||
"check:representation-browser": "node scripts/check-representation-browser.mjs",
|
||||
"check:deepseek-browser": "node scripts/check-deepseek-browser.mjs",
|
||||
"check:deepseek-sampling-browser": "node scripts/check-deepseek-sampling-browser.mjs",
|
||||
"check:deepseek-cross-source-sampling-browser": "node scripts/check-deepseek-cross-source-sampling-browser.mjs",
|
||||
"check:k3-browser": "node scripts/check-k3-browser.mjs"
|
||||
},
|
||||
"dependencies": {
|
||||
|
||||
@@ -0,0 +1,725 @@
|
||||
# DeepSeek-V2-Lite-Chat 跨来源采样审计
|
||||
|
||||
> 执行日期:2026-07-30
|
||||
>
|
||||
> 协议:`llm-atlas-deepseek-chat-cross-source-sampling-v1`
|
||||
>
|
||||
> 模型:`deepseek-ai/DeepSeek-V2-Lite-Chat`
|
||||
>
|
||||
> revision:`85864749cd611b4353ce1decdb286193298f64c7`
|
||||
>
|
||||
> 预注册:`research/DEEPSEEK_V2_LITE_CHAT_CROSS_SOURCE_SAMPLING_PROTOCOL.md`
|
||||
>
|
||||
> 前序:Round 06 单来源 × 八 seed × 八条件采样
|
||||
|
||||
## 0. 一句话结论
|
||||
|
||||
本轮保持 256 条正式生成的总预算不变,把网格从 Round 06 的:
|
||||
|
||||
```text
|
||||
4 sources × 8 seeds × 8 conditions
|
||||
```
|
||||
|
||||
改成:
|
||||
|
||||
```text
|
||||
16 sources × 4 seeds × 4 conditions
|
||||
```
|
||||
|
||||
结果为:
|
||||
|
||||
```text
|
||||
256 sampled outputs
|
||||
├─ 250 natural EOS
|
||||
├─ 6 budget truncated
|
||||
├─ 247 个不同的完整 token trajectory hashes
|
||||
├─ Math:47 / 64 strict exact
|
||||
├─ Code:52 / 64 official tests pass
|
||||
├─ R0 / R1:63 / 64 同 source-condition 轨迹分叉
|
||||
└─ 新进程 R0:64 / 64 格的八项合同字段 exact
|
||||
```
|
||||
|
||||
最重要的不是总正确数下降,而是总数背后的题目异质性终于可见:
|
||||
|
||||
```text
|
||||
Math tasks:14 / 16 · 16 / 16 · 8 / 16 · 9 / 16
|
||||
Code tasks:16 / 16 · 8 / 16 · 12 / 16 · 16 / 16
|
||||
```
|
||||
|
||||
Round 06 的 Math `62 / 64` 与 Code `63 / 64` 各来自同一道题的重复采样,不能代表
|
||||
64 道独立任务。本轮也仍不是完整 benchmark;它只把覆盖单位从每域 1 条扩大到每域
|
||||
4 条,并第一次允许我们描述这 4 条 source 之间的离散性。
|
||||
|
||||
---
|
||||
|
||||
## 1. 这轮修复的统计对象错误
|
||||
|
||||
### 1.1 seed 是题内重复,不是新题
|
||||
|
||||
对固定 Math source 抽 64 个 seed,能够回答:
|
||||
|
||||
- 同一道题的有限采样轨迹怎样分叉;
|
||||
- 这道题的任务结果是否对 seed 敏感;
|
||||
- 固定 seed 与执行环境时能否复跑出同一轨迹。
|
||||
|
||||
它不能回答:
|
||||
|
||||
- 模型在 64 道不同 Math 题上的能力;
|
||||
- source 换成另一题后,条件方向是否保留;
|
||||
- 这道题是否容易、典型或能代表 GSM8K;
|
||||
- 一个由 seed 数制造的窄区间是否具有 benchmark 意义。
|
||||
|
||||
本轮因此冻结:
|
||||
|
||||
```text
|
||||
主要覆盖单位:source
|
||||
source 内重复:seed
|
||||
```
|
||||
|
||||
所有方向先在每条 source 内,由四个 seed 的 cell summary 计算;然后只描述四条
|
||||
source 的正、零、负方向数量。没有把题内 seed 当作独立 source。
|
||||
|
||||
### 1.2 为什么仍保留 micro total
|
||||
|
||||
`47 / 64` 与 `52 / 64` 仍有用,因为它们精确描述这份固定网格中通过的输出数。但它们
|
||||
必须与逐题矩阵一起展示:
|
||||
|
||||
```text
|
||||
micro total
|
||||
≠ 题目难度分布
|
||||
≠ benchmark accuracy
|
||||
≠ population estimate
|
||||
```
|
||||
|
||||
### 1.3 本轮不是盲确认
|
||||
|
||||
`period` 是在看到 Round 06 结果后选出的定向 follow-up:上一轮句点格在单来源上出现
|
||||
明显的长度收缩与轨迹集合变化。因此:
|
||||
|
||||
- 本轮可以检查前序方向在新增 source 上是否保留;
|
||||
- 不能称为完全独立、未见结果的 confirmatory experiment;
|
||||
- 不能从 `11 / 16` 同向构造事后显著性故事。
|
||||
|
||||
---
|
||||
|
||||
## 2. 冻结执行合同
|
||||
|
||||
### 2.1 模型与软件
|
||||
|
||||
| 对象 | 冻结值 |
|
||||
|---|---|
|
||||
| checkpoint | `deepseek-ai/DeepSeek-V2-Lite-Chat` |
|
||||
| revision | `85864749cd611b4353ce1decdb286193298f64c7` |
|
||||
| dtype | BF16 |
|
||||
| attention | eager |
|
||||
| PyTorch | `2.11.0+cu128` |
|
||||
| Transformers | `4.41.2` |
|
||||
| CUDA placement | embedding + layers 0–23 |
|
||||
| CPU placement | layers 24–26 + final norm + `lm_head` |
|
||||
| padding | left padding;`PAD=EOS`;padding mask 为 0 |
|
||||
| cache | `use_cache=true` |
|
||||
|
||||
CPU offload 是本机执行拓扑,不是模型结构或能力属性。
|
||||
|
||||
### 2.2 解码参数
|
||||
|
||||
固定 revision 的官方 `generation_config.json` 给出:
|
||||
|
||||
```text
|
||||
do_sample = true
|
||||
temperature = 0.3
|
||||
top_p = 0.95
|
||||
```
|
||||
|
||||
正式执行显式传入:
|
||||
|
||||
```text
|
||||
do_sample = true
|
||||
temperature = 0.3
|
||||
top_p = 0.95
|
||||
top_k = 0
|
||||
max_new_tokens = 512
|
||||
use_cache = true
|
||||
```
|
||||
|
||||
`top_k=0` 用于关闭 Transformers 通用默认 top-k,不是额外的 checkpoint 作者结论。
|
||||
|
||||
### 2.3 四格条件
|
||||
|
||||
```text
|
||||
s0_eos
|
||||
s1_eos
|
||||
s0_period
|
||||
s1_period
|
||||
```
|
||||
|
||||
- `s0/s1`:system off/on;
|
||||
- `eos`:官方 assistant 历史边界;
|
||||
- `period`:把该位置的 EOS 单 ID 替换为普通句点 ID 13。
|
||||
|
||||
句点格是单 ID 反事实,不是官方有效聊天序列。每个 source × replicate 的四格放在同一个
|
||||
batch,行顺序固定。
|
||||
|
||||
### 2.4 四个 seed
|
||||
|
||||
| replicate | base seed |
|
||||
|---:|---:|
|
||||
| R0 | 2,101,325,316 |
|
||||
| R1 | 2,511,573,438 |
|
||||
| R2 | 1,677,220,094 |
|
||||
| R3 | 2,412,346,607 |
|
||||
|
||||
base seed 在结果之前由协议字符串的 SHA-256 派生。每条 source 再从 base seed 与
|
||||
source ID 派生 run seed。四行共享 batch seed 与随机调用时序,但各行消费不同 RNG
|
||||
子流,因此仍是:
|
||||
|
||||
```text
|
||||
batch-seed aligned
|
||||
≠ common-random-number paired
|
||||
```
|
||||
|
||||
本轮不报告 paired p-value。
|
||||
|
||||
---
|
||||
|
||||
## 3. source 选择与输入身份
|
||||
|
||||
source 不是按 Round 06 的正确率、完成率或文本质量挑选,而是来自已经冻结的公开语料
|
||||
选择合同:
|
||||
|
||||
```text
|
||||
sample salt = llm-atlas-deepseek-routing-template-control-v1
|
||||
每域按 SHA-256 selection_rank 排序,取 within_domain_index 0–3
|
||||
```
|
||||
|
||||
### 3.1 四域各四条
|
||||
|
||||
| domain | index 0 | index 1 | index 2 | index 3 |
|
||||
|---|---|---|---|---|
|
||||
| English | WikiText `0443` | `0030` | `2909` | `2746` |
|
||||
| Chinese | TNEWS `4855` | `8935` | `3448` | `1059` |
|
||||
| Code | HumanEval `31` | `44` | `133` | `23` |
|
||||
| Math | GSM8K `1069` | `1228` | `0144` | `1251` |
|
||||
|
||||
### 3.2 prompt hash 闸门
|
||||
|
||||
本轮的:
|
||||
|
||||
```text
|
||||
16 sources × 4 conditions = 64 prompt cells
|
||||
```
|
||||
|
||||
与 Round 05 已执行的对应输入逐格比较:
|
||||
|
||||
```text
|
||||
64 / 64 prompt token SHA-256 exact
|
||||
```
|
||||
|
||||
HumanEval canonical solution/tests 与 GSM8K gold 只在生成冻结后进入独立 evaluator,
|
||||
从未进入模型 prompt。
|
||||
|
||||
---
|
||||
|
||||
## 4. smoke 与正式执行
|
||||
|
||||
### 4.1 16-token smoke
|
||||
|
||||
```text
|
||||
16 sources × R0/R1 × 4 conditions = 128 short outputs
|
||||
16 sources × R0 replay × 4 conditions = 64 replay outputs
|
||||
```
|
||||
|
||||
| 闸门 | 结果 |
|
||||
|---|---:|
|
||||
| prompt hash exact | 64 / 64 |
|
||||
| R0 同进程重放全合同 exact | 64 / 64 |
|
||||
| R0/R1 可比格 | 64 |
|
||||
| R0/R1 trajectory 分叉 | 34 / 64 |
|
||||
| OOM / NaN / exception | 0 |
|
||||
|
||||
16-token smoke 只检查执行合同,没有进入正式统计。
|
||||
|
||||
### 4.2 正式网格
|
||||
|
||||
```text
|
||||
16 sources
|
||||
× 4 seeds
|
||||
× 4 conditions
|
||||
× 512 new-token cap
|
||||
= 256 sampled outputs
|
||||
```
|
||||
|
||||
| 检查 | 结果 |
|
||||
|---|---:|
|
||||
| sources | 16 |
|
||||
| sources / domain | 4 |
|
||||
| conditions / source | 4 |
|
||||
| seeds / source-condition | 4 |
|
||||
| total outputs | 256 |
|
||||
| missing / duplicate cells | 0 |
|
||||
| prompt hash exact | 64 / 64 |
|
||||
|
||||
执行记录:
|
||||
|
||||
| 指标 | 值 |
|
||||
|---|---:|
|
||||
| generation seconds | `3,520.086` |
|
||||
| peak CUDA allocated | `29,694,417,408 bytes` |
|
||||
| unique full trajectory hashes | `247 / 256` |
|
||||
| R0/R1 trajectory different | `63 / 64` |
|
||||
|
||||
这里的耗时属于 CPU-offloaded eager 机制审计,不是服务吞吐 benchmark。
|
||||
|
||||
---
|
||||
|
||||
## 5. stopping 与轨迹总账
|
||||
|
||||
| 指标 | 结果 |
|
||||
|---|---:|
|
||||
| natural EOS | 250 / 256 |
|
||||
| budget truncated | 6 / 256 |
|
||||
| unique full token trajectory hashes | 247 / 256 |
|
||||
| R0/R1 comparable cells | 64 |
|
||||
| R0/R1 different trajectories | 63 |
|
||||
|
||||
`247 unique` 只说明完整 token ID 序列的 SHA-256 不同,不说明存在 247 种语义、策略或
|
||||
推理方法。
|
||||
|
||||
高 EOS 率也不等于高正确率:
|
||||
|
||||
```text
|
||||
Math 的 17 条失败全部 natural EOS
|
||||
Code 有自然 EOS 的 runtime / assertion failure
|
||||
```
|
||||
|
||||
因此仍按四张账处理:
|
||||
|
||||
```text
|
||||
stopping
|
||||
→ task terminal
|
||||
→ evaluator coverage
|
||||
→ correctness
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. Math:47 / 64 怎样分布
|
||||
|
||||
独立 evaluator 使用 GSM8K 官方 gold,抽取最终数值;严格成功要求:
|
||||
|
||||
```text
|
||||
自然完成
|
||||
+ evaluator covered
|
||||
+ predicted numeric answer == gold numeric answer
|
||||
```
|
||||
|
||||
### 6.1 4 tasks × 4 conditions
|
||||
|
||||
| source | S0·EOS | S1·EOS | S0·句点 | S1·句点 | total |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| GSM8K/1069 | 3/4 | 4/4 | 3/4 | 4/4 | 14/16 |
|
||||
| GSM8K/1228 | 4/4 | 4/4 | 4/4 | 4/4 | 16/16 |
|
||||
| GSM8K/0144 | 2/4 | 2/4 | 3/4 | 1/4 | 8/16 |
|
||||
| GSM8K/1251 | 1/4 | 4/4 | 1/4 | 3/4 | 9/16 |
|
||||
|
||||
总计:
|
||||
|
||||
```text
|
||||
47 / 64 strict exact
|
||||
```
|
||||
|
||||
### 6.2 这张矩阵教会什么
|
||||
|
||||
同样是预先冻结的 GSM8K source:
|
||||
|
||||
- `/1228` 在 16 次生成中全对;
|
||||
- `/0144` 只有 8 / 16;
|
||||
- `/1251` 只有 9 / 16;
|
||||
- `/1069` 的 14 / 16 更接近 Round 06 单题高正确率。
|
||||
|
||||
所以 Round 06 的 `62 / 64` 首先是 `/1069` 这一道题的局部性质,不能推广成模型的
|
||||
GSM8K 采样正确率。
|
||||
|
||||
### 6.3 source-blocked task contrast
|
||||
|
||||
system main effect 的四条 source 方向:
|
||||
|
||||
```text
|
||||
2 positive · 1 zero · 1 negative
|
||||
```
|
||||
|
||||
句点减 EOS 的 main effect:
|
||||
|
||||
```text
|
||||
0 positive · 3 zero · 1 negative
|
||||
```
|
||||
|
||||
interaction:
|
||||
|
||||
```text
|
||||
0 positive · 2 zero · 2 negative
|
||||
```
|
||||
|
||||
这些是固定四题的描述量,不是总体效应。
|
||||
|
||||
### 6.4 失败身份
|
||||
|
||||
Math 有 17 条 strict failure:
|
||||
|
||||
- `/1069`:2;
|
||||
- `/0144`:8;
|
||||
- `/1251`:7;
|
||||
- `/1228`:0。
|
||||
|
||||
17 条全部自然 EOS,全部为 evaluator 已覆盖但最终数值与 gold 不同。失败不是正则抽取器
|
||||
无法找到答案,也不是长度上限截断。
|
||||
|
||||
---
|
||||
|
||||
## 7. Code:52 / 64 怎样分布
|
||||
|
||||
生成完成后,独立 evaluator:
|
||||
|
||||
1. 抽取 Python candidate;
|
||||
2. 做 AST parse;
|
||||
3. 在固定容器中执行 HumanEval official tests;
|
||||
4. 记录 passed、runtime error、assertion failed;
|
||||
5. 以 candidate/test/harness hashes 做安全缓存。
|
||||
|
||||
### 7.1 沙箱合同
|
||||
|
||||
| 对象 | 值 |
|
||||
|---|---|
|
||||
| image | `python:3.11-alpine@sha256:25976e9d…7702a4` |
|
||||
| network | none |
|
||||
| filesystem | read-only |
|
||||
| user | `65534:65534` |
|
||||
| capabilities | all dropped |
|
||||
| no new privileges | true |
|
||||
| memory / swap | 256 MiB / 256 MiB |
|
||||
| pids | 64 |
|
||||
| cpus | 0.5 |
|
||||
| timeout | 5 seconds |
|
||||
| unique code cache entries | 47 |
|
||||
| cache hits | 17 |
|
||||
|
||||
64 / 64 candidates 完成 AST 检查与执行身份记录。
|
||||
|
||||
### 7.2 4 tasks × 4 conditions
|
||||
|
||||
| source | S0·EOS | S1·EOS | S0·句点 | S1·句点 | total |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| HumanEval/31 | 4/4 | 4/4 | 4/4 | 4/4 | 16/16 |
|
||||
| HumanEval/44 | 2/4 | 1/4 | 1/4 | 4/4 | 8/16 |
|
||||
| HumanEval/133 | 4/4 | 4/4 | 4/4 | 0/4 | 12/16 |
|
||||
| HumanEval/23 | 4/4 | 4/4 | 4/4 | 4/4 | 16/16 |
|
||||
|
||||
总计:
|
||||
|
||||
```text
|
||||
52 / 64 official tests pass
|
||||
```
|
||||
|
||||
### 7.3 总 interaction 为 0,却不代表每题为 0
|
||||
|
||||
Code strict success 的 source-level interaction:
|
||||
|
||||
| source | interaction |
|
||||
|---|---:|
|
||||
| HumanEval/31 | 0 |
|
||||
| HumanEval/44 | +1.0 |
|
||||
| HumanEval/133 | −1.0 |
|
||||
| HumanEval/23 | 0 |
|
||||
| domain mean | 0 |
|
||||
|
||||
这是本轮最清楚的聚合陷阱:
|
||||
|
||||
```text
|
||||
domain mean = 0
|
||||
```
|
||||
|
||||
不是“每道题都没有 interaction”,而是 `/44` 与 `/133` 正负完全抵消。
|
||||
|
||||
### 7.4 失败身份
|
||||
|
||||
12 条 Code strict failure:
|
||||
|
||||
```text
|
||||
HumanEval/44
|
||||
├─ 5 runtime_error
|
||||
└─ 3 assertion_failed
|
||||
|
||||
HumanEval/133
|
||||
└─ 4 assertion_failed(全部 s1_period)
|
||||
```
|
||||
|
||||
其中只有 1 条属于 budget-truncated unresolved;其余失败不能归因于 512-token 上限。
|
||||
AST 合法、自然 EOS 或说明流畅都不能替代 official tests。
|
||||
|
||||
---
|
||||
|
||||
## 8. 句点是否普遍缩短生成
|
||||
|
||||
每条 source 内先计算:
|
||||
|
||||
```text
|
||||
period main
|
||||
= mean(S0·period, S1·period)
|
||||
− mean(S0·EOS, S1·EOS)
|
||||
```
|
||||
|
||||
负数表示句点格更短。
|
||||
|
||||
### 8.1 四域方向
|
||||
|
||||
| domain | shorter sources | domain mean | source median |
|
||||
|---|---:|---:|---:|
|
||||
| English | 1 / 4 | −8.5 | +21.2 |
|
||||
| Chinese | 3 / 4 | −13.9 | −13.0 |
|
||||
| Code | 4 / 4 | −130.0 | −166.1 |
|
||||
| Math | 3 / 4 | −19.8 | −9.6 |
|
||||
| total | 11 / 16 | — | — |
|
||||
|
||||
### 8.2 English 直接修正 Round 06 的单来源印象
|
||||
|
||||
| source | period − EOS mean tokens |
|
||||
|---|---:|
|
||||
| WikiText/0443 | −121.0 |
|
||||
| WikiText/0030 | +44.5 |
|
||||
| WikiText/2909 | +41.0 |
|
||||
| WikiText/2746 | +1.375 |
|
||||
|
||||
如果只看 Round 06 的 `/0443`,会写成:
|
||||
|
||||
```text
|
||||
句点大幅缩短 English 输出
|
||||
```
|
||||
|
||||
加入三条预先冻结的 English source 后:
|
||||
|
||||
```text
|
||||
1 条更短
|
||||
3 条更长
|
||||
domain mean 仍是 −8.5
|
||||
source median 却是 +21.2
|
||||
```
|
||||
|
||||
均值与多数方向相反,是因为 `/0443` 的 `−121` 足以压过另外三条正值。这正是为什么
|
||||
本轮在前端同时展示四个 source 点、均值和中位数。
|
||||
|
||||
### 8.3 其余三域
|
||||
|
||||
Chinese:
|
||||
|
||||
```text
|
||||
+0.375 · −30.0 · −7.5 · −18.5
|
||||
```
|
||||
|
||||
Code:
|
||||
|
||||
```text
|
||||
−178.625 · −164.375 · −167.875 · −9.0
|
||||
```
|
||||
|
||||
Math:
|
||||
|
||||
```text
|
||||
−4.625 · −68.75 · +8.875 · −14.5
|
||||
```
|
||||
|
||||
Code 的四条 source 同向,是当前固定网格里最一致的长度模式;仍不能写成对
|
||||
HumanEval 总体或所有代码任务的普遍规律。
|
||||
|
||||
---
|
||||
|
||||
## 9. system main effect 也依赖 source
|
||||
|
||||
平均生成长度的 system main effect:
|
||||
|
||||
| domain | negative | zero | positive | mean |
|
||||
|---|---:|---:|---:|---:|
|
||||
| English | 4 | 0 | 0 | −137.0 |
|
||||
| Chinese | 2 | 0 | 2 | −6.9 |
|
||||
| Code | 4 | 0 | 0 | −143.3 |
|
||||
| Math | 3 | 0 | 1 | −13.6 |
|
||||
|
||||
English 与 Code 在这四条 source 上都呈 system-on 更短;Chinese 则正负各半。不能把
|
||||
一个域的方向复制到另一个域。
|
||||
|
||||
自然 EOS 已接近天花板:
|
||||
|
||||
- system-on 提升 EOS 的 source:English 2 / 4;
|
||||
- Chinese、Code、Math:均 0 / 4;
|
||||
- 这主要说明多数 cell 已经 4 / 4 EOS,不是 system 没有其它行为差异。
|
||||
|
||||
---
|
||||
|
||||
## 10. source 间与 source 内离散
|
||||
|
||||
每个 domain-condition 同时保存:
|
||||
|
||||
```text
|
||||
within-source seed length range
|
||||
between-source mean length range
|
||||
within-source seed success range
|
||||
between-source success-rate range
|
||||
```
|
||||
|
||||
这些量回答不同问题:
|
||||
|
||||
- within-source:固定一道题,四个 seed 能让长度或成功怎样变化;
|
||||
- between-source:固定域与 condition,四道题的均值或成功率跨度多大。
|
||||
|
||||
它们没有被压成一个“哪种随机性更大”的全局数字,因为:
|
||||
|
||||
- length 与 correctness 量纲不同;
|
||||
- 四条 source 太少,不适合稳定估计方差分量;
|
||||
- seed 并不是逐行 common random numbers;
|
||||
- source 不是从总体随机抽样。
|
||||
|
||||
前端允许逐域逐 condition 检查两个范围,但不输出伪精确的总体方差比例。
|
||||
|
||||
---
|
||||
|
||||
## 11. 独立新进程复跑
|
||||
|
||||
正式生成后,全新进程只重跑:
|
||||
|
||||
```text
|
||||
16 sources × R0 × 4 conditions = 64 outputs
|
||||
```
|
||||
|
||||
逐格比较八项预注册字段:
|
||||
|
||||
| 字段 | exact |
|
||||
|---|---:|
|
||||
| run seed | 64 / 64 |
|
||||
| prompt hash | 64 / 64 |
|
||||
| complete generated token IDs | 64 / 64 |
|
||||
| decoded text | 64 / 64 |
|
||||
| EOS state | 64 / 64 |
|
||||
| truncation state | 64 / 64 |
|
||||
| CPU RNG pre-state hash | 64 / 64 |
|
||||
| CUDA RNG pre-state hash | 64 / 64 |
|
||||
|
||||
因此这份固定合同同时满足:
|
||||
|
||||
```text
|
||||
不同 seed:63 / 64 同格轨迹分叉
|
||||
相同合同:64 / 64 新进程轨迹 exact
|
||||
```
|
||||
|
||||
“采样会变化”和“采样可复现”并不矛盾。
|
||||
|
||||
exact replay 不承诺跨 PyTorch、Transformers、CUDA kernel、硬件或 batch 合同复现。
|
||||
|
||||
---
|
||||
|
||||
## 12. 工件与 hash chain
|
||||
|
||||
| 工件 | bytes | SHA-256 |
|
||||
|---|---:|---|
|
||||
| formal sampling | 1,980,601 | `f013132485f27adce008f03f781bed9982efc0d7939f13faede01c9f6f3d7f7c` |
|
||||
| independent evaluation | 370,580 | `e88b274599fc5951561f9e5e7438fb4d6d25f341ae3bc4a2c4121877a0db8975` |
|
||||
| fresh-process R0 rerun | 597,066 | `143dc9d0f7c914db4781e36b1401cdc9fbc2971a0dca71188bab0f8cedb002a6` |
|
||||
| reproduction comparison | 30,037 | `ec4a47894953f5d73bb62211b588632ef7032da0e915218992b7fc3ef9c2a556` |
|
||||
| source-blocked analysis | 168,001 | `d0dece388998fee419d34ff33f140695a9fedef6e79799cdf42eb283b047bc84` |
|
||||
|
||||
前端压缩工件:
|
||||
|
||||
```text
|
||||
src/data/deepseek-v2-lite-chat-cross-source-sampling-compact.json
|
||||
SHA-256 d0ab65646c6119bdeafeb451103dc6afebff3624a1e05f13a45a52ad965be1af
|
||||
```
|
||||
|
||||
builder 在生成前端数据之前验证:
|
||||
|
||||
1. evaluator 输入 hash 指向正式 sampling;
|
||||
2. reproduction 的 formal/rerun hashes 指向对应文件;
|
||||
3. analysis 的三项输入 hashes 全部一致;
|
||||
4. 64 格八项复现字段全部 exact;
|
||||
5. `16 × 4 × 4 = 256` 网格合同成立。
|
||||
|
||||
---
|
||||
|
||||
## 13. 一手来源与实现身份
|
||||
|
||||
### 13.1 官方 checkpoint 配置
|
||||
|
||||
- DeepSeek-V2-Lite-Chat generation config:
|
||||
<https://huggingface.co/deepseek-ai/DeepSeek-V2-Lite-Chat/blob/main/generation_config.json>
|
||||
|
||||
这里只把固定 revision 文件中的 `.3/.95` 记为作者发布配置;`top_k=0` 是本实验为关闭
|
||||
Transformers 通用 top-k 而显式加入的执行参数。
|
||||
|
||||
### 13.2 nucleus sampling
|
||||
|
||||
- Holtzman et al., *The Curious Case of Neural Text Degeneration*:
|
||||
<https://arxiv.org/abs/1904.09751>
|
||||
|
||||
本轮使用 nucleus sampling 的有限样本,不声称恢复完整生成分布。
|
||||
|
||||
### 13.3 任务来源
|
||||
|
||||
- Chen et al., *Evaluating Large Language Models Trained on Code*:
|
||||
<https://arxiv.org/abs/2107.03374>
|
||||
- Cobbe et al., *Training Verifiers to Solve Math Word Problems*:
|
||||
<https://arxiv.org/abs/2110.14168>
|
||||
|
||||
HumanEval official tests 与 GSM8K gold 只服务于本轮固定四题的独立评测,不生成
|
||||
benchmark 总体分数。
|
||||
|
||||
---
|
||||
|
||||
## 14. 可以说与不能说
|
||||
|
||||
### 可以说
|
||||
|
||||
- 在固定 16 条 source 的网格中,250 / 256 条生成自然 EOS;
|
||||
- 247 / 256 条完整 token trajectories 唯一;
|
||||
- R0/R1 在 63 / 64 同 source-condition cells 分叉;
|
||||
- 新进程 R0 在 64 / 64 格、八项预注册字段 exact;
|
||||
- 四道 Math 的 strict pass 从 8 / 16 到 16 / 16;
|
||||
- 四道 Code 的 tests pass 从 8 / 16 到 16 / 16;
|
||||
- 句点在 11 / 16 条 source 上缩短平均生成长度;
|
||||
- English 的域均值与多数 source 方向相反;
|
||||
- Code 的 task interaction 域均值为 0,但两条非零 source 正负抵消。
|
||||
|
||||
### 不能说
|
||||
|
||||
- DeepSeek-V2-Lite-Chat 的 GSM8K 准确率是 47 / 64;
|
||||
- DeepSeek-V2-Lite-Chat 的 HumanEval pass@1 是 52 / 64;
|
||||
- 16 条 source 随机代表四个数据集总体;
|
||||
- 4 个 seed 是 4 道独立题;
|
||||
- 11 / 16 构成总体显著性;
|
||||
- period 是官方有效聊天边界;
|
||||
- period 普遍缩短所有 English、Math 或中文输出;
|
||||
- domain mean 为 0 说明每条 source 都没有效应;
|
||||
- natural EOS、AST parse 或可执行等同于正确;
|
||||
- exact replay 会跨软件、硬件和 batch 合同自动成立。
|
||||
|
||||
---
|
||||
|
||||
## 15. 下一步
|
||||
|
||||
本轮修复了“单题多 seed 冒充多题覆盖”,但仍只有每域 4 条 source。下一轮优先级应是:
|
||||
|
||||
1. 把任务覆盖扩大到足以报告 task-level bootstrap,并在结果前冻结抽样框;
|
||||
2. 为每条 source 使用独立 per-row RNG stream,构造真正可解释的 common-random-number
|
||||
条件对;
|
||||
3. 分离 ordinary boundary token 的词法身份、频率与位置作用,不只复查句点;
|
||||
4. 对 HumanEval/44 与 /133 的相反 interaction 做预注册 failure taxonomy;
|
||||
5. 在足够 source 覆盖后,再决定是否值得进行干预式 mediation;
|
||||
6. 继续保持生成、evaluator、统计分析与 reproduction 四套工件分离。
|
||||
|
||||
本轮已经回答的是:
|
||||
|
||||
```text
|
||||
同样的 sampling 配方与输入干预,换一条 source 后,方向会不会变?
|
||||
```
|
||||
|
||||
答案是:
|
||||
|
||||
```text
|
||||
会,而且聚合均值有时会与多数 source 方向相反。
|
||||
```
|
||||
@@ -0,0 +1,454 @@
|
||||
# DeepSeek-V2-Lite-Chat 跨来源采样协议
|
||||
|
||||
> 状态:已按预注册协议执行;正式结果见
|
||||
> `research/DEEPSEEK_V2_LITE_CHAT_CROSS_SOURCE_SAMPLING_AUDIT.md`
|
||||
>
|
||||
> 注册日期:2026-07-30
|
||||
>
|
||||
> 协议 ID:`llm-atlas-deepseek-chat-cross-source-sampling-v1`
|
||||
>
|
||||
> 模型:`deepseek-ai/DeepSeek-V2-Lite-Chat`
|
||||
>
|
||||
> revision:`85864749cd611b4353ce1decdb286193298f64c7`
|
||||
>
|
||||
> 前序实验:Round 06 单来源 × 八 seed × 八条件采样
|
||||
|
||||
## 0. 这一轮补什么,不补什么
|
||||
|
||||
Round 06 固定四条 source,每条 source 在八个 seed 和八种边界条件下生成:
|
||||
|
||||
```text
|
||||
4 sources × 8 seeds × 8 conditions = 256 outputs
|
||||
```
|
||||
|
||||
它证明了两件看似矛盾但可以同时成立的事:
|
||||
|
||||
1. 不同 seed 会让同一格的 sampled trajectory 分叉;
|
||||
2. 固定 checkpoint、输入、batch 行顺序、软件与 seed 后,同一 trajectory 可以跨进程逐
|
||||
token 复现。
|
||||
|
||||
但 Math 和 Code 各只有一道题。即使那一道题有 64 个 sampled outputs,也仍然只有一个
|
||||
任务对象。把 64 个 seed 重复当成 64 道独立 benchmark 题,会制造伪样本量。
|
||||
|
||||
Round 07 因此把预算从“同题更多 seed”移动到“更多预先选定的 source”:
|
||||
|
||||
```text
|
||||
16 sources × 4 seeds × 4 conditions = 256 outputs
|
||||
```
|
||||
|
||||
本轮第一次允许描述固定 4 道 Math / 4 道 Code 之间的离散性,但仍不是完整
|
||||
GSM8K/HumanEval benchmark,更不是模型总体能力估计。
|
||||
|
||||
---
|
||||
|
||||
## 1. 研究问题与证据层级
|
||||
|
||||
### 1.1 主要问题
|
||||
|
||||
在固定 source selection、官方 sampling 参数与四行 batch 合同下:
|
||||
|
||||
1. Round 06 观察到的 completion-length / EOS 模式能否在另外三条同域 source 上出现?
|
||||
2. Math / Code 的 pass count 是否由单一道题主导?
|
||||
3. `system off/on × EOS/period` 四格的方向是否跨 source 一致,还是明显依赖题目?
|
||||
4. source 间差异与 source 内 seed 差异,哪一个在当前固定网格中更大?
|
||||
|
||||
### 1.2 探索与确认的边界
|
||||
|
||||
选择 `period` 作为唯一 ordinary-token 对照,受 Round 06 的结果启发:它在单来源运行中
|
||||
显示了明显的输出长度与 exact-trajectory 收缩。因此本轮是**定向复查**,不是完全独立、
|
||||
未见前序结果的 confirmatory experiment。
|
||||
|
||||
本轮不会:
|
||||
|
||||
- 把 Round 06 与 Round 07 合并后计算“独立复现率”;
|
||||
- 把四道题的方向一致写成总体显著性;
|
||||
- 根据本轮结果继续替换 source、condition 或 seed;
|
||||
- 把 period 序列称为官方有效聊天格式;
|
||||
- 把题内四个 seed 当成四道独立题。
|
||||
|
||||
---
|
||||
|
||||
## 2. 模型、依赖与解码合同
|
||||
|
||||
继承冻结对象:
|
||||
|
||||
- 官方 SFT Chat checkpoint;
|
||||
- revision `85864749cd611b4353ce1decdb286193298f64c7`;
|
||||
- checkpoint 12 个文件及各自 SHA-256;
|
||||
- Transformers `4.41.2`;
|
||||
- PyTorch `2.11.0+cu128`;
|
||||
- BF16、官方 remote modeling code、eager attention;
|
||||
- CUDA resident:embedding + layers 0–23;
|
||||
- CPU offload:layers 24–26 + final norm + `lm_head`;
|
||||
- `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`;
|
||||
- GPU static placement `28GiB`、CPU placement `80GiB`;
|
||||
- 左填充,`PAD=EOS`,padding mask 为 0;
|
||||
- `use_cache=true`、统一 `max_new_tokens=512`。
|
||||
|
||||
固定 revision 的官方 `generation_config.json` 给出:
|
||||
|
||||
```text
|
||||
do_sample = true
|
||||
temperature = 0.3
|
||||
top_p = 0.95
|
||||
bos_id = 100000
|
||||
eos_id = 100001
|
||||
```
|
||||
|
||||
正式执行显式传入:
|
||||
|
||||
```text
|
||||
do_sample = true
|
||||
temperature = 0.3
|
||||
top_p = 0.95
|
||||
top_k = 0
|
||||
max_new_tokens = 512
|
||||
use_cache = true
|
||||
```
|
||||
|
||||
`top_k=0` 用于关闭 Transformers 通用默认 top-k;它不是 checkpoint 文件中的额外作者
|
||||
结论。
|
||||
|
||||
---
|
||||
|
||||
## 3. source 如何在结果之前冻结
|
||||
|
||||
source 不按 Round 06 的正确率、完成率或文本质量挑选。它们取自已冻结的
|
||||
`deepseek-v2-lite-routing-special-token-family-control.json`:
|
||||
|
||||
```text
|
||||
sample salt = llm-atlas-deepseek-routing-template-control-v1
|
||||
每域按 SHA-256 selection_rank 排序,取 within_domain_index 0–3
|
||||
```
|
||||
|
||||
### 3.1 English
|
||||
|
||||
| index | source ID |
|
||||
|---:|---|
|
||||
| 0 | `wikitext2/raw-validation/0443` |
|
||||
| 1 | `wikitext2/raw-validation/0030` |
|
||||
| 2 | `wikitext2/raw-validation/2909` |
|
||||
| 3 | `wikitext2/raw-validation/2746` |
|
||||
|
||||
### 3.2 Chinese
|
||||
|
||||
| index | source ID |
|
||||
|---:|---|
|
||||
| 0 | `tnews/test/4855` |
|
||||
| 1 | `tnews/test/8935` |
|
||||
| 2 | `tnews/test/3448` |
|
||||
| 3 | `tnews/test/1059` |
|
||||
|
||||
### 3.3 Code
|
||||
|
||||
| index | source ID |
|
||||
|---:|---|
|
||||
| 0 | `HumanEval/31` |
|
||||
| 1 | `HumanEval/44` |
|
||||
| 2 | `HumanEval/133` |
|
||||
| 3 | `HumanEval/23` |
|
||||
|
||||
### 3.4 Math
|
||||
|
||||
| index | source ID |
|
||||
|---:|---|
|
||||
| 0 | `gsm8k/test/1069` |
|
||||
| 1 | `gsm8k/test/1228` |
|
||||
| 2 | `gsm8k/test/0144` |
|
||||
| 3 | `gsm8k/test/1251` |
|
||||
|
||||
这些正是 Round 05 greedy completion 已执行的 16 条 source。本轮在加载模型前,必须验证
|
||||
16 × 4 = 64 个 prompt token hashes 与 Round 05 对应格逐格 exact。
|
||||
|
||||
---
|
||||
|
||||
## 4. 四格因子与 batch 合同
|
||||
|
||||
固定四格:
|
||||
|
||||
```text
|
||||
s0_eos, s1_eos, s0_period, s1_period
|
||||
```
|
||||
|
||||
其中:
|
||||
|
||||
- `s0/s1`:system off/on;
|
||||
- `eos`:官方 assistant 历史边界;
|
||||
- `period`:把该位置的 EOS 单 ID 替换为普通句点 ID 13;
|
||||
- system 文本、one-shot 内容、target source、角色词头、目标位置与 prompt 其余 token 均继承
|
||||
前序协议。
|
||||
|
||||
每个 source × replicate 的四格在同一个 batch,行顺序永久固定。四行 batch 与 Round 06
|
||||
的八行 batch 不同,因此即使复用相同 base seed,也不应期待 trajectory exact;本轮使用
|
||||
新协议派生的新 seed,避免把两种 batch 合同伪装成直接重复。
|
||||
|
||||
Transformers `4.41.2` 对整个 batch 调用一次 `torch.multinomial`,不接受逐行
|
||||
`Generator`。四格只共享 batch seed 与调用时序,每行消费不同 RNG 子流,因此仍是:
|
||||
|
||||
```text
|
||||
batch-seed aligned
|
||||
≠ common-random-number paired
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. seed 冻结
|
||||
|
||||
四个 base seed 由以下字符串做 SHA-256,取前 4 bytes big-endian unsigned integer:
|
||||
|
||||
```text
|
||||
llm-atlas-deepseek-chat-cross-source-sampling-v1/seed/{index}
|
||||
```
|
||||
|
||||
固定结果:
|
||||
|
||||
| replicate | base seed |
|
||||
|---:|---:|
|
||||
| R0 | 2,101,325,316 |
|
||||
| R1 | 2,511,573,438 |
|
||||
| R2 | 1,677,220,094 |
|
||||
| R3 | 2,412,346,607 |
|
||||
|
||||
实际 run seed:
|
||||
|
||||
```text
|
||||
SHA256(
|
||||
protocol_id + "/run\0"
|
||||
+ decimal(base_seed)
|
||||
+ "\0"
|
||||
+ source_id
|
||||
)[0:8] big-endian
|
||||
modulo (2^63 - 1)
|
||||
```
|
||||
|
||||
每次 source × replicate batch 前:
|
||||
|
||||
```python
|
||||
torch.manual_seed(run_seed)
|
||||
torch.cuda.manual_seed_all(run_seed)
|
||||
```
|
||||
|
||||
同时记录 CPU / CUDA RNG state 的执行前后 SHA-256。
|
||||
|
||||
---
|
||||
|
||||
## 6. 输入文件身份
|
||||
|
||||
| 输入 | SHA-256 |
|
||||
|---|---|
|
||||
| source selection contract | `c372c1b03a8b15f615b54ded5d9257a8fc2cdb7728001735d3d4c1d8534af5bf` |
|
||||
| Round 05 greedy baseline | `6af40512c5868caab0ef58356aaa7728384f2b7acc7c502aa89478fdf5a2a468` |
|
||||
| HumanEval gzip | `b796127e635a67f93fb35c04f4cb03cf06f38c8072ee7cee8833d7bee06979ef` |
|
||||
| GSM8K test JSONL | `3730d312f6e3440559ace48831e51066acaca737f6eabec99bccb9e4b3c39d14` |
|
||||
| TNEWS test JSONL | `74f199325768fbf2d6020711edfff23d653e99e0f8ac31a126a54db29c3a0ca8` |
|
||||
| TNEWS archive | `77c476e70cfe0b014a81b84c6e1db2142a8a2f52f4ae0a8216aa75e673933462` |
|
||||
| WikiText-2 validation parquet | `204929b7ff9d6184953f867dedb860e40aa69c078fc1e54b3baaa8fb28511c4c` |
|
||||
|
||||
HumanEval `canonical_solution` / tests 与 GSM8K answer 只进入生成后的独立 evaluator,绝不
|
||||
进入模型 prompt。
|
||||
|
||||
---
|
||||
|
||||
## 7. 执行网格与闸门
|
||||
|
||||
### 7.1 16-token smoke
|
||||
|
||||
```text
|
||||
16 sources × R0/R1 × 4 conditions × 16 tokens = 128 short outputs
|
||||
16 sources × R0 replay × 4 conditions × 16 tokens = 64 replay outputs
|
||||
```
|
||||
|
||||
通过条件:
|
||||
|
||||
1. 64 / 64 prompt hashes 与 Round 05 exact;
|
||||
2. 128 条 smoke 无 OOM、NaN 或 exception;
|
||||
3. R0 同进程 replay 的 run seed、prompt hash、token IDs、text 与 stop state 64 / 64 exact;
|
||||
4. R0/R1 的 64 个同 source-condition cells 至少一格 trajectory 分叉;
|
||||
5. 正式参数确实是 `.3/.95/top-k 0`。
|
||||
|
||||
smoke 输出不进入正式统计。
|
||||
|
||||
### 7.2 正式网格
|
||||
|
||||
```text
|
||||
16 sources
|
||||
× 4 seeds
|
||||
× 4 conditions
|
||||
× 512 new-token cap
|
||||
= 256 sampled outputs
|
||||
```
|
||||
|
||||
source 顺序固定为 English → Chinese → Code → Math,各域按 `within_domain_index` 升序;
|
||||
replicate 固定 R0 → R3;每个 replicate 内行顺序固定为四格顺序。不能根据中间输出提前
|
||||
停止或只续写截断格。
|
||||
|
||||
### 7.3 新进程复跑
|
||||
|
||||
正式完成后重新启动 Python、重新加载模型,只复跑 R0:
|
||||
|
||||
```text
|
||||
16 sources × 1 seed × 4 conditions = 64 outputs
|
||||
```
|
||||
|
||||
逐格核对:
|
||||
|
||||
- run seed;
|
||||
- prompt token hash;
|
||||
- 完整 generated token IDs;
|
||||
- decoded text;
|
||||
- EOS state;
|
||||
- truncation state;
|
||||
- CPU RNG pre-state hash;
|
||||
- CUDA RNG pre-state hash。
|
||||
|
||||
只能写“64 / 64 R0 cells independently reproduced”,不能写 256 / 256。
|
||||
|
||||
---
|
||||
|
||||
## 8. 独立 evaluator
|
||||
|
||||
仍把四张账分开:
|
||||
|
||||
1. stopping:natural EOS / budget truncated;
|
||||
2. task terminal:明确答案、闭合 code fence 或 tests pass;
|
||||
3. evaluator coverage:数值可抽取,或代码 AST + sandbox 已执行;
|
||||
4. correctness:GSM8K strict numeric exact / HumanEval official tests pass。
|
||||
|
||||
HumanEval 每个独特 candidate 使用固定镜像:
|
||||
|
||||
```text
|
||||
python:3.11-alpine@
|
||||
sha256:25976e9d34a0fab1f278cae931f34c8303d97bf0c0d7f85b6b4dcf641d7702a4
|
||||
```
|
||||
|
||||
沙箱:
|
||||
|
||||
```text
|
||||
network none · read-only filesystem · user 65534:65534
|
||||
cap-drop ALL · no-new-privileges · no host mounts
|
||||
256 MiB memory/swap · pids 64 · cpus .5 · timeout 5s
|
||||
```
|
||||
|
||||
完全相同的 candidate + task + tests + harness 可以复用 execution cache,但保留每个生成格
|
||||
自己的 evaluator row。
|
||||
|
||||
---
|
||||
|
||||
## 9. 预注册统计
|
||||
|
||||
### 9.1 source × condition
|
||||
|
||||
每个集合有 4 个 seed,报告:
|
||||
|
||||
- natural EOS / 4;
|
||||
- truncated / 4;
|
||||
- generated length mean / min / max;
|
||||
- unique full-token trajectory hashes / 4;
|
||||
- 六对 seed pair 的 token similarity;
|
||||
- Math:4 个抽取答案、严格正确数、唯一绝对多数答案(若存在);
|
||||
- Code:AST、executed、official tests pass、observed any-pass。
|
||||
|
||||
`observed any-pass` 只表示当前四次抽样至少一次通过,不称为标准 HumanEval pass@4。
|
||||
|
||||
### 9.2 domain × condition
|
||||
|
||||
每个 domain 的四条 source 等权,报告:
|
||||
|
||||
- micro count:16 个 outputs 的合计;
|
||||
- source macro mean:先算每条 source 的四 seed 比例,再对四条 source 等权;
|
||||
- source range:四条 source 的最小/最大;
|
||||
- 4 × 4 source-condition 矩阵。
|
||||
|
||||
因为每条 source 的 seed 数相同,micro 与 macro 点估计可能数值相同;两者仍分开呈现,
|
||||
避免未来不等重复数时悄悄改变权重。
|
||||
|
||||
### 9.3 source-blocked 四格 contrast
|
||||
|
||||
对每条 source 先算四格 cell mean,再形成描述性 contrast:
|
||||
|
||||
```text
|
||||
system main =
|
||||
mean(s1_eos, s1_period) - mean(s0_eos, s0_period)
|
||||
|
||||
boundary main =
|
||||
mean(s0_period, s1_period) - mean(s0_eos, s1_eos)
|
||||
|
||||
interaction =
|
||||
(s1_period - s1_eos) - (s0_period - s0_eos)
|
||||
```
|
||||
|
||||
分别对:
|
||||
|
||||
- natural EOS rate;
|
||||
- generated-token length;
|
||||
- Math / Code fixed-budget success;
|
||||
- Math / Code strict-complete success。
|
||||
|
||||
每个 domain 展示四条 source 的 contrast dots、median、min、max 与正/零/负方向计数。不算
|
||||
p-value,不给总体置信区间。
|
||||
|
||||
### 9.4 两层离散性
|
||||
|
||||
对 generated-token length 与 task success:
|
||||
|
||||
- `within-source seed range`:同 source-condition 四 seed 的范围;
|
||||
- `between-source cell-mean range`:同 domain-condition 四条 source 均值的范围。
|
||||
|
||||
这只是固定网格中的描述,不作随机效应方差分解,也不把四条 source 当作领域总体随机样本。
|
||||
|
||||
### 9.5 前序方向复查
|
||||
|
||||
只针对 Round 06 已观察到的方向登记:
|
||||
|
||||
```text
|
||||
period 相对 EOS 是否缩短 mean generated tokens?
|
||||
system on 相对 off 是否提高 natural EOS rate?
|
||||
```
|
||||
|
||||
按 source 报告方向,不因结果改写为“改善质量”。Math / Code correctness 单独展示,不能由
|
||||
长度或 EOS 代替。
|
||||
|
||||
---
|
||||
|
||||
## 10. 失败与修订规则
|
||||
|
||||
- smoke 前可修实现错误;正式 source/seed/condition/metrics 不随输出改动;
|
||||
- 正式 JSON 完整写出前失败,整轮重新开始,不保留“表现较好”的部分 source;
|
||||
- OOM 优先降低 static GPU placement,并用同 seed smoke 做 placement exact 闸门;
|
||||
- 不拆四行 batch;拆 batch 或改行顺序意味着新协议;
|
||||
- evaluator bug 可以修复并重跑 evaluator,但 sampled output 文件冻结;
|
||||
- 新进程 R0 任一 preregistered 字段不 exact,则结果只作复现失败诊断;
|
||||
- 某题 evaluator uncovered 仍保留该格,不能删题或换题;
|
||||
- 沙箱超时、AST 失败、assertion failure 与 runtime error 分开记录。
|
||||
|
||||
---
|
||||
|
||||
## 11. 永久禁止的结论
|
||||
|
||||
- 4 道 GSM8K / HumanEval 代表完整 benchmark;
|
||||
- 64 个 task-domain outputs 等于 64 道独立题;
|
||||
- 4 seeds 足以估计完整生成分布;
|
||||
- observed any-pass 等于标准 pass@4;
|
||||
- period 是官方有效聊天边界;
|
||||
- system 或 period 让模型“更聪明”“更稳定”;
|
||||
- natural EOS、语义终点、可评测与正确是同一个指标;
|
||||
- 四条 source 的方向计数是总体显著性;
|
||||
- batch-seed aligned 是 common-random-number pair;
|
||||
- exact same-seed replay 可跨软件、kernel 或硬件保证;
|
||||
- CPU-offloaded eager 延迟等于生产服务吞吐;
|
||||
- 输出关联已经定位到 hidden-state / router mediation。
|
||||
|
||||
---
|
||||
|
||||
## 12. 一手来源
|
||||
|
||||
- DeepSeek-V2-Lite-Chat pinned `generation_config.json`:
|
||||
<https://huggingface.co/deepseek-ai/DeepSeek-V2-Lite-Chat/blob/85864749cd611b4353ce1decdb286193298f64c7/generation_config.json>
|
||||
- Holtzman et al., *The Curious Case of Neural Text Degeneration*:
|
||||
<https://arxiv.org/abs/1904.09751>
|
||||
- Chen et al., *Evaluating Large Language Models Trained on Code*:
|
||||
<https://arxiv.org/abs/2107.03374>
|
||||
- Cobbe et al., *Training Verifiers to Solve Math Word Problems*:
|
||||
<https://arxiv.org/abs/2110.14168>
|
||||
- OpenAI HumanEval pinned repository:
|
||||
<https://github.com/openai/human-eval/tree/6d43fb980f9fee3c892a914eda09951f772ad10d>
|
||||
@@ -0,0 +1,347 @@
|
||||
import { createHash } from "node:crypto";
|
||||
import { readFileSync, statSync, writeFileSync } from "node:fs";
|
||||
import { resolve } from "node:path";
|
||||
|
||||
const root = resolve(import.meta.dirname, "..");
|
||||
const paths = {
|
||||
sampling: resolve(
|
||||
root,
|
||||
"src/data/deepseek-v2-lite-chat-cross-source-sampling.json",
|
||||
),
|
||||
evaluation: resolve(
|
||||
root,
|
||||
"src/data/deepseek-v2-lite-chat-cross-source-sampling-eval.json",
|
||||
),
|
||||
rerun: resolve(
|
||||
root,
|
||||
"src/data/deepseek-v2-lite-chat-cross-source-sampling-repro-r0.json",
|
||||
),
|
||||
reproduction: resolve(
|
||||
root,
|
||||
"src/data/deepseek-v2-lite-chat-cross-source-sampling-reproduction.json",
|
||||
),
|
||||
analysis: resolve(
|
||||
root,
|
||||
"src/data/deepseek-v2-lite-chat-cross-source-sampling-analysis.json",
|
||||
),
|
||||
output: resolve(
|
||||
root,
|
||||
"src/data/deepseek-v2-lite-chat-cross-source-sampling-compact.json",
|
||||
),
|
||||
};
|
||||
|
||||
const readJson = (path) => JSON.parse(readFileSync(path, "utf8"));
|
||||
const sha256 = (path) => createHash("sha256")
|
||||
.update(readFileSync(path))
|
||||
.digest("hex");
|
||||
const artifact = (path) => ({
|
||||
bytes: statSync(path).size,
|
||||
sha256: sha256(path),
|
||||
});
|
||||
|
||||
const sampling = readJson(paths.sampling);
|
||||
const evaluation = readJson(paths.evaluation);
|
||||
const reproduction = readJson(paths.reproduction);
|
||||
const analysis = readJson(paths.analysis);
|
||||
const samplingArtifact = artifact(paths.sampling);
|
||||
const evaluationArtifact = artifact(paths.evaluation);
|
||||
const rerunArtifact = artifact(paths.rerun);
|
||||
const reproductionArtifact = artifact(paths.reproduction);
|
||||
|
||||
if (evaluation.input.sampling_sha256 !== samplingArtifact.sha256) {
|
||||
throw new Error("evaluation → sampling hash contract failed");
|
||||
}
|
||||
if (reproduction.formal.sha256 !== samplingArtifact.sha256) {
|
||||
throw new Error("reproduction → formal hash contract failed");
|
||||
}
|
||||
if (reproduction.rerun.sha256 !== rerunArtifact.sha256) {
|
||||
throw new Error("reproduction → rerun hash contract failed");
|
||||
}
|
||||
if (
|
||||
reproduction.summary.cells !== 64
|
||||
|| reproduction.summary.all_preregistered_fields_exact !== 64
|
||||
|| Object.values(reproduction.summary.by_field).some(
|
||||
(value) => value !== 64,
|
||||
)
|
||||
) {
|
||||
throw new Error("64-cell reproduction gate failed");
|
||||
}
|
||||
if (
|
||||
analysis.inputs.sampling.sha256 !== samplingArtifact.sha256
|
||||
|| analysis.inputs.evaluation.sha256 !== evaluationArtifact.sha256
|
||||
|| analysis.inputs.reproduction.sha256 !== reproductionArtifact.sha256
|
||||
) {
|
||||
throw new Error("analysis input hash chain failed");
|
||||
}
|
||||
if (
|
||||
analysis.contract.sources !== 16
|
||||
|| analysis.contract.sources_per_domain !== 4
|
||||
|| analysis.contract.seeds_per_source_condition !== 4
|
||||
|| analysis.contract.outputs !== 256
|
||||
) {
|
||||
throw new Error("source-blocked analysis grid contract failed");
|
||||
}
|
||||
|
||||
const conditions = analysis.contract.conditions;
|
||||
const conditionLabels = {
|
||||
s0_eos: "S0 · EOS",
|
||||
s1_eos: "S1 · EOS",
|
||||
s0_period: "S0 · 句点",
|
||||
s1_period: "S1 · 句点",
|
||||
};
|
||||
const domainLabels = {
|
||||
english: "English · WikiText-2",
|
||||
chinese: "中文 · TNEWS",
|
||||
code: "Code · HumanEval",
|
||||
math: "Math · GSM8K",
|
||||
};
|
||||
const sourceLabel = (sourceId) => {
|
||||
if (sourceId.startsWith("HumanEval/")) {
|
||||
return sourceId;
|
||||
}
|
||||
if (sourceId.startsWith("gsm8k/")) {
|
||||
return `GSM8K/${sourceId.split("/").at(-1)}`;
|
||||
}
|
||||
if (sourceId.startsWith("tnews/")) {
|
||||
return `TNEWS/${sourceId.split("/").at(-1)}`;
|
||||
}
|
||||
if (sourceId.startsWith("wikitext2/")) {
|
||||
return `WikiText/${sourceId.split("/").at(-1)}`;
|
||||
}
|
||||
return sourceId;
|
||||
};
|
||||
|
||||
const sourceMeta = Object.fromEntries(
|
||||
Object.entries(analysis.source_metadata).map(([sourceId, row]) => [
|
||||
sourceId,
|
||||
{
|
||||
...row,
|
||||
label: sourceLabel(sourceId),
|
||||
},
|
||||
]),
|
||||
);
|
||||
|
||||
const evalRows = evaluation.rows.map((row) => {
|
||||
let task = null;
|
||||
if (row.domain === "math") {
|
||||
task = {
|
||||
predicted: row.task_evaluation.predicted_final,
|
||||
gold: row.task_evaluation.gold_final,
|
||||
covered: row.task_evaluation.evaluator_covered,
|
||||
fixedPass: row.task_evaluation.fixed_budget_numeric_exact,
|
||||
strictPass: row.task_evaluation.strict_complete_numeric_exact,
|
||||
method: row.task_evaluation.extraction_method,
|
||||
};
|
||||
} else if (row.domain === "code") {
|
||||
task = {
|
||||
ast: row.task_evaluation.python_ast_parse,
|
||||
executed: row.task_evaluation.execution.status !== "not_run",
|
||||
status: row.task_evaluation.execution.status,
|
||||
fixedPass: row.task_evaluation.fixed_budget_tests_pass,
|
||||
strictPass: row.task_evaluation.strict_complete_tests_pass,
|
||||
candidateHash: row.task_evaluation.candidate_sha256,
|
||||
cacheHit: row.task_evaluation.execution_cache_hit,
|
||||
};
|
||||
}
|
||||
return {
|
||||
sourceId: row.source_id,
|
||||
domain: row.domain,
|
||||
condition: row.condition,
|
||||
replicate: row.replicate_label,
|
||||
baseSeed: row.base_seed,
|
||||
generatedTokens: row.generated_tokens,
|
||||
hitEos: row.hit_eos,
|
||||
truncated: row.stopped_at_max_new_tokens,
|
||||
trajectoryHash: row.generated_token_ids_sha256,
|
||||
completionClass: row.completion_class,
|
||||
task,
|
||||
};
|
||||
});
|
||||
|
||||
const taskFailures = evalRows
|
||||
.filter((row) => row.task && !row.task.strictPass)
|
||||
.map((row) => ({
|
||||
sourceId: row.sourceId,
|
||||
sourceLabel: sourceLabel(row.sourceId),
|
||||
domain: row.domain,
|
||||
condition: row.condition,
|
||||
replicate: row.replicate,
|
||||
generatedTokens: row.generatedTokens,
|
||||
hitEos: row.hitEos,
|
||||
completionClass: row.completionClass,
|
||||
detail: row.domain === "math"
|
||||
? {
|
||||
predicted: row.task.predicted,
|
||||
gold: row.task.gold,
|
||||
failure: "wrong_numeric_answer",
|
||||
}
|
||||
: {
|
||||
status: row.task.status,
|
||||
candidateHash: row.task.candidateHash,
|
||||
failure: row.task.status,
|
||||
},
|
||||
}));
|
||||
|
||||
const taskTotals = Object.fromEntries(
|
||||
Object.entries(analysis.task_matrix).map(([domain, tasks]) => [
|
||||
domain,
|
||||
tasks.map((task) => ({
|
||||
sourceId: task.source_id,
|
||||
label: sourceLabel(task.source_id),
|
||||
withinDomainIndex: task.within_domain_index,
|
||||
conditions: Object.fromEntries(
|
||||
Object.entries(task.conditions).map(([condition, cell]) => [
|
||||
condition,
|
||||
{
|
||||
pass: cell.strict_complete_success,
|
||||
outputs: 4,
|
||||
anyPass: cell.observed_any_strict_complete_pass,
|
||||
naturalEos: cell.natural_eos,
|
||||
meanGeneratedTokens: cell.generated_tokens.mean,
|
||||
minGeneratedTokens: cell.generated_tokens.min,
|
||||
maxGeneratedTokens: cell.generated_tokens.max,
|
||||
},
|
||||
]),
|
||||
),
|
||||
totalPass: Object.values(task.conditions).reduce(
|
||||
(total, cell) => total + cell.strict_complete_success,
|
||||
0,
|
||||
),
|
||||
totalOutputs: 16,
|
||||
})),
|
||||
]),
|
||||
);
|
||||
|
||||
const periodDirection = Object.fromEntries(
|
||||
Object.entries(analysis.prior_direction_check).map(([domain, row]) => {
|
||||
const values = Object.values(
|
||||
row.period_shortens_mean_tokens.by_source,
|
||||
);
|
||||
return [domain, {
|
||||
shorter: values.filter(Boolean).length,
|
||||
sources: values.length,
|
||||
bySource: row.period_shortens_mean_tokens.by_source,
|
||||
}];
|
||||
}),
|
||||
);
|
||||
const systemEosDirection = Object.fromEntries(
|
||||
Object.entries(analysis.prior_direction_check).map(([domain, row]) => {
|
||||
const values = Object.values(
|
||||
row.system_on_raises_natural_eos_rate.by_source,
|
||||
);
|
||||
return [domain, {
|
||||
raises: values.filter(Boolean).length,
|
||||
sources: values.length,
|
||||
bySource: row.system_on_raises_natural_eos_rate.by_source,
|
||||
}];
|
||||
}),
|
||||
);
|
||||
|
||||
const result = {
|
||||
schemaVersion: 1,
|
||||
capturedAt: sampling.captured_at,
|
||||
contract: {
|
||||
protocolId: sampling.protocol_id,
|
||||
model: sampling.model.repo,
|
||||
revision: sampling.model.revision,
|
||||
conditions,
|
||||
conditionLabels,
|
||||
domains: Object.keys(analysis.contract.sources_by_domain),
|
||||
domainLabels,
|
||||
sources: analysis.contract.sources,
|
||||
sourcesPerDomain: analysis.contract.sources_per_domain,
|
||||
seedsPerCell: analysis.contract.seeds_per_source_condition,
|
||||
outputs: analysis.contract.outputs,
|
||||
baseSeeds: sampling.seed_contract.executed_base_seeds,
|
||||
decode: {
|
||||
doSample: sampling.generation_contract.do_sample,
|
||||
temperature: sampling.generation_contract.temperature,
|
||||
topP: sampling.generation_contract.top_p,
|
||||
topK: sampling.generation_contract.top_k,
|
||||
maxNewTokens: sampling.generation_contract.max_new_tokens,
|
||||
},
|
||||
sourceIsPrimaryCoverageUnit: (
|
||||
analysis.contract.source_is_primary_coverage_unit
|
||||
),
|
||||
seedIsWithinSourceRepeat: (
|
||||
analysis.contract.seed_is_within_source_repeat
|
||||
),
|
||||
noPopulationConfidenceIntervals: (
|
||||
analysis.contract.no_population_confidence_intervals
|
||||
),
|
||||
noPValues: analysis.contract.no_p_values,
|
||||
batchSeedAlignedNotCommonRandomNumbers: (
|
||||
sampling.seed_contract.batch_seed_aligned_not_common_random_numbers
|
||||
),
|
||||
},
|
||||
headline: {
|
||||
...analysis.overall,
|
||||
promptHashExact: sampling.source_contract.prompt_hash_audit.exact,
|
||||
promptHashCells: sampling.source_contract.prompt_hash_audit.cells,
|
||||
firstTwoSeedComparableCells: (
|
||||
sampling.summary.first_two_seeds.comparable_cells
|
||||
),
|
||||
firstTwoSeedDifferentTrajectories: (
|
||||
sampling.summary.first_two_seeds.different_trajectories
|
||||
),
|
||||
reproducedCells: (
|
||||
reproduction.summary.all_preregistered_fields_exact
|
||||
),
|
||||
reproductionCells: reproduction.summary.cells,
|
||||
codeAstParse: evaluation.summary.code.ast_parse,
|
||||
codeExecuted: evaluation.summary.code.executed,
|
||||
codeStatuses: evaluation.summary.code.execution_statuses,
|
||||
codeUniqueExecutionKeys: (
|
||||
evaluation.sandbox.unique_code_cache_entries
|
||||
),
|
||||
},
|
||||
sourceMeta,
|
||||
sourcesByDomain: analysis.contract.sources_by_domain,
|
||||
sourceCells: analysis.source_condition_cells,
|
||||
domainConditions: analysis.domain_condition_summary,
|
||||
sourceContrasts: analysis.source_contrasts,
|
||||
domainContrasts: analysis.domain_contrasts,
|
||||
periodDirection,
|
||||
systemEosDirection,
|
||||
taskTotals,
|
||||
taskFailures,
|
||||
evalRows,
|
||||
reproduction: reproduction.summary,
|
||||
artifacts: {
|
||||
sampling: samplingArtifact,
|
||||
evaluation: evaluationArtifact,
|
||||
rerun: rerunArtifact,
|
||||
reproduction: reproductionArtifact,
|
||||
analysis: artifact(paths.analysis),
|
||||
},
|
||||
execution: {
|
||||
generationSeconds: sampling.sources.reduce(
|
||||
(total, source) => total + source.runs.reduce(
|
||||
(subtotal, run) => subtotal + run.generation_seconds,
|
||||
0,
|
||||
),
|
||||
0,
|
||||
),
|
||||
peakCudaMemoryAllocatedBytes: (
|
||||
sampling.execution.peak_cuda_memory_allocated_bytes
|
||||
),
|
||||
torch: sampling.execution.torch,
|
||||
transformers: sampling.execution.transformers,
|
||||
deviceMap: sampling.execution.device_map,
|
||||
sandbox: evaluation.sandbox,
|
||||
},
|
||||
claimBoundary: [
|
||||
...sampling.claim_boundary,
|
||||
...evaluation.claim_boundary,
|
||||
...analysis.claim_boundary,
|
||||
...reproduction.claim_boundary,
|
||||
],
|
||||
};
|
||||
|
||||
writeFileSync(paths.output, `${JSON.stringify(result, null, 2)}\n`);
|
||||
console.log(JSON.stringify({
|
||||
output: paths.output,
|
||||
...artifact(paths.output),
|
||||
headline: result.headline,
|
||||
periodDirection: result.periodDirection,
|
||||
}, null, 2));
|
||||
@@ -82,6 +82,8 @@ const overview = await evaluate(`(() => ({
|
||||
behaviorEdges: document.querySelectorAll("[data-behavior-map-edge]").length,
|
||||
completionDepthTabs: document.querySelectorAll("[data-cd-tab]").length,
|
||||
completionDepthPanels: document.querySelectorAll("[data-cd-panel]").length,
|
||||
crossSourceTabs: document.querySelectorAll("[data-cs-tab]").length,
|
||||
crossSourcePanels: document.querySelectorAll("[data-cs-panel]").length,
|
||||
hiddenStages: document.querySelectorAll("[data-hidden-stage]").length,
|
||||
routerLayers: document.querySelectorAll("[data-router-layer]").length,
|
||||
branches: document.querySelectorAll(".branch-grid > a").length,
|
||||
@@ -1210,14 +1212,15 @@ console.log(JSON.stringify(report, null, 2));
|
||||
const numeric = (text) => Number.parseFloat(text.replaceAll(",", ""));
|
||||
const failures = [];
|
||||
if (!overview.title.includes("为什么转向")) failures.push("专题标题异常");
|
||||
if (overview.sections !== 29 || overview.tocLinks !== 29) failures.push("二十八个编号专题加阅读链的目录结构异常");
|
||||
if (overview.sections !== 30 || overview.tocLinks !== 30) failures.push("二十九个编号专题加阅读链的目录结构异常");
|
||||
if (overview.ledgers !== 24 || overview.waves !== 10) failures.push("二十四张问题账或十次转向结构异常");
|
||||
if (overview.paperLinks !== 60 || overview.branches !== 5 || overview.followups !== 1) failures.push("论文链、旁支或公开后续标记异常");
|
||||
if (overview.labTabs !== 4 || overview.labPanels !== 4) failures.push("四联实验结构异常");
|
||||
if (overview.artifactTabs !== 13 || overview.artifactPanels !== 13 || overview.artifactLayers !== 27) failures.push("真实权重十三联实验结构异常");
|
||||
if (overview.behaviorTabs !== 4 || overview.behaviorPanels !== 4 || overview.behaviorSources !== 16 || overview.behaviorEdges !== 10) failures.push("Chat 行为实验结构异常");
|
||||
if (overview.completionDepthTabs !== 4 || overview.completionDepthPanels !== 4 || overview.hiddenStages !== 29 || overview.routerLayers !== 26) failures.push("Chat 完成度与全深度实验结构异常");
|
||||
if (overview.heroLabs !== "20 个可操作实验") failures.push("DeepSeek 实验总数账异常");
|
||||
if (overview.crossSourceTabs !== 4 || overview.crossSourcePanels !== 4) failures.push("跨来源采样实验结构异常");
|
||||
if (overview.heroLabs !== "21 个可操作实验") failures.push("DeepSeek 实验总数账异常");
|
||||
if (overview.navLinks !== 20 || home.navLinks !== 20 || mobile.mobileLinks !== 20 || overview.activeNav !== "DeepSeek") failures.push("全站导航未同步 DeepSeek");
|
||||
if (overview.documentOverflow > 1 || mobile.documentOverflow > 1) failures.push("桌面或移动端存在文档级横向溢出");
|
||||
if (capacity.initial.panel !== "capacity" || capacity.initial.total !== "32.1× FFN" || capacity.initial.active !== "1.13× FFN") failures.push("V3 稀疏容量初始账异常");
|
||||
|
||||
@@ -0,0 +1,314 @@
|
||||
import { writeFileSync } from "node:fs";
|
||||
|
||||
const cdpPort = process.env.CDP_PORT ?? "9230";
|
||||
const baseUrl = process.env.SITE_URL ?? "http://127.0.0.1:4327";
|
||||
const pages = await fetch(`http://127.0.0.1:${cdpPort}/json/list`)
|
||||
.then((response) => response.json());
|
||||
const page = pages.find((entry) => entry.type === "page");
|
||||
if (!page) throw new Error(`CDP ${cdpPort} 没有可用页面`);
|
||||
|
||||
const socket = new WebSocket(page.webSocketDebuggerUrl);
|
||||
await new Promise((resolve, reject) => {
|
||||
socket.addEventListener("open", resolve, { once: true });
|
||||
socket.addEventListener("error", reject, { once: true });
|
||||
});
|
||||
|
||||
let nextId = 0;
|
||||
const pending = new Map();
|
||||
const exceptions = [];
|
||||
socket.addEventListener("message", (event) => {
|
||||
const message = JSON.parse(event.data);
|
||||
if (message.id && pending.has(message.id)) {
|
||||
const { resolve, reject } = pending.get(message.id);
|
||||
pending.delete(message.id);
|
||||
if (message.error) reject(new Error(message.error.message));
|
||||
else resolve(message.result);
|
||||
}
|
||||
if (message.method === "Runtime.exceptionThrown") {
|
||||
exceptions.push(
|
||||
message.params.exceptionDetails.exception?.description
|
||||
?? message.params.exceptionDetails.text,
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
const command = (method, params = {}) => new Promise((resolve, reject) => {
|
||||
const id = ++nextId;
|
||||
pending.set(id, { resolve, reject });
|
||||
socket.send(JSON.stringify({ id, method, params }));
|
||||
});
|
||||
const pause = (milliseconds) => new Promise(
|
||||
(resolve) => setTimeout(resolve, milliseconds),
|
||||
);
|
||||
const evaluate = async (expression) => {
|
||||
const result = await command("Runtime.evaluate", {
|
||||
expression,
|
||||
returnByValue: true,
|
||||
awaitPromise: true,
|
||||
});
|
||||
if (result.exceptionDetails) {
|
||||
throw new Error(
|
||||
result.exceptionDetails.exception?.description
|
||||
?? result.exceptionDetails.text,
|
||||
);
|
||||
}
|
||||
return result.result.value;
|
||||
};
|
||||
const screenshot = async (path) => {
|
||||
const result = await command("Page.captureScreenshot", {
|
||||
format: "png",
|
||||
captureBeyondViewport: false,
|
||||
});
|
||||
writeFileSync(path, Buffer.from(result.data, "base64"));
|
||||
};
|
||||
const assert = (condition, message) => {
|
||||
if (!condition) throw new Error(message);
|
||||
};
|
||||
|
||||
await command("Page.enable");
|
||||
await command("Runtime.enable");
|
||||
await command("Emulation.setDeviceMetricsOverride", {
|
||||
width: 1440,
|
||||
height: 1100,
|
||||
deviceScaleFactor: 1,
|
||||
mobile: false,
|
||||
});
|
||||
await command("Page.navigate", { url: `${baseUrl}/deepseek/` });
|
||||
for (let attempt = 0; attempt < 100; attempt += 1) {
|
||||
await pause(100);
|
||||
if (await evaluate("document.readyState === 'complete'")) break;
|
||||
}
|
||||
|
||||
const overview = await evaluate(`(() => {
|
||||
const root = document.querySelector("[data-cross-source-lab]");
|
||||
if (!root) return null;
|
||||
document.documentElement.style.scrollBehavior = "auto";
|
||||
window.scrollTo(0, root.getBoundingClientRect().top + window.scrollY);
|
||||
const facts = Object.fromEntries(
|
||||
[...root.querySelectorAll(".cs-ledger article")].map((node) => [
|
||||
node.querySelector("span").textContent.trim(),
|
||||
node.querySelector("b").textContent.trim(),
|
||||
]),
|
||||
);
|
||||
return {
|
||||
tabs: root.querySelectorAll("[data-cs-tab]").length,
|
||||
panels: root.querySelectorAll("[data-cs-panel]").length,
|
||||
active: root.querySelector("[data-cs-panel]:not([hidden])")?.dataset.csPanel,
|
||||
sourceCards: root.querySelectorAll("[data-cs-source-cards] > article").length,
|
||||
seedDots: root.querySelectorAll(".seed-dots > i").length,
|
||||
outputCount: root.querySelector("[data-cs-output-count]").textContent.trim(),
|
||||
naturalEos: root.querySelector("[data-cs-natural-eos]").textContent.trim(),
|
||||
facts,
|
||||
labs: [...document.querySelectorAll(".page-facts > div")]
|
||||
.find((node) => node.querySelector("dt")?.textContent.trim() === "LABS")
|
||||
?.querySelector("dd")?.textContent.trim(),
|
||||
status: [...document.querySelectorAll(".page-facts > div")]
|
||||
.find((node) => node.querySelector("dt")?.textContent.trim() === "STATUS")
|
||||
?.querySelector("dd")?.textContent.trim(),
|
||||
overflow: document.documentElement.scrollWidth
|
||||
- document.documentElement.clientWidth,
|
||||
};
|
||||
})()`);
|
||||
await pause(250);
|
||||
await screenshot("/tmp/llm-atlas-cross-source-desktop.png");
|
||||
|
||||
const hierarchySwitch = await evaluate(`(() => {
|
||||
const root = document.querySelector("[data-cross-source-lab]");
|
||||
const domain = root.querySelector("[data-cs-domain]");
|
||||
const condition = root.querySelector("[data-cs-condition]");
|
||||
domain.value = "code";
|
||||
domain.dispatchEvent(new Event("change", { bubbles: true }));
|
||||
condition.value = "s1_period";
|
||||
condition.dispatchEvent(new Event("change", { bubbles: true }));
|
||||
return {
|
||||
cards: [...root.querySelectorAll("[data-cs-source-cards] > article")].map((node) => ({
|
||||
id: node.dataset.sourceId,
|
||||
eos: node.querySelector("header em").textContent.trim(),
|
||||
footer: node.querySelector(":scope > p").textContent.trim(),
|
||||
dots: node.querySelectorAll(".seed-dots > i").length,
|
||||
})),
|
||||
outputCount: root.querySelector("[data-cs-output-count]").textContent.trim(),
|
||||
naturalEos: root.querySelector("[data-cs-natural-eos]").textContent.trim(),
|
||||
};
|
||||
})()`);
|
||||
|
||||
const taskMatrices = await evaluate(`(() => {
|
||||
const root = document.querySelector("[data-cross-source-lab]");
|
||||
root.querySelector('[data-cs-tab="tasks"]').click();
|
||||
const read = () => ({
|
||||
active: root.querySelector("[data-cs-panel]:not([hidden])").dataset.csPanel,
|
||||
total: root.querySelector("[data-cs-task-total]").textContent.trim(),
|
||||
rows: [...root.querySelectorAll("[data-cs-task-matrix] > div")].map((row) => ({
|
||||
id: row.dataset.taskId,
|
||||
cells: [...row.querySelectorAll("b")].map((cell) => cell.textContent.trim()),
|
||||
total: row.querySelector("strong").textContent.trim(),
|
||||
})),
|
||||
range: root.querySelector("[data-cs-task-range]").textContent.trim(),
|
||||
interaction: root.querySelector("[data-cs-task-interaction]").textContent.trim(),
|
||||
failure: root.querySelector("[data-cs-task-failure]").textContent.trim(),
|
||||
});
|
||||
const math = read();
|
||||
root.querySelector('[data-cs-task-domain="code"]').click();
|
||||
const code = read();
|
||||
return { math, code };
|
||||
})()`);
|
||||
await pause(150);
|
||||
await screenshot("/tmp/llm-atlas-cross-source-task-matrix.png");
|
||||
|
||||
const directions = await evaluate(`(() => {
|
||||
const root = document.querySelector("[data-cross-source-lab]");
|
||||
root.querySelector('[data-cs-tab="directions"]').click();
|
||||
const select = root.querySelector("[data-cs-direction-domain]");
|
||||
const read = () => ({
|
||||
active: root.querySelector("[data-cs-panel]:not([hidden])").dataset.csPanel,
|
||||
shorter: root.querySelector("[data-cs-shorter]").textContent.trim(),
|
||||
mean: root.querySelector("[data-cs-domain-mean]").textContent.trim(),
|
||||
median: root.querySelector("[data-cs-domain-median]").textContent.trim(),
|
||||
rows: [...root.querySelectorAll("[data-cs-direction-rows] > article")].map((row) => ({
|
||||
id: row.dataset.directionSource,
|
||||
value: row.querySelector("b").textContent.trim(),
|
||||
dotClass: row.querySelector("u").className,
|
||||
})),
|
||||
});
|
||||
select.value = "english";
|
||||
select.dispatchEvent(new Event("change", { bubbles: true }));
|
||||
const english = read();
|
||||
select.value = "code";
|
||||
select.dispatchEvent(new Event("change", { bubbles: true }));
|
||||
const code = read();
|
||||
return { english, code };
|
||||
})()`);
|
||||
|
||||
const keyboard = await evaluate(`(() => {
|
||||
const root = document.querySelector("[data-cross-source-lab]");
|
||||
const first = root.querySelector('[data-cs-tab="hierarchy"]');
|
||||
first.click();
|
||||
first.focus();
|
||||
first.dispatchEvent(new KeyboardEvent("keydown", {
|
||||
key: "ArrowRight", bubbles: true,
|
||||
}));
|
||||
return {
|
||||
selected: root.querySelector('[data-cs-tab][aria-selected="true"]').dataset.csTab,
|
||||
active: root.querySelector("[data-cs-panel]:not([hidden])").dataset.csPanel,
|
||||
focused: document.activeElement.dataset.csTab,
|
||||
};
|
||||
})()`);
|
||||
|
||||
const reproduction = await evaluate(`(() => {
|
||||
const root = document.querySelector("[data-cross-source-lab]");
|
||||
root.querySelector('[data-cs-tab="reproduction"]').click();
|
||||
return {
|
||||
fields: [...root.querySelectorAll(".repro-fields article")].map((node) => (
|
||||
node.querySelector("b").textContent.trim()
|
||||
)),
|
||||
artifacts: [...root.querySelectorAll(".hash-chain article")].map((node) => ({
|
||||
label: node.querySelector("span").textContent.trim(),
|
||||
hash: node.querySelector("b").textContent.trim(),
|
||||
})),
|
||||
active: root.querySelector("[data-cs-panel]:not([hidden])").dataset.csPanel,
|
||||
};
|
||||
})()`);
|
||||
|
||||
await command("Emulation.setDeviceMetricsOverride", {
|
||||
width: 390,
|
||||
height: 844,
|
||||
deviceScaleFactor: 1,
|
||||
mobile: true,
|
||||
});
|
||||
await pause(300);
|
||||
const mobile = await evaluate(`(() => {
|
||||
const root = document.querySelector("[data-cross-source-lab]");
|
||||
root.querySelector('[data-cs-tab="tasks"]').click();
|
||||
root.scrollIntoView();
|
||||
return {
|
||||
documentOverflow: document.documentElement.scrollWidth
|
||||
- document.documentElement.clientWidth,
|
||||
rootOverflow: root.scrollWidth - root.clientWidth,
|
||||
matrixWidth: root.querySelector("[data-cs-task-matrix]").getBoundingClientRect().width,
|
||||
viewport: window.innerWidth,
|
||||
};
|
||||
})()`);
|
||||
await screenshot("/tmp/llm-atlas-cross-source-mobile.png");
|
||||
|
||||
assert(overview, "找不到跨来源采样实验");
|
||||
assert(overview.tabs === 4 && overview.panels === 4, "四页签/面板合同异常");
|
||||
assert(overview.active === "hierarchy", "初始面板不是 hierarchy");
|
||||
assert(overview.sourceCards === 4 && overview.seedDots === 16, "source/seed 层级渲染异常");
|
||||
assert(overview.outputCount === "16 / 16", "初始输出总账异常");
|
||||
assert(overview.facts.SOURCES === "16", "source headline 异常");
|
||||
assert(overview.facts["SAMPLED OUTPUTS"] === "256", "output headline 异常");
|
||||
assert(overview.facts["NATURAL EOS"] === "250 / 256", "EOS headline 异常");
|
||||
assert(overview.facts["MATH · STRICT"] === "47 / 64", "Math headline 异常");
|
||||
assert(overview.facts["CODE · TESTS"] === "52 / 64", "Code headline 异常");
|
||||
assert(overview.labs === "21 个可操作实验", "DeepSeek LABS 总账异常");
|
||||
assert(overview.status === "七轮 · 512 条采样", "DeepSeek STATUS 总账异常");
|
||||
assert(overview.overflow <= 1, `桌面横向溢出 ${overview.overflow}px`);
|
||||
|
||||
assert(hierarchySwitch.cards.length === 4, "Code source cards 数量异常");
|
||||
assert(hierarchySwitch.cards.every((row) => row.dots === 4), "每题 seed 数不为 4");
|
||||
assert(
|
||||
hierarchySwitch.cards.find((row) => row.id === "HumanEval/133")
|
||||
?.footer.includes("0 / 4 strict task pass"),
|
||||
"HumanEval/133 × s1_period 应为 0/4",
|
||||
);
|
||||
assert(hierarchySwitch.naturalEos === "16 / 16", "Code s1_period EOS 异常");
|
||||
|
||||
assert(taskMatrices.math.total === "47 / 64", "Math total 异常");
|
||||
assert(
|
||||
taskMatrices.math.rows.map((row) => row.cells.join(",")).join("|")
|
||||
=== "3 / 4,4 / 4,3 / 4,4 / 4|4 / 4,4 / 4,4 / 4,4 / 4|2 / 4,2 / 4,3 / 4,1 / 4|1 / 4,4 / 4,1 / 4,3 / 4",
|
||||
"Math 4×4 pass matrix 异常",
|
||||
);
|
||||
assert(taskMatrices.math.range === "8 → 16 / 16", "Math source range 异常");
|
||||
assert(taskMatrices.math.failure === "17 / 64", "Math failure 账异常");
|
||||
assert(taskMatrices.code.total === "52 / 64", "Code total 异常");
|
||||
assert(
|
||||
taskMatrices.code.rows.map((row) => row.cells.join(",")).join("|")
|
||||
=== "4 / 4,4 / 4,4 / 4,4 / 4|2 / 4,1 / 4,1 / 4,4 / 4|4 / 4,4 / 4,4 / 4,0 / 4|4 / 4,4 / 4,4 / 4,4 / 4",
|
||||
"Code 4×4 pass matrix 异常",
|
||||
);
|
||||
assert(taskMatrices.code.interaction === "1↑ · 2= · 1↓", "Code interaction 方向异常");
|
||||
assert(taskMatrices.code.failure === "12 / 64", "Code failure 账异常");
|
||||
|
||||
assert(directions.english.shorter === "1 / 4", "English 方向数异常");
|
||||
assert(directions.english.mean === "−8.5 tokens".replace("−", "-"), "English mean 异常");
|
||||
assert(directions.english.median === "+21.2 tokens", "English median 异常");
|
||||
assert(
|
||||
directions.english.rows.map((row) => row.value).join("|")
|
||||
=== "-121.0 tokens|+44.5 tokens|+41.0 tokens|+1.4 tokens",
|
||||
"English source contrast 异常",
|
||||
);
|
||||
assert(directions.code.shorter === "4 / 4", "Code 方向数异常");
|
||||
assert(directions.code.rows.every((row) => row.dotClass === "negative"), "Code 应四条全负");
|
||||
|
||||
assert(
|
||||
keyboard.selected === "tasks"
|
||||
&& keyboard.active === "tasks"
|
||||
&& keyboard.focused === "tasks",
|
||||
"页签键盘导航异常",
|
||||
);
|
||||
assert(reproduction.active === "reproduction", "复现面板切换异常");
|
||||
assert(reproduction.fields.length === 8, "复现字段数不为 8");
|
||||
assert(reproduction.fields.every((value) => value === "64 / 64"), "复现字段未全部 exact");
|
||||
assert(reproduction.artifacts.length === 5, "hash chain 工件数不为 5");
|
||||
assert(mobile.documentOverflow <= 1, `移动端 document 横向溢出 ${mobile.documentOverflow}px`);
|
||||
assert(mobile.rootOverflow <= 1, `移动端实验横向溢出 ${mobile.rootOverflow}px`);
|
||||
assert(mobile.matrixWidth <= mobile.viewport, "移动端任务矩阵超出 viewport");
|
||||
assert(exceptions.length === 0, `浏览器异常:${exceptions.join(" | ")}`);
|
||||
|
||||
console.log(JSON.stringify({
|
||||
overview,
|
||||
hierarchySwitch,
|
||||
taskMatrices,
|
||||
directions,
|
||||
keyboard,
|
||||
reproduction,
|
||||
mobile,
|
||||
screenshots: [
|
||||
"/tmp/llm-atlas-cross-source-desktop.png",
|
||||
"/tmp/llm-atlas-cross-source-task-matrix.png",
|
||||
"/tmp/llm-atlas-cross-source-mobile.png",
|
||||
],
|
||||
}, null, 2));
|
||||
|
||||
socket.close();
|
||||
@@ -0,0 +1,670 @@
|
||||
---
|
||||
import rawLab from "@/data/deepseek-v2-lite-chat-cross-source-sampling-compact.json";
|
||||
|
||||
const lab = rawLab as any;
|
||||
const json = JSON.stringify(lab).replaceAll("<", "\\u003c");
|
||||
const domains = lab.contract.domains as string[];
|
||||
const conditions = lab.contract.conditions as string[];
|
||||
---
|
||||
|
||||
<figure class="cross-source-lab" data-cross-source-lab>
|
||||
<header class="cs-head">
|
||||
<div>
|
||||
<p>ROUND 07 / SOURCE-BLOCKED FOLLOW-UP</p>
|
||||
<h3>seed 不是题目:把同题重复与跨题覆盖拆成两层</h3>
|
||||
</div>
|
||||
<p>
|
||||
总预算仍是 256 条,但从上一轮的 <code>4 sources × 8 seeds × 8 cells</code>
|
||||
改成 <code>16 sources × 4 seeds × 4 cells</code>。source 是覆盖单位,seed
|
||||
只是题内重复;所有方向都先逐 source 计算,再描述四条 source 是否同向。
|
||||
</p>
|
||||
</header>
|
||||
|
||||
<div class="cs-ledger">
|
||||
<article><span>SOURCES</span><b>16</b><p>四域各 4 条,结果前冻结</p></article>
|
||||
<article><span>SAMPLED OUTPUTS</span><b>256</b><p>16 × 4 seeds × 4 cells</p></article>
|
||||
<article class="pass"><span>NATURAL EOS</span><b>250 / 256</b><p>6 条在 512-token 触顶</p></article>
|
||||
<article><span>MATH · STRICT</span><b>47 / 64</b><p>四道 GSM8K,不是同题 64 次</p></article>
|
||||
<article><span>CODE · TESTS</span><b>52 / 64</b><p>四道 HumanEval,沙箱执行</p></article>
|
||||
<article class="pass"><span>FRESH PROCESS</span><b>64 / 64</b><p>八项合同逐字段 exact</p></article>
|
||||
</div>
|
||||
|
||||
<div class="cs-tabs" role="tablist" aria-label="选择跨来源采样实验视图">
|
||||
<button type="button" role="tab" data-cs-tab="hierarchy" aria-selected="true">
|
||||
<span>01</span><b>先分清 source 与 seed</b><small>coverage hierarchy</small>
|
||||
</button>
|
||||
<button type="button" role="tab" data-cs-tab="tasks" aria-selected="false" tabindex="-1">
|
||||
<span>02</span><b>总分怎样藏住题目差异</b><small>task × condition matrix</small>
|
||||
</button>
|
||||
<button type="button" role="tab" data-cs-tab="directions" aria-selected="false" tabindex="-1">
|
||||
<span>03</span><b>平均值为什么会误导</b><small>source-level contrasts</small>
|
||||
</button>
|
||||
<button type="button" role="tab" data-cs-tab="reproduction" aria-selected="false" tabindex="-1">
|
||||
<span>04</span><b>随机轨迹如何被审计</b><small>hash chain + replay</small>
|
||||
</button>
|
||||
</div>
|
||||
|
||||
<section class="cs-panel" data-cs-panel="hierarchy">
|
||||
<div class="cs-panel-lead">
|
||||
<div><span>I / COVERAGE HIERARCHY</span><h4>64 个 seed 输出,不会自动变成 64 道题</h4></div>
|
||||
<p>
|
||||
先选择一个域与条件。每张卡是一条独立 source,四个圆点才是该题内的四次采样。
|
||||
圆点高度表示生成长度,颜色表示自然 EOS 或预算截断。
|
||||
</p>
|
||||
</div>
|
||||
|
||||
<div class="unit-equation" aria-label="实验单位层级">
|
||||
<article><span>DOMAIN</span><b>4 类语料</b><p>English · 中文 · Code · Math</p></article>
|
||||
<i>→</i>
|
||||
<article class="primary"><span>PRIMARY COVERAGE UNIT</span><b>4 sources / 域</b><p>跨题方向只在这一层数</p></article>
|
||||
<i>→</i>
|
||||
<article><span>WITHIN-SOURCE REPEAT</span><b>4 seeds / 格</b><p>观察同题采样离散,不扩充题数</p></article>
|
||||
</div>
|
||||
|
||||
<div class="cs-controls">
|
||||
<label><span>DOMAIN</span>
|
||||
<select data-cs-domain aria-label="选择跨来源域">
|
||||
{domains.map((domain) => (
|
||||
<option value={domain}>{lab.contract.domainLabels[domain]}</option>
|
||||
))}
|
||||
</select>
|
||||
</label>
|
||||
<label><span>CONDITION</span>
|
||||
<select data-cs-condition aria-label="选择跨来源 condition">
|
||||
{conditions.map((condition) => (
|
||||
<option value={condition}>{lab.contract.conditionLabels[condition]}</option>
|
||||
))}
|
||||
</select>
|
||||
</label>
|
||||
<div><span>FIXED DECODE</span><b>T .3 · P .95 · K 0 · CAP 512</b></div>
|
||||
</div>
|
||||
|
||||
<div class="source-summary">
|
||||
<article><span>OUTPUTS</span><b data-cs-output-count>—</b><p>4 sources × 4 seeds</p></article>
|
||||
<article><span>NATURAL EOS</span><b data-cs-natural-eos>—</b><p>micro count;不等于正确</p></article>
|
||||
<article><span>BETWEEN-SOURCE RANGE</span><b data-cs-between-range>—</b><p>四条 source 的平均长度跨度</p></article>
|
||||
<article><span>WITHIN-SOURCE RANGE</span><b data-cs-within-range>—</b><p>每题 seed 长度跨度的均值</p></article>
|
||||
</div>
|
||||
|
||||
<div class="source-card-grid" data-cs-source-cards></div>
|
||||
|
||||
<aside class="cs-note">
|
||||
<b>读法:先横向比较四张卡,再纵向看每张卡里的四个 seed</b>
|
||||
<p>
|
||||
如果四张卡差很多,增加同一题的 seed 只能更精细地描出那一道题,不能弥补 source
|
||||
覆盖不足。这里保留 micro 总账,但主结论使用 source-blocked 描述。
|
||||
</p>
|
||||
</aside>
|
||||
</section>
|
||||
|
||||
<section class="cs-panel" data-cs-panel="tasks" hidden>
|
||||
<div class="cs-panel-lead">
|
||||
<div><span>II / TASK MATRIX</span><h4>47 / 64 与 52 / 64 背后,是八条完全不同的题目轨迹</h4></div>
|
||||
<p>
|
||||
每格分母固定为四个 seed。绿色越深表示通过越多;自然结束、数值抽取、AST
|
||||
解析与官方测试各自记账,不用“看起来完成了”替代任务正确。
|
||||
</p>
|
||||
</div>
|
||||
|
||||
<div class="task-switch">
|
||||
<div role="group" aria-label="选择任务域">
|
||||
<button type="button" data-cs-task-domain="math" aria-pressed="true">MATH · GSM8K</button>
|
||||
<button type="button" data-cs-task-domain="code" aria-pressed="false">CODE · HUMANEVAL</button>
|
||||
</div>
|
||||
<p><span data-cs-task-total-label>MATH · STRICT</span><b data-cs-task-total>47 / 64</b></p>
|
||||
</div>
|
||||
|
||||
<div class="task-matrix" data-cs-task-matrix></div>
|
||||
|
||||
<div class="matrix-reading">
|
||||
<article>
|
||||
<span>题目差异</span>
|
||||
<b data-cs-task-range>—</b>
|
||||
<p data-cs-task-range-copy>—</p>
|
||||
</article>
|
||||
<article>
|
||||
<span>条件交互</span>
|
||||
<b data-cs-task-interaction>—</b>
|
||||
<p data-cs-task-interaction-copy>—</p>
|
||||
</article>
|
||||
<article>
|
||||
<span>失败身份</span>
|
||||
<b data-cs-task-failure>—</b>
|
||||
<p data-cs-task-failure-copy>—</p>
|
||||
</article>
|
||||
</div>
|
||||
|
||||
<aside class="cs-note dark">
|
||||
<b>最关键的反例:Code 的总体 interaction 恰好是 0</b>
|
||||
<p>
|
||||
但逐题看,HumanEval/44 是 +1.0,HumanEval/133 是 −1.0,另外两题为 0。
|
||||
“总体没有交互”在这里不是“每题都没有”,而是两道题正负抵消。
|
||||
</p>
|
||||
</aside>
|
||||
</section>
|
||||
|
||||
<section class="cs-panel" data-cs-panel="directions" hidden>
|
||||
<div class="cs-panel-lead">
|
||||
<div><span>III / SOURCE-LEVEL DIRECTIONS</span><h4>一个极端 source,可以让均值与多数方向相反</h4></div>
|
||||
<p>
|
||||
下图是每条 source 的“句点 − EOS”平均生成长度差。负数代表句点更短;
|
||||
四个点才是四个覆盖单位,域均值只放在旁边,不拿它替代点的方向。
|
||||
</p>
|
||||
</div>
|
||||
|
||||
<div class="direction-controls">
|
||||
<label><span>DOMAIN</span>
|
||||
<select data-cs-direction-domain aria-label="选择方向审计域">
|
||||
{domains.map((domain) => (
|
||||
<option value={domain}>{lab.contract.domainLabels[domain]}</option>
|
||||
))}
|
||||
</select>
|
||||
</label>
|
||||
<article><span>PERIOD SHORTER</span><b data-cs-shorter>—</b><p>四条 source 中负方向数量</p></article>
|
||||
<article><span>DOMAIN MEAN</span><b data-cs-domain-mean>—</b><p>tokens · period − EOS</p></article>
|
||||
<article><span>SOURCE MEDIAN</span><b data-cs-domain-median>—</b><p>比均值更不易被单点拖动</p></article>
|
||||
</div>
|
||||
|
||||
<div class="zero-axis">
|
||||
<header><span>PERIOD SHORTER</span><b>0 · NO LENGTH CHANGE</b><span>PERIOD LONGER</span></header>
|
||||
<div data-cs-direction-rows></div>
|
||||
</div>
|
||||
|
||||
<div class="round-correction">
|
||||
<article>
|
||||
<span>ROUND 06 · SINGLE ENGLISH SOURCE</span>
|
||||
<b>WikiText/0443:−121.0 tokens</b>
|
||||
<p>如果只看这一条,会得到“句点大幅缩短 English”的印象。</p>
|
||||
</article>
|
||||
<i>→</i>
|
||||
<article class="result">
|
||||
<span>ROUND 07 · FOUR ENGLISH SOURCES</span>
|
||||
<b>1 shorter · 3 longer</b>
|
||||
<p>另三条分别 +44.5、+41.0、+1.4;均值仍为 −8.5,只因首题权重大。</p>
|
||||
</article>
|
||||
</div>
|
||||
|
||||
<div class="direction-all">
|
||||
{domains.map((domain) => (
|
||||
<article>
|
||||
<span>{lab.contract.domainLabels[domain]}</span>
|
||||
<b>{lab.periodDirection[domain].shorter} / 4 shorter</b>
|
||||
<p>
|
||||
mean {
|
||||
lab.domainContrasts[domain].mean_generated_tokens
|
||||
.boundary_main_period_minus_eos.mean.toFixed(1)
|
||||
} · median {
|
||||
lab.domainContrasts[domain].mean_generated_tokens
|
||||
.boundary_main_period_minus_eos.median.toFixed(1)
|
||||
}
|
||||
</p>
|
||||
</article>
|
||||
))}
|
||||
</div>
|
||||
|
||||
<aside class="cs-note">
|
||||
<b>11 / 16 source 的句点格平均更短,但这不是总体显著性</b>
|
||||
<p>
|
||||
句点是受 Round 06 结果启发的定向复查,不是未见前序结果的盲确认;16 条 source
|
||||
也没有随机代表任何总体。因此不报告 p-value 或 population confidence interval。
|
||||
</p>
|
||||
</aside>
|
||||
</section>
|
||||
|
||||
<section class="cs-panel" data-cs-panel="reproduction" hidden>
|
||||
<div class="cs-panel-lead">
|
||||
<div><span>IV / REPLAY + HASH CHAIN</span><h4>正式生成、独立评测、分析与复跑彼此锁定</h4></div>
|
||||
<p>
|
||||
R0 的 16 sources × 4 conditions 在全新进程重跑。比较不只看 headline,
|
||||
而是逐格检查 seed、prompt、完整 token IDs、文本、停止状态与 RNG pre-state。
|
||||
</p>
|
||||
</div>
|
||||
|
||||
<div class="repro-equation">
|
||||
<article><span>DIFFERENT SEEDS</span><b>63 / 64 分叉</b><p>R0 与 R1 的同 source-condition 轨迹</p></article>
|
||||
<i>+</i>
|
||||
<article><span>SAME FULL CONTRACT</span><b>64 / 64 exact</b><p>R0 与全新进程 R0′</p></article>
|
||||
<i>=</i>
|
||||
<article class="result"><span>AUDITABLE SAMPLING</span><b>随机 ≠ 漂移</b><p>分叉与复现同时过闸</p></article>
|
||||
</div>
|
||||
|
||||
<div class="repro-fields">
|
||||
{Object.entries(lab.reproduction.by_field).map(([field, count]) => (
|
||||
<article>
|
||||
<span>{String(field).replaceAll("_", " ").toUpperCase()}</span>
|
||||
<b>{String(count)} / 64</b><i>EXACT</i>
|
||||
</article>
|
||||
))}
|
||||
</div>
|
||||
|
||||
<div class="hash-chain">
|
||||
{Object.entries(lab.artifacts).map(([name, artifact]: [string, any], index) => (
|
||||
<>
|
||||
<article>
|
||||
<span>{String(index + 1).padStart(2, "0")} / {name.toUpperCase()}</span>
|
||||
<b>{artifact.sha256.slice(0, 12)}…{artifact.sha256.slice(-8)}</b>
|
||||
<p>{artifact.bytes.toLocaleString()} bytes</p>
|
||||
</article>
|
||||
{index < Object.keys(lab.artifacts).length - 1 && <i>→</i>}
|
||||
</>
|
||||
))}
|
||||
</div>
|
||||
|
||||
<div class="execution-contract">
|
||||
<article><span>CHECKPOINT</span><b>85864749…f64c7</b><p>官方 DeepSeek-V2-Lite-Chat revision</p></article>
|
||||
<article><span>SOFTWARE</span><b>Torch 2.11 · TF 4.41.2</b><p>BF16 · eager attention</p></article>
|
||||
<article><span>PLACEMENT</span><b>CUDA L0–23 · CPU L24–26</b><p>norm 与 LM head 也在 CPU</p></article>
|
||||
<article><span>CODE SANDBOX</span><b>network none · read-only</b><p>47 个唯一候选;17 次 cache hit</p></article>
|
||||
</div>
|
||||
|
||||
<aside class="cs-note">
|
||||
<b>exact replay 只属于这份执行合同</b>
|
||||
<p>
|
||||
更换软件、kernel、硬件、batch 行数或顺序都可能改变采样轨迹。这里证明的是固定合同内
|
||||
可复现,不承诺跨版本 bit-exact。
|
||||
</p>
|
||||
</aside>
|
||||
</section>
|
||||
|
||||
<figcaption>
|
||||
<b>证据边界</b>
|
||||
<span>
|
||||
16 sources 仍不是 benchmark;四题方向不推断总体。seed 是题内重复;句点不是官方聊天
|
||||
边界;source-blocked contrast 是固定网格描述量,没有 p-value 与总体置信区间。
|
||||
</span>
|
||||
<code>FORMAL f013132…d7f7c · EVAL e88b274…b8975 · REPLAY 143dc9d…02a6</code>
|
||||
</figcaption>
|
||||
|
||||
<script is:inline type="application/json" data-cs-data set:html={json}></script>
|
||||
</figure>
|
||||
|
||||
<script>
|
||||
document.querySelectorAll<HTMLElement>("[data-cross-source-lab]").forEach((root) => {
|
||||
const payload = root.querySelector<HTMLScriptElement>("[data-cs-data]");
|
||||
if (!payload) return;
|
||||
const data = JSON.parse(payload.textContent ?? "{}");
|
||||
const one = <T extends Element>(selector: string) => root.querySelector<T>(selector);
|
||||
const set = (selector: string, value: string) => {
|
||||
const node = one<HTMLElement>(selector);
|
||||
if (node) node.textContent = value;
|
||||
};
|
||||
const signed = (value: number, digits = 1) => (
|
||||
`${value > 0 ? "+" : ""}${value.toFixed(digits)}`
|
||||
);
|
||||
|
||||
const tabs = [...root.querySelectorAll<HTMLButtonElement>("[data-cs-tab]")];
|
||||
const panels = [...root.querySelectorAll<HTMLElement>("[data-cs-panel]")];
|
||||
tabs.forEach((button, index) => {
|
||||
button.addEventListener("click", () => {
|
||||
const target = button.dataset.csTab;
|
||||
tabs.forEach((candidate) => {
|
||||
const active = candidate === button;
|
||||
candidate.setAttribute("aria-selected", String(active));
|
||||
candidate.tabIndex = active ? 0 : -1;
|
||||
});
|
||||
panels.forEach((panel) => {
|
||||
panel.hidden = panel.dataset.csPanel !== target;
|
||||
});
|
||||
});
|
||||
button.addEventListener("keydown", (event) => {
|
||||
if (!["ArrowLeft", "ArrowRight"].includes(event.key)) return;
|
||||
event.preventDefault();
|
||||
const delta = event.key === "ArrowRight" ? 1 : -1;
|
||||
const target = tabs[(index + delta + tabs.length) % tabs.length];
|
||||
target.click();
|
||||
target.focus();
|
||||
});
|
||||
});
|
||||
|
||||
const domainSelect = one<HTMLSelectElement>("[data-cs-domain]");
|
||||
const conditionSelect = one<HTMLSelectElement>("[data-cs-condition]");
|
||||
const sourceCards = one<HTMLElement>("[data-cs-source-cards]");
|
||||
const renderHierarchy = () => {
|
||||
const domain = domainSelect?.value ?? "english";
|
||||
const condition = conditionSelect?.value ?? "s0_eos";
|
||||
const summary = data.domainConditions[domain][condition];
|
||||
set("[data-cs-output-count]", `${summary.outputs} / 16`);
|
||||
set("[data-cs-natural-eos]", `${summary.micro.natural_eos} / 16`);
|
||||
set(
|
||||
"[data-cs-between-range]",
|
||||
`${summary.variability.between_source_mean_length_range.toFixed(1)} tokens`,
|
||||
);
|
||||
set(
|
||||
"[data-cs-within-range]",
|
||||
`${summary.variability.within_source_seed_length_range.mean.toFixed(1)} tokens`,
|
||||
);
|
||||
if (!sourceCards) return;
|
||||
sourceCards.replaceChildren();
|
||||
data.sourcesByDomain[domain].forEach((sourceId: string) => {
|
||||
const cell = data.sourceCells[sourceId][condition];
|
||||
const samples = data.evalRows.filter((row: any) => (
|
||||
row.sourceId === sourceId && row.condition === condition
|
||||
));
|
||||
const article = document.createElement("article");
|
||||
article.dataset.sourceId = sourceId;
|
||||
const header = document.createElement("header");
|
||||
const heading = document.createElement("div");
|
||||
const index = document.createElement("span");
|
||||
index.textContent = `SOURCE ${String(data.sourceMeta[sourceId].within_domain_index + 1).padStart(2, "0")}`;
|
||||
const title = document.createElement("b");
|
||||
title.textContent = data.sourceMeta[sourceId].label;
|
||||
heading.append(index, title);
|
||||
const eos = document.createElement("em");
|
||||
eos.textContent = `${cell.natural_eos} / 4 EOS`;
|
||||
header.append(heading, eos);
|
||||
const dots = document.createElement("div");
|
||||
dots.className = "seed-dots";
|
||||
samples.forEach((sample: any) => {
|
||||
const dot = document.createElement("i");
|
||||
dot.className = sample.hitEos ? "eos" : "truncated";
|
||||
dot.style.setProperty("--seed-height", `${Math.max(12, sample.generatedTokens / 512 * 100)}%`);
|
||||
dot.title = `${sample.replicate}: ${sample.generatedTokens} tokens`;
|
||||
const label = document.createElement("span");
|
||||
label.textContent = sample.replicate;
|
||||
const value = document.createElement("b");
|
||||
value.textContent = String(sample.generatedTokens);
|
||||
dot.append(label, value);
|
||||
dots.append(dot);
|
||||
});
|
||||
const footer = document.createElement("p");
|
||||
const task = cell.strict_complete_success === null
|
||||
? `${cell.unique_trajectory_hashes} / 4 unique token trajectories`
|
||||
: `${cell.strict_complete_success} / 4 strict task pass`;
|
||||
footer.textContent = `mean ${cell.generated_tokens.mean.toFixed(1)} · range ${cell.generated_tokens.min}–${cell.generated_tokens.max} · ${task}`;
|
||||
article.append(header, dots, footer);
|
||||
sourceCards.append(article);
|
||||
});
|
||||
};
|
||||
domainSelect?.addEventListener("change", renderHierarchy);
|
||||
conditionSelect?.addEventListener("change", renderHierarchy);
|
||||
renderHierarchy();
|
||||
|
||||
const taskButtons = [
|
||||
...root.querySelectorAll<HTMLButtonElement>("[data-cs-task-domain]"),
|
||||
];
|
||||
const taskMatrix = one<HTMLElement>("[data-cs-task-matrix]");
|
||||
const renderTasks = (domain: "math" | "code") => {
|
||||
taskButtons.forEach((button) => button.setAttribute(
|
||||
"aria-pressed",
|
||||
String(button.dataset.csTaskDomain === domain),
|
||||
));
|
||||
const tasks = data.taskTotals[domain];
|
||||
const total = tasks.reduce((sum: number, task: any) => sum + task.totalPass, 0);
|
||||
set("[data-cs-task-total-label]", domain === "math" ? "MATH · STRICT" : "CODE · OFFICIAL TESTS");
|
||||
set("[data-cs-task-total]", `${total} / 64`);
|
||||
if (taskMatrix) {
|
||||
taskMatrix.replaceChildren();
|
||||
const header = document.createElement("header");
|
||||
const blank = document.createElement("b");
|
||||
blank.textContent = "SOURCE / TASK";
|
||||
header.append(blank);
|
||||
data.contract.conditions.forEach((condition: string) => {
|
||||
const label = document.createElement("b");
|
||||
label.textContent = data.contract.conditionLabels[condition];
|
||||
header.append(label);
|
||||
});
|
||||
const sum = document.createElement("b");
|
||||
sum.textContent = "TOTAL";
|
||||
header.append(sum);
|
||||
taskMatrix.append(header);
|
||||
tasks.forEach((task: any) => {
|
||||
const row = document.createElement("div");
|
||||
row.dataset.taskId = task.sourceId;
|
||||
const label = document.createElement("span");
|
||||
label.textContent = task.label;
|
||||
row.append(label);
|
||||
data.contract.conditions.forEach((condition: string) => {
|
||||
const cell = task.conditions[condition];
|
||||
const item = document.createElement("b");
|
||||
item.textContent = `${cell.pass} / 4`;
|
||||
item.dataset.pass = String(cell.pass);
|
||||
item.style.setProperty("--pass", String(cell.pass));
|
||||
item.title = `${cell.naturalEos}/4 EOS · mean ${cell.meanGeneratedTokens.toFixed(1)} tokens`;
|
||||
row.append(item);
|
||||
});
|
||||
const totalCell = document.createElement("strong");
|
||||
totalCell.textContent = `${task.totalPass} / 16`;
|
||||
row.append(totalCell);
|
||||
taskMatrix.append(row);
|
||||
});
|
||||
}
|
||||
const totals = tasks.map((task: any) => task.totalPass);
|
||||
set("[data-cs-task-range]", `${Math.min(...totals)} → ${Math.max(...totals)} / 16`);
|
||||
set(
|
||||
"[data-cs-task-range-copy]",
|
||||
domain === "math"
|
||||
? "GSM8K/0144 只有 8/16;GSM8K/1228 是 16/16。"
|
||||
: "HumanEval/44 只有 8/16;HumanEval/31 与 /23 都是 16/16。",
|
||||
);
|
||||
const interaction = data.domainContrasts[domain]
|
||||
.strict_complete_success_rate.interaction;
|
||||
set(
|
||||
"[data-cs-task-interaction]",
|
||||
`${interaction.directions.positive}↑ · ${interaction.directions.zero}= · ${interaction.directions.negative}↓`,
|
||||
);
|
||||
set(
|
||||
"[data-cs-task-interaction-copy]",
|
||||
domain === "math"
|
||||
? "两题为负、两题为零;domain mean −0.188。"
|
||||
: "一题正、一题负、两题为零;domain mean 恰好 0。",
|
||||
);
|
||||
const failures = data.taskFailures.filter((row: any) => row.domain === domain);
|
||||
set("[data-cs-task-failure]", `${failures.length} / 64`);
|
||||
set(
|
||||
"[data-cs-task-failure-copy]",
|
||||
domain === "math"
|
||||
? "17 条都是数值答案错误;它们仍全部自然 EOS。"
|
||||
: "5 次 runtime error、7 次 assertion failed;其中 1 条输出截断。",
|
||||
);
|
||||
};
|
||||
taskButtons.forEach((button) => button.addEventListener("click", () => (
|
||||
renderTasks((button.dataset.csTaskDomain ?? "math") as "math" | "code")
|
||||
)));
|
||||
renderTasks("math");
|
||||
|
||||
const directionSelect = one<HTMLSelectElement>("[data-cs-direction-domain]");
|
||||
const directionRows = one<HTMLElement>("[data-cs-direction-rows]");
|
||||
const renderDirections = () => {
|
||||
const domain = directionSelect?.value ?? "english";
|
||||
const contrast = data.domainContrasts[domain].mean_generated_tokens
|
||||
.boundary_main_period_minus_eos;
|
||||
set(
|
||||
"[data-cs-shorter]",
|
||||
`${data.periodDirection[domain].shorter} / ${data.periodDirection[domain].sources}`,
|
||||
);
|
||||
set("[data-cs-domain-mean]", `${signed(contrast.mean)} tokens`);
|
||||
set("[data-cs-domain-median]", `${signed(contrast.median)} tokens`);
|
||||
if (!directionRows) return;
|
||||
directionRows.replaceChildren();
|
||||
data.sourcesByDomain[domain].forEach((sourceId: string) => {
|
||||
const value = contrast.by_source[sourceId];
|
||||
const row = document.createElement("article");
|
||||
row.dataset.directionSource = sourceId;
|
||||
const label = document.createElement("span");
|
||||
label.textContent = data.sourceMeta[sourceId].label;
|
||||
const track = document.createElement("i");
|
||||
const dot = document.createElement("u");
|
||||
const position = Math.max(2, Math.min(98, 50 + value / 5));
|
||||
dot.style.setProperty("--dot-position", `${position}%`);
|
||||
dot.className = value < 0 ? "negative" : "positive";
|
||||
track.append(dot);
|
||||
const output = document.createElement("b");
|
||||
output.textContent = `${signed(value)} tokens`;
|
||||
row.append(label, track, output);
|
||||
directionRows.append(row);
|
||||
});
|
||||
};
|
||||
directionSelect?.addEventListener("change", renderDirections);
|
||||
renderDirections();
|
||||
});
|
||||
</script>
|
||||
|
||||
<style>
|
||||
.cross-source-lab { margin: 1.6rem 0 0; overflow: hidden; border: 1px solid var(--line); background: #f7f3ea; }
|
||||
.cs-head { display: grid; grid-template-columns: 1.08fr .92fr; gap: 1.5rem; padding: 1.55rem; color: #edf3f1; background: #263f47; }
|
||||
.cs-head p { margin: 0; font-size: .62rem; line-height: 1.65; }
|
||||
.cs-head > div > p { color: #8fcbbf; font: 690 .48rem/1.2 var(--font-mono); letter-spacing: .08em; }
|
||||
.cs-head h3 { margin: .6rem 0 0; max-width: 25ch; color: #edf3f1; font: 760 1.22rem/1.14 var(--font-display); }
|
||||
.cs-head code { color: #f0c7aa; font-size: .53rem; }
|
||||
.cs-ledger { display: grid; grid-template-columns: repeat(6,1fr); border-bottom: 1px solid var(--line); background: var(--line); gap: 1px; }
|
||||
.cs-ledger article { min-width: 0; padding: .82rem; background: #eee9df; }
|
||||
.cs-ledger article.pass { background: #dcebe5; }
|
||||
.cs-ledger span, .source-summary span, .direction-controls span { display: block; color: #687772; font: .44rem/1.2 var(--font-mono); }
|
||||
.cs-ledger b { display: block; margin-top: .38rem; font: 760 .72rem/1.1 var(--font-mono); }
|
||||
.cs-ledger p { margin: .32rem 0 0; color: #717975; font-size: .46rem; line-height: 1.4; }
|
||||
.cs-tabs { display: grid; grid-template-columns: repeat(4,1fr); gap: 1px; background: var(--line); border-bottom: 1px solid var(--line); }
|
||||
.cs-tabs button { position: relative; min-width: 0; padding: .88rem; text-align: left; color: #52615e; border: 0; background: #e5e0d5; cursor: pointer; }
|
||||
.cs-tabs button[aria-selected="true"] { color: #f0f5f3; background: #2e776c; }
|
||||
.cs-tabs span, .cs-tabs b, .cs-tabs small { display: block; }
|
||||
.cs-tabs span { color: var(--orange); font: 730 .44rem/1 var(--font-mono); }
|
||||
.cs-tabs button[aria-selected="true"] span { color: #f2c6a8; }
|
||||
.cs-tabs b { margin-top: .38rem; font-size: .62rem; }
|
||||
.cs-tabs small { margin-top: .25rem; opacity: .65; font: .42rem/1 var(--font-mono); }
|
||||
.cs-panel { padding: 1.25rem; }
|
||||
.cs-panel-lead { display: grid; grid-template-columns: 1.1fr .9fr; gap: 1.4rem; margin-bottom: 1rem; }
|
||||
.cs-panel-lead span { color: var(--orange); font: 700 .46rem/1.2 var(--font-mono); }
|
||||
.cs-panel-lead h4 { margin: .42rem 0 0; font: 750 .96rem/1.15 var(--font-display); }
|
||||
.cs-panel-lead p { margin: 0; color: #63706c; font-size: .57rem; line-height: 1.65; }
|
||||
.unit-equation { display: grid; grid-template-columns: 1fr auto 1.2fr auto 1fr; gap: .6rem; align-items: center; }
|
||||
.unit-equation article { padding: .82rem; border: 1px solid var(--line); background: #eee9df; }
|
||||
.unit-equation article.primary { color: #eef4f2; border: 0; background: #2f776c; }
|
||||
.unit-equation span { font: .43rem/1.2 var(--font-mono); opacity: .72; }
|
||||
.unit-equation b { display: block; margin-top: .4rem; font: 720 .67rem/1.15 var(--font-mono); }
|
||||
.unit-equation p { margin: .35rem 0 0; opacity: .72; font-size: .48rem; }
|
||||
.unit-equation > i { color: var(--orange); font: 740 .72rem/1 var(--font-mono); }
|
||||
.cs-controls { display: grid; grid-template-columns: 1fr 1fr 1.15fr; gap: 1px; margin-top: 1rem; border: 1px solid var(--line); background: var(--line); }
|
||||
.cs-controls label, .cs-controls > div { padding: .72rem; background: #eee9df; }
|
||||
.cs-controls span { display: block; color: #6e7975; font: .43rem/1 var(--font-mono); }
|
||||
.cs-controls select { width: 100%; margin-top: .38rem; padding: .42rem; font: 650 .54rem/1.2 var(--font-mono); border: 1px solid var(--line); background: #fffaf2; }
|
||||
.cs-controls b { display: block; margin-top: .48rem; font: 680 .54rem/1.2 var(--font-mono); }
|
||||
.source-summary { display: grid; grid-template-columns: repeat(4,1fr); gap: 1px; margin-top: 1rem; border: 1px solid var(--line); background: var(--line); }
|
||||
.source-summary article { padding: .72rem; background: #f9f5ec; }
|
||||
.source-summary b { display: block; margin-top: .38rem; color: #2f6f65; font: 740 .65rem/1.1 var(--font-mono); }
|
||||
.source-summary p { margin: .32rem 0 0; color: #737d79; font-size: .46rem; }
|
||||
.source-card-grid { display: grid; grid-template-columns: repeat(4,1fr); gap: .7rem; margin-top: 1rem; }
|
||||
.cs-note { margin-top: 1rem; padding: .85rem 1rem; border-left: .24rem solid var(--orange); background: #eee9df; }
|
||||
.cs-note.dark { color: #e8efed; border-left-color: #e0a17c; background: #29434a; }
|
||||
.cs-note b { font: 710 .6rem/1.3 var(--font-mono); }
|
||||
.cs-note p { margin: .42rem 0 0; opacity: .76; font-size: .53rem; line-height: 1.6; }
|
||||
.task-switch { display: flex; justify-content: space-between; gap: 1rem; align-items: stretch; padding: .65rem; border: 1px solid var(--line); background: #eee9df; }
|
||||
.task-switch > div { display: flex; gap: .35rem; }
|
||||
.task-switch button { padding: .55rem .75rem; color: #53625f; border: 1px solid var(--line); background: #fffaf2; font: 690 .48rem/1 var(--font-mono); cursor: pointer; }
|
||||
.task-switch button[aria-pressed="true"] { color: #edf3f1; border-color: #2f776c; background: #2f776c; }
|
||||
.task-switch p { margin: 0; padding: .35rem .5rem; text-align: right; }
|
||||
.task-switch p span { display: block; color: #6a7773; font: .42rem/1 var(--font-mono); }
|
||||
.task-switch p b { display: block; margin-top: .32rem; font: 750 .66rem/1 var(--font-mono); }
|
||||
.task-matrix { margin-top: 1rem; overflow: hidden; border: 1px solid var(--line); }
|
||||
.matrix-reading { display: grid; grid-template-columns: repeat(3,1fr); gap: 1px; margin-top: 1rem; border: 1px solid var(--line); background: var(--line); }
|
||||
.matrix-reading article { padding: .82rem; background: #eee9df; }
|
||||
.matrix-reading span { color: var(--orange); font: .45rem/1.2 var(--font-mono); }
|
||||
.matrix-reading b { display: block; margin-top: .4rem; font: 740 .68rem/1.15 var(--font-mono); }
|
||||
.matrix-reading p { margin: .4rem 0 0; color: #69736f; font-size: .51rem; line-height: 1.5; }
|
||||
.direction-controls { display: grid; grid-template-columns: 1.25fr repeat(3,1fr); gap: 1px; border: 1px solid var(--line); background: var(--line); }
|
||||
.direction-controls label, .direction-controls article { padding: .75rem; background: #eee9df; }
|
||||
.direction-controls select { width: 100%; margin-top: .4rem; padding: .42rem; border: 1px solid var(--line); background: #fffaf2; font: 650 .53rem/1.2 var(--font-mono); }
|
||||
.direction-controls b { display: block; margin-top: .4rem; font: 740 .65rem/1.1 var(--font-mono); }
|
||||
.direction-controls p { margin: .28rem 0 0; color: #6c7672; font-size: .44rem; }
|
||||
.zero-axis { margin-top: 1rem; overflow: hidden; border: 1px solid var(--line); }
|
||||
.zero-axis > header { display: grid; grid-template-columns: 1fr auto 1fr; padding: .6rem .8rem; color: #e9f0ee; background: #29434a; font: .43rem/1 var(--font-mono); }
|
||||
.zero-axis > header span:last-child { text-align: right; }
|
||||
.zero-axis > header b { color: #a8d2c9; }
|
||||
.round-correction { display: grid; grid-template-columns: 1fr auto 1fr; gap: .7rem; align-items: center; margin-top: 1rem; }
|
||||
.round-correction article { padding: .9rem; border: 1px solid var(--line); background: #eee9df; }
|
||||
.round-correction article.result { color: #eaf1ef; border: 0; background: #2e776c; }
|
||||
.round-correction span { font: .43rem/1.2 var(--font-mono); opacity: .72; }
|
||||
.round-correction b { display: block; margin-top: .42rem; font: 720 .65rem/1.2 var(--font-mono); }
|
||||
.round-correction p { margin: .38rem 0 0; opacity: .72; font-size: .51rem; line-height: 1.5; }
|
||||
.round-correction > i { color: var(--orange); font: 760 .72rem/1 var(--font-mono); }
|
||||
.direction-all { display: grid; grid-template-columns: repeat(4,1fr); gap: 1px; margin-top: 1rem; border: 1px solid var(--line); background: var(--line); }
|
||||
.direction-all article { padding: .75rem; background: #eee9df; }
|
||||
.direction-all span { color: #66736f; font: .43rem/1.2 var(--font-mono); }
|
||||
.direction-all b { display: block; margin-top: .38rem; font: 720 .58rem/1.2 var(--font-mono); }
|
||||
.direction-all p { margin: .32rem 0 0; color: #717b77; font: .44rem/1.4 var(--font-mono); }
|
||||
.repro-equation { display: grid; grid-template-columns: 1fr auto 1fr auto 1fr; gap: .65rem; align-items: center; }
|
||||
.repro-equation article { padding: .9rem; border: 1px solid var(--line); background: #eee9df; }
|
||||
.repro-equation article.result { color: #eaf2ef; border: 0; background: #2e776c; }
|
||||
.repro-equation span { font: .43rem/1.2 var(--font-mono); opacity: .7; }
|
||||
.repro-equation b { display: block; margin-top: .4rem; font: 730 .67rem/1.1 var(--font-mono); }
|
||||
.repro-equation p { margin: .35rem 0 0; opacity: .72; font-size: .49rem; }
|
||||
.repro-equation > i { color: var(--orange); font: 760 .78rem/1 var(--font-mono); }
|
||||
.repro-fields { display: grid; grid-template-columns: repeat(4,1fr); gap: 1px; margin-top: 1rem; border: 1px solid var(--line); background: var(--line); }
|
||||
.repro-fields article { padding: .72rem; background: #dfece7; }
|
||||
.repro-fields span { display: block; color: #58716b; font: .4rem/1.2 var(--font-mono); }
|
||||
.repro-fields b { display: block; margin-top: .34rem; color: #28675c; font: 740 .59rem/1.1 var(--font-mono); }
|
||||
.repro-fields i { display: block; margin-top: .25rem; color: #568178; font: .38rem/1 var(--font-mono); }
|
||||
.hash-chain { display: grid; grid-template-columns: 1fr auto 1fr auto 1fr auto 1fr auto 1fr; gap: .35rem; align-items: center; margin-top: 1rem; padding: .75rem; border: 1px solid var(--line); background: #eee9df; }
|
||||
.hash-chain article { min-width: 0; padding: .55rem; background: #fffaf2; }
|
||||
.hash-chain span { display: block; color: var(--orange); font: .39rem/1.2 var(--font-mono); }
|
||||
.hash-chain b { display: block; margin-top: .34rem; overflow: hidden; font: 650 .43rem/1.25 var(--font-mono); text-overflow: ellipsis; }
|
||||
.hash-chain p { margin: .28rem 0 0; color: #6f7975; font: .38rem/1.2 var(--font-mono); }
|
||||
.hash-chain > i { color: var(--orange); font: 700 .55rem/1 var(--font-mono); }
|
||||
.execution-contract { display: grid; grid-template-columns: repeat(4,1fr); gap: 1px; margin-top: 1rem; border: 1px solid var(--line); background: var(--line); }
|
||||
.execution-contract article { min-width: 0; padding: .76rem; background: #eee9df; }
|
||||
.execution-contract span { color: #64716d; font: .42rem/1.2 var(--font-mono); }
|
||||
.execution-contract b { display: block; margin-top: .38rem; overflow-wrap: anywhere; font: 680 .54rem/1.3 var(--font-mono); }
|
||||
.execution-contract p { margin: .34rem 0 0; color: #707975; font-size: .46rem; }
|
||||
.cross-source-lab > figcaption { display: grid; grid-template-columns: auto 1fr auto; gap: 1rem; align-items: center; padding: .9rem 1.1rem; color: #e2eae8; background: #203a42; }
|
||||
.cross-source-lab > figcaption b { color: #8dc8bd; font: 720 .5rem/1.2 var(--font-mono); }
|
||||
.cross-source-lab > figcaption span { font-size: .52rem; line-height: 1.5; }
|
||||
.cross-source-lab > figcaption code { color: #dda988; font: .4rem/1.3 var(--font-mono); }
|
||||
@media (max-width: 900px) {
|
||||
.cs-head, .cs-panel-lead { grid-template-columns: 1fr; }
|
||||
.cs-ledger { grid-template-columns: repeat(3,1fr); }
|
||||
.cs-tabs { grid-template-columns: 1fr 1fr; }
|
||||
.source-card-grid, .source-summary, .direction-all, .execution-contract { grid-template-columns: 1fr 1fr; }
|
||||
.hash-chain { grid-template-columns: 1fr; }
|
||||
.hash-chain > i { transform: rotate(90deg); text-align: center; }
|
||||
.cross-source-lab > figcaption { grid-template-columns: 1fr; }
|
||||
}
|
||||
@media (max-width: 640px) {
|
||||
.cs-panel { padding: .9rem; }
|
||||
.cs-head { padding: 1.2rem; }
|
||||
.cs-tabs, .cs-controls, .unit-equation, .task-switch, .matrix-reading,
|
||||
.direction-controls, .round-correction, .repro-equation { grid-template-columns: 1fr; }
|
||||
.task-switch { display: grid; }
|
||||
.task-switch > div { display: grid; grid-template-columns: 1fr 1fr; }
|
||||
.task-switch p { text-align: left; }
|
||||
.unit-equation > i, .round-correction > i, .repro-equation > i { transform: rotate(90deg); text-align: center; }
|
||||
.source-card-grid, .source-summary, .direction-all, .execution-contract, .repro-fields { grid-template-columns: 1fr 1fr; }
|
||||
}
|
||||
</style>
|
||||
|
||||
<style is:global>
|
||||
[data-cross-source-lab] [data-cs-source-cards] > article { min-width: 0; overflow: hidden; border: 1px solid var(--line); background: #eee9df; }
|
||||
[data-cross-source-lab] [data-cs-source-cards] header { display: flex; justify-content: space-between; gap: .5rem; padding: .7rem; border-bottom: 1px solid var(--line); }
|
||||
[data-cross-source-lab] [data-cs-source-cards] header span,
|
||||
[data-cross-source-lab] [data-cs-source-cards] header b { display: block; }
|
||||
[data-cross-source-lab] [data-cs-source-cards] header span { color: var(--orange); font: .4rem/1 var(--font-mono); }
|
||||
[data-cross-source-lab] [data-cs-source-cards] header b { margin-top: .3rem; font: 700 .54rem/1.15 var(--font-mono); }
|
||||
[data-cross-source-lab] [data-cs-source-cards] header em { color: #57736c; font: 650 .42rem/1.2 var(--font-mono); }
|
||||
[data-cross-source-lab] .seed-dots { display: grid; grid-template-columns: repeat(4,1fr); gap: 1px; height: 8.2rem; padding: .6rem; background: rgba(37,61,67,.12); }
|
||||
[data-cross-source-lab] .seed-dots > i { position: relative; display: flex; flex-direction: column; justify-content: space-between; align-self: end; height: var(--seed-height); min-height: 1.3rem; padding: .28rem; color: #f2f6f4; background: linear-gradient(to top,#2f776c,#72aea3); font-style: normal; }
|
||||
[data-cross-source-lab] .seed-dots > i.truncated { background: linear-gradient(to top,#a94f31,#d4936c); }
|
||||
[data-cross-source-lab] .seed-dots span { font: .37rem/1 var(--font-mono); }
|
||||
[data-cross-source-lab] .seed-dots b { font: 690 .43rem/1 var(--font-mono); }
|
||||
[data-cross-source-lab] [data-cs-source-cards] > article > p { margin: 0; padding: .6rem .7rem; color: #65716d; font: .4rem/1.45 var(--font-mono); }
|
||||
[data-cross-source-lab] [data-cs-task-matrix] > header,
|
||||
[data-cross-source-lab] [data-cs-task-matrix] > div { display: grid; grid-template-columns: 1.35fr repeat(4,1fr) .85fr; gap: 1px; align-items: stretch; background: var(--line); }
|
||||
[data-cross-source-lab] [data-cs-task-matrix] > header > *,
|
||||
[data-cross-source-lab] [data-cs-task-matrix] > div > * { min-width: 0; padding: .65rem .45rem; }
|
||||
[data-cross-source-lab] [data-cs-task-matrix] > header > * { color: #e7efed; background: #29434a; font: 650 .4rem/1.2 var(--font-mono); text-align: center; }
|
||||
[data-cross-source-lab] [data-cs-task-matrix] > header > b:first-child { text-align: left; }
|
||||
[data-cross-source-lab] [data-cs-task-matrix] > div > span { background: #eee9df; font: 670 .48rem/1.2 var(--font-mono); }
|
||||
[data-cross-source-lab] [data-cs-task-matrix] > div > b { color: #244a43; background: color-mix(in srgb,#338274 calc(var(--pass) * 20%),#f7f3ea); font: 740 .54rem/1 var(--font-mono); text-align: center; }
|
||||
[data-cross-source-lab] [data-cs-task-matrix] > div > strong { color: #eaf1ef; background: #2f776c; font: 740 .54rem/1 var(--font-mono); text-align: center; }
|
||||
[data-cross-source-lab] [data-cs-direction-rows] > article { display: grid; grid-template-columns: 1fr 2.8fr .8fr; gap: .7rem; align-items: center; padding: .7rem .8rem; border-bottom: 1px solid var(--line); background: #eee9df; }
|
||||
[data-cross-source-lab] [data-cs-direction-rows] > article:last-child { border-bottom: 0; }
|
||||
[data-cross-source-lab] [data-cs-direction-rows] span { font: 650 .5rem/1.2 var(--font-mono); }
|
||||
[data-cross-source-lab] [data-cs-direction-rows] > article > i { position: relative; height: .44rem; background: linear-gradient(to right,#d99a75 0 49.7%,#345f59 49.7% 50.3%,#8fc8bd 50.3% 100%); }
|
||||
[data-cross-source-lab] [data-cs-direction-rows] u { position: absolute; top: 50%; left: var(--dot-position); width: .8rem; height: .8rem; border: .14rem solid #f7f3ea; border-radius: 50%; background: #2f776c; box-shadow: 0 0 0 1px #29434a; transform: translate(-50%,-50%); }
|
||||
[data-cross-source-lab] [data-cs-direction-rows] u.negative { background: #b65c38; }
|
||||
[data-cross-source-lab] [data-cs-direction-rows] b { font: 700 .49rem/1.2 var(--font-mono); text-align: right; }
|
||||
@media (max-width: 640px) {
|
||||
[data-cross-source-lab] [data-cs-task-matrix] > header,
|
||||
[data-cross-source-lab] [data-cs-task-matrix] > div { grid-template-columns: 1.2fr repeat(4,.7fr) .72fr; }
|
||||
[data-cross-source-lab] [data-cs-task-matrix] > header > *,
|
||||
[data-cross-source-lab] [data-cs-task-matrix] > div > * { padding: .5rem .2rem; font-size: .36rem; }
|
||||
[data-cross-source-lab] [data-cs-direction-rows] > article { grid-template-columns: 1fr; }
|
||||
[data-cross-source-lab] [data-cs-direction-rows] b { text-align: left; }
|
||||
}
|
||||
</style>
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,941 @@
|
||||
{
|
||||
"schema_version": 1,
|
||||
"protocol_id": "llm-atlas-deepseek-chat-cross-source-sampling-v1",
|
||||
"formal": {
|
||||
"path": "/tmp/deepseek-v2-lite-chat-cross-source-sampling-formal.json",
|
||||
"sha256": "f013132485f27adce008f03f781bed9982efc0d7939f13faede01c9f6f3d7f7c",
|
||||
"content_hash": "8fa14db0e9e6f1fd649a500796953ca948606a2c525d9c92c4b5f848beeaa3a7",
|
||||
"base_seeds": [
|
||||
2101325316,
|
||||
2511573438,
|
||||
1677220094,
|
||||
2412346607
|
||||
]
|
||||
},
|
||||
"rerun": {
|
||||
"path": "/tmp/deepseek-v2-lite-chat-cross-source-sampling-repro-r0.json",
|
||||
"sha256": "143dc9d0f7c914db4781e36b1401cdc9fbc2971a0dca71188bab0f8cedb002a6",
|
||||
"content_hash": "dc2706ed500e2ec3f7922b4aecbfd19247353d3b973162ea4bff44188b93e7f9",
|
||||
"base_seeds": [
|
||||
2101325316
|
||||
]
|
||||
},
|
||||
"rows": [
|
||||
{
|
||||
"source_id": "HumanEval/133",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "HumanEval/133",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "HumanEval/133",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "HumanEval/133",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "HumanEval/23",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "HumanEval/23",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "HumanEval/23",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "HumanEval/23",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "HumanEval/31",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "HumanEval/31",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "HumanEval/31",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "HumanEval/31",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "HumanEval/44",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "HumanEval/44",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "HumanEval/44",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "HumanEval/44",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/0144",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/0144",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/0144",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/0144",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/1069",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/1069",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/1069",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/1069",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/1228",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/1228",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/1228",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/1228",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/1251",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/1251",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/1251",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "gsm8k/test/1251",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/1059",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/1059",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/1059",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/1059",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/3448",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/3448",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/3448",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/3448",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/4855",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/4855",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/4855",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/4855",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/8935",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/8935",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/8935",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "tnews/test/8935",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/0030",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/0030",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/0030",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/0030",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/0443",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/0443",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/0443",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/0443",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/2746",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/2746",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/2746",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/2746",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/2909",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/2909",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s0_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/2909",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_eos",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
},
|
||||
{
|
||||
"source_id": "wikitext2/raw-validation/2909",
|
||||
"base_seed": 2101325316,
|
||||
"condition": "s1_period",
|
||||
"run_seed_exact": true,
|
||||
"prompt_hash_exact": true,
|
||||
"generated_token_ids_exact": true,
|
||||
"decoded_text_exact": true,
|
||||
"eos_state_exact": true,
|
||||
"truncation_state_exact": true,
|
||||
"cpu_rng_pre_state_exact": true,
|
||||
"cuda_rng_pre_state_exact": true,
|
||||
"all_preregistered_fields_exact": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"cells": 64,
|
||||
"all_preregistered_fields_exact": 64,
|
||||
"by_field": {
|
||||
"run_seed_exact": 64,
|
||||
"prompt_hash_exact": 64,
|
||||
"generated_token_ids_exact": 64,
|
||||
"decoded_text_exact": 64,
|
||||
"eos_state_exact": 64,
|
||||
"truncation_state_exact": 64,
|
||||
"cpu_rng_pre_state_exact": 64,
|
||||
"cuda_rng_pre_state_exact": 64
|
||||
}
|
||||
},
|
||||
"claim_boundary": [
|
||||
"Only the rerun seed subset is independently reproduced.",
|
||||
"Exact replay is scoped to the pinned software and hardware contract.",
|
||||
"Reproduction does not imply trajectories are seed-invariant."
|
||||
],
|
||||
"content_hash": "35562a6c5b5bdafd9472d8d638975b84e5ecf06d7d2061b64c341ddf28790da1"
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -6,6 +6,7 @@ import DeepSeekArtifactLab from "@/components/DeepSeekArtifactLab.astro";
|
||||
import DeepSeekBehaviorLab from "@/components/DeepSeekBehaviorLab.astro";
|
||||
import DeepSeekCompletionDepthLab from "@/components/DeepSeekCompletionDepthLab.astro";
|
||||
import DeepSeekSamplingLab from "@/components/DeepSeekSamplingLab.astro";
|
||||
import DeepSeekCrossSourceSamplingLab from "@/components/DeepSeekCrossSourceSamplingLab.astro";
|
||||
import { deepseekBranches, deepseekLedgers, deepseekPaperChain, deepseekWaves } from "@/data/deepseek";
|
||||
|
||||
const toc = [
|
||||
@@ -35,21 +36,22 @@ const toc = [
|
||||
["23", "behavior", "Chat:最终生成行为"],
|
||||
["24", "completion-depth", "Chat:完成度与全深度"],
|
||||
["25", "sampling", "Chat:多种子采样稳健性"],
|
||||
["26", "branches", "别漏掉旁支"],
|
||||
["27", "audit", "事实、推导与教学模型"],
|
||||
["26", "cross-source-sampling", "Chat:跨题采样与统计单位"],
|
||||
["27", "branches", "别漏掉旁支"],
|
||||
["28", "audit", "事实、推导与教学模型"],
|
||||
["↳", "papers", "六十节点阅读链"],
|
||||
];
|
||||
---
|
||||
|
||||
<BaseLayout
|
||||
title="DeepSeek 技术谱系与真实权重深读:从 Dense、MoE、MLA 到 R1 与 V4"
|
||||
description="用二十四张问题账、十次技术转向、二十个交互实验、真实 V2-Lite Base / Chat 权重、512-token 完成度评测、29 阶段隐藏状态、26 层 MoE 路由追踪与 256 条多种子采样,完整理解 DeepSeek 的 MoE、MLA、FP8、DualPipe、GRPO、R1、V3.2 与 V4。"
|
||||
description="用二十四张问题账、十次技术转向、二十一个交互实验、真实 V2-Lite Base / Chat 权重、512-token 完成度评测、29 阶段隐藏状态、26 层 MoE 路由追踪,以及单题与跨题两轮各 256 条采样,完整理解 DeepSeek 的 MoE、MLA、FP8、DualPipe、GRPO、R1、V3.2 与 V4。"
|
||||
section="deepseek"
|
||||
>
|
||||
<header class="page-hero deepseek-hero">
|
||||
<div class="page-hero-inner">
|
||||
<div>
|
||||
<p class="eyebrow"><span>SPOTLIGHT / DEEPSEEK · ROUND 06</span> SAMPLING × COMPLETION × FULL DEPTH</p>
|
||||
<p class="eyebrow"><span>SPOTLIGHT / DEEPSEEK · ROUND 07</span> SOURCE COVERAGE × SAMPLING × FULL DEPTH</p>
|
||||
<h1>不要背模型名<br />要看懂每次为什么转向</h1>
|
||||
<p class="lead">
|
||||
这不是七篇报告的摘要,而是一套可追问、可计算、可反驳的技术谱系:
|
||||
@@ -61,9 +63,9 @@ const toc = [
|
||||
<div><dt>SPAN</dt><dd>2024.01 → 2026.06</dd></div>
|
||||
<div><dt>LEDGERS</dt><dd>24 张问题账</dd></div>
|
||||
<div><dt>LINEAGE</dt><dd>10 次技术转向</dd></div>
|
||||
<div><dt>LABS</dt><dd>20 个可操作实验</dd></div>
|
||||
<div><dt>LABS</dt><dd>21 个可操作实验</dd></div>
|
||||
<div><dt>EVIDENCE</dt><dd>60 个一手 / 官方节点</dd></div>
|
||||
<div><dt>STATUS</dt><dd>六轮 · 256 条采样</dd></div>
|
||||
<div><dt>STATUS</dt><dd>七轮 · 512 条采样</dd></div>
|
||||
</dl>
|
||||
</div>
|
||||
</header>
|
||||
@@ -825,8 +827,21 @@ const toc = [
|
||||
<DeepSeekSamplingLab />
|
||||
</section>
|
||||
|
||||
<section class="article-section" id="cross-source-sampling">
|
||||
<p class="eyebrow"><span>26</span> SEEDS ARE NOT TASKS</p>
|
||||
<h2>同一道题抽 64 次,仍然只有一道题:把预算移到 16 条预先冻结的 source</h2>
|
||||
<p class="lede">
|
||||
Round 06 证明了 sampled trajectory 会跨 seed 分叉,也能在固定执行合同下逐 token
|
||||
复现;但 Math 与 Code 各只有一道题。Round 07 保持 256 条正式输出的总预算不变,
|
||||
改为四域各四条 source、每格四个 seed,只保留
|
||||
<code>system off/on × EOS/句点</code> 四格。source 是主要覆盖单位,seed 是题内重复;
|
||||
方向先逐 source 计算,不把 64 个生成误写成 64 道独立任务。
|
||||
</p>
|
||||
<DeepSeekCrossSourceSamplingLab />
|
||||
</section>
|
||||
|
||||
<section class="article-section" id="branches">
|
||||
<p class="eyebrow"><span>26</span> THE MAIN LINE IS NOT THE WHOLE TREE</p>
|
||||
<p class="eyebrow"><span>27</span> THE MAIN LINE IS NOT THE WHOLE TREE</p>
|
||||
<h2>如果只读 V2 → V3 → R1 → V4,会漏掉五条反过来影响主线的旁支</h2>
|
||||
<div class="branch-grid">
|
||||
{deepseekBranches.map(([name, line, text, url]) => (
|
||||
@@ -846,7 +861,7 @@ const toc = [
|
||||
</section>
|
||||
|
||||
<section class="article-section" id="audit">
|
||||
<p class="eyebrow"><span>27</span> EVIDENCE AUDIT</p>
|
||||
<p class="eyebrow"><span>28</span> EVIDENCE AUDIT</p>
|
||||
<h2>同一张页面里有三种知识,它们的语气必须不同</h2>
|
||||
<div class="audit-grid">
|
||||
<article class="reported">
|
||||
|
||||
@@ -145,18 +145,18 @@ const paths = [
|
||||
</a>
|
||||
<a class="release-card deepseek-release" href="/deepseek/">
|
||||
<div>
|
||||
<p class="eyebrow"><span>NEW / DEEPSEEK ROUND 06</span> MULTI-SEED SAMPLING · EXACT REPLAY</p>
|
||||
<p class="eyebrow"><span>NEW / DEEPSEEK ROUND 07</span> CROSS-SOURCE SAMPLING · SOURCE-BLOCKED AUDIT</p>
|
||||
<h2>从 Dense 到百万上下文:每次创新都在偿还上一代最贵的一张账</h2>
|
||||
<p>
|
||||
在 512-token greedy 与 29-stage / 26-gate 全深度 trace 之后,再按预注册的
|
||||
8 个 seed 生成 256 条官方 nucleus samples:251 条自然 EOS、242 条不同完整
|
||||
token 轨迹;Math 62 / 64、Code 63 / 64,并在新进程复现 R0/R1 的 64 / 64 格。
|
||||
单题抽 64 次仍然只有一道题。新一轮保持 256 条预算不变,改用 16 条预先冻结的
|
||||
source:Math 四题从 8 / 16 到 16 / 16,Code 四题也从 8 / 16 到 16 / 16;
|
||||
English 甚至出现域均值与 3 / 4 source 方向相反。新进程 R0 仍 64 / 64 格 exact。
|
||||
</p>
|
||||
</div>
|
||||
<dl>
|
||||
<div><dt>LINEAGE</dt><dd>1991 → 2026 · 10 次转向</dd></div>
|
||||
<div><dt>NODES</dt><dd>60 个一手 / 官方节点</dd></div>
|
||||
<div><dt>LAB</dt><dd>20 · Base / Chat / sampling</dd></div>
|
||||
<div><dt>LAB</dt><dd>21 · Base / Chat / sampling</dd></div>
|
||||
</dl>
|
||||
<span class="release-arrow" aria-hidden="true">进入 DeepSeek 完整技术谱系 →</span>
|
||||
</a>
|
||||
|
||||
@@ -15,7 +15,7 @@ const workstreams = [
|
||||
{ label: "表示、位置与残差高速公路", value: 81, next: "加入真实 hidden-state / norm traces、长上下文位置外推复现与更多深层稳定性消融" },
|
||||
{ label: "Scaling Laws", value: 74, next: "加入真实拟合复现、置信区间与更多模型族对照" },
|
||||
{ label: "数据工程与预训练配方", value: 73, next: "逐图精读 FineWeb / DCLM,加入真实去重与 mixture traces" },
|
||||
{ label: "DeepSeek 专题", value: 99, next: "扩大 sampling 的 source/task 覆盖并推进干预式 mediation、SM90 FlashMLA、FP8/pipeline 与 R1-like RL" },
|
||||
{ label: "DeepSeek 专题", value: 99, next: "把跨题采样扩大到可做 task-level bootstrap,再推进 per-row RNG、干预式 mediation、SM90 FlashMLA 与 R1-like RL" },
|
||||
{ label: "指令微调与人类偏好", value: 75, next: "加入真实偏好分歧样本、RM 长度偏置与 PPO/DPO 小模型复现" },
|
||||
{ label: "推理与测试时扩展", value: 76, next: "真实模型采样曲线、PRM 案例与逐篇图表精读" },
|
||||
{ label: "工具使用与长程 Agent", value: 74, next: "补真实环境 traces、cross-harness 对照、Agent RL 训练曲线与安全案例" },
|
||||
@@ -50,7 +50,7 @@ const workstreams = [
|
||||
<div><dt>OVERALL</dt><dd>专题平均 {average}%</dd></div>
|
||||
<div><dt>READABLE</dt><dd>{published} 个首版可读专题</dd></div>
|
||||
<div><dt>ACTIVE</dt><dd>{researching} 个研究/写作中</dd></div>
|
||||
<div><dt>UPDATED</dt><dd>2026-07-30 01:50 CST</dd></div>
|
||||
<div><dt>UPDATED</dt><dd>2026-07-30 04:00 CST</dd></div>
|
||||
<div><dt>MODE</dt><dd>持续迭代,不锁死版本</dd></div>
|
||||
</dl>
|
||||
</div>
|
||||
@@ -97,12 +97,12 @@ const workstreams = [
|
||||
<article><span>✓</span><h3>K3 报告已结构化拆解</h3><p>47 页报告目录、151 条参考来源和架构/后训练/系统主线已经提取。</p></article>
|
||||
<article><span>✓</span><h3>17 专题知识图</h3><p>从语言模型基础到评测安全,包含先修依赖和三条贯穿案例。</p></article>
|
||||
<article><span>✓</span><h3>编辑式网站系统</h3><p>响应式导航、章节模板、侧栏、进度、论文链和证据提示组件。</p></article>
|
||||
<article><span>✓</span><h3>八十七个原创交互视图</h3><p>K3 三轴图、八联报告实验与四联开放工件实验,DeepSeek 四联公式实验、十三联 Base 工件实验、Chat 行为、completion/full-depth 与 multi-seed sampling 三轮实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
|
||||
<article><span>✓</span><h3>八十八个原创交互视图</h3><p>K3 三轴图、八联报告实验与四联开放工件实验,DeepSeek 四联公式实验、十三联 Base 工件实验、Chat 行为、completion/full-depth、multi-seed 与 cross-source sampling 四轮实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。</p></article>
|
||||
<article><span>✓</span><h3>十七篇首版长文</h3><p>K3、语言模型前史、Transformer、表示/位置/残差、DeepSeek、Scaling、数据工程、长上下文、MoE、后训练、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全专题。</p></article>
|
||||
<article><span>✓</span><h3>语言模型前史深度专题</h3><p>八张独立问题账、33 个正式节点、20 段长文与概率—向量—记忆—对齐四联实验。</p></article>
|
||||
<article><span>✓</span><h3>Transformer 深度专题</h3><p>十张独立问题账、40 个正式节点、21 段正文与 QKV—Mask—多头位置—Block 成本四联实验。</p></article>
|
||||
<article><span>✓</span><h3>表示、位置与残差高速公路深度专题</h3><p>二十张问题账、66 个一手节点、DeepSeek/Kimi 双谱系,以及 Token—位置—Norm—Residual/FFN 四联实验。</p></article>
|
||||
<article><span>✓</span><h3>DeepSeek 六轮真实权重里程碑</h3><p>在 512-token greedy 与 29-stage / 26-gate 全深度 trace 之后,预注册 8 个 SHA-256 seed,生成 256 条 official nucleus samples:251 条 natural EOS、242 个 unique trajectory hashes;GSM8K 62 / 64 strict exact、HumanEval 63 / 64 tests pass,新进程 R0/R1 的八项合同字段 64 / 64 exact。</p></article>
|
||||
<article><span>✓</span><h3>DeepSeek 七轮真实权重里程碑</h3><p>继单题多 seed 之后,保持 256 条预算并把覆盖扩大到 16 条预先冻结的 source:250 条 natural EOS、247 个 unique trajectories;Math 四题为 14/16、16/16、8/16、9/16,Code 四题为 16/16、8/16、12/16、16/16。source-blocked 方向揭示 English 均值与多数题相反,新进程 R0 八项合同字段 64 / 64 exact。</p></article>
|
||||
<article><span>✓</span><h3>Kimi K3 技术报告二轮深读</h3><p>三十二张问题账、Figure 1–16 / Table 1–5 审计、100 节点阅读链,以及 Delta—Decay—AttnRes—LatentMoE—SiTU—QB—MOPD—Cache 八联实验。</p></article>
|
||||
<article><span>✓</span><h3>Kimi K3 三轮开放工件里程碑</h3><p>固定官方 revisions,审计 96 个 shards、497,220 个 tensor entries 与真实 KDA / MLA / MoE / MoonViT shapes;四联实验分开显示层型、tensor anatomy、参数范围和复现边界。</p></article>
|
||||
<article><span>✓</span><h3>FlashKDA RTX 5090 执行闸门</h3><p>隔离 CUDA 13.0 / glibc 2.39 编译 sm_120a wheel;6/6 官方参考逐元素相等,并完成 fixed / varlen、三种 state mode 的 1,800 个 CUDA Event samples。</p></article>
|
||||
@@ -134,7 +134,7 @@ const workstreams = [
|
||||
<div class="queue-table">
|
||||
<div class="head"><b>优先级</b><b>专题</b><b>本轮交付</b><b>完成闸门</b></div>
|
||||
<div><span>P0</span><strong>K3 三轮</strong><p>开放权重 traces → FlashKDA / AttnRes / MoE 真实行为 → Figure 1–16 数值重绘与独立复现</p><em>运行证据 + 逐图复现</em></div>
|
||||
<div><span>P0</span><strong>DeepSeek 六轮后续</strong><p>扩大任务与语言 source 的 sampling 覆盖 → 干预式 mediation → SM90 FlashMLA / FP8 / pipeline traces → R1-like RL 小模型复现</p><em>运行证据 + 独立复现</em></div>
|
||||
<div><span>P0</span><strong>DeepSeek 七轮后续</strong><p>扩大到 task-level bootstrap → per-row RNG 对照 → 干预式 mediation → SM90 FlashMLA / FP8 / pipeline traces → R1-like RL 小模型复现</p><em>运行证据 + 独立复现</em></div>
|
||||
<div><span>P0</span><strong>Transformer 二轮</strong><p>多头电路逐图 → Pre/Post-LN 真实 traces → Flash/KV 配置与 kernel 对照</p><em>逐图笔记 + 实测边界</em></div>
|
||||
<div><span>P0</span><strong>表示、位置与残差二轮</strong><p>真实 hidden-state / norm traces → 长上下文位置外推 → mHC / AttnRes 深层稳定性消融</p><em>可复现实验 + 逐图笔记</em></div>
|
||||
<div><span>P0</span><strong>语言模型前史二轮</strong><p>Kneser–Ney / LSTM / Bahdanau 逐图 → 真实小语料复现 → tokenizer 公平性</p><em>可复现实验 + 逐图笔记</em></div>
|
||||
@@ -233,6 +233,10 @@ const workstreams = [
|
||||
<div><time>2026-07-29</time><b>停止、终点、可评测与正确分四张账</b><p>自然 EOS 不等于答对;fallback 不冒充 strict completion;HumanEval tests pass 不冒充代码安全。</p></div>
|
||||
<div><time>2026-07-29</time><b>全深度比较只保留 exact interior tokens</b><p>1,537 个 content tokens / condition 在八格中 ID exact;56 个跨字符边界 token 排除,不拿不同 token 比隐藏状态。</p></div>
|
||||
<div><time>2026-07-29</time><b>表示与路由分叉不是中介因果</b><p>29-stage hidden 与 26-gate route 曲线描述传播;没有干预式 mediation 前不解释输出或能力因果。</p></div>
|
||||
<div><time>2026-07-30</time><b>seed 是题内重复,不是任务覆盖</b><p>Round 07 用 16 sources × 4 seeds × 4 cells 保持 256 条预算;所有方向先在 source 内计算,再数四题方向,不把 seed 当独立 benchmark 题。</p></div>
|
||||
<div><time>2026-07-30</time><b>任务总数必须展开为逐题矩阵</b><p>Math 与 Code 的四题都从 8/16 跨到 16/16;跨不同 GSM8K gold 的 final-answer frequency 禁止聚合。</p></div>
|
||||
<div><time>2026-07-30</time><b>均值与 source 方向同时展示</b><p>English 句点 contrast 均值 −8.5,但 3/4 source 为正;Code 两道题 +1/−1 interaction 在域均值 0 中抵消。</p></div>
|
||||
<div><time>2026-07-30</time><b>跨题 sampling 仍要求精确复跑</b><p>R0/R1 63/64 同格分叉;新进程 R0 的 seed、prompt、token、text、stop 与 CPU/CUDA RNG pre-state 八字段 64/64 exact。</p></div>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
|
||||
Reference in New Issue
Block a user