Files
llm-atlas/experiments/deepseek/README.md
T
2026-07-30 00:24:27 +08:00

675 lines
25 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DeepSeek real-weight execution probes
The probes in this directory use official DeepSeek artifacts and keep their
scope deliberately narrower than a full-model benchmark.
## DeepSeek-V2-Lite truncated trace
`v2_lite_trace.py` executes layers 0–6 from the official BF16 checkpoint. Those
seven layers are fully contained in shard 1; layer 7 is split across shards 1
and 2 and is therefore outside the default evidence boundary.
Pinned model:
```text
deepseek-ai/DeepSeek-V2-Lite@604d5664dddd88a0433dbae533b7fe9472482de0
```
Required Python stack:
```text
torch==2.11.0+cu128
transformers==4.41.2
safetensors==0.8.0
```
The 2024 remote code does not import under Transformers 5.5 because
`is_torch_fx_available` was removed. The probe imports the official files as a
read-only local package; it does not patch the model source.
Download the metadata, tokenizer, remote code, index, and first shard with the
Hugging Face CLI, then run:
```bash
python experiments/deepseek/v2_lite_trace.py \
--artifact-dir /path/to/deepseek-v2-lite \
--output src/data/deepseek-v2-lite-trace.json
```
The result contains:
- real tokenizer pieces and model-derived hidden states;
- the actual `[B,T,576]` MLA compressed projection at each executed layer;
- the expanded key/value tensors stored by the Hugging Face eager cache;
- token-level top-6 routed expert IDs and weights for six MoE layers;
- per-layer and per-prompt expert-load summaries;
- explicit boundaries against global load, expert semantics, training traces,
full-model generation, and production serving claims.
Run the probe twice and compare deterministic evidence while excluding timing:
```bash
python experiments/deepseek/compare_v2_lite_traces.py \
--first /path/to/trace-1.json \
--second /path/to/trace-2.json \
--output src/data/deepseek-v2-lite-trace-repro.json
```
## Fixed public routing corpus
`v2_lite_routing_corpus.py` keeps the same official layer 0–6 execution boundary
but replaces the four authored prompts with 128 source-addressable public
prompts:
- 32 WikiText-2 raw validation passages;
- 32 CLUE TNEWS public-test sentences;
- 32 OpenAI HumanEval prompts, without solutions/tests or code execution;
- 32 OpenAI GSM8K test questions, without answers.
Selection is the ascending SHA-256 rank of a fixed salt, domain, and source ID.
Inputs are truncated to 96 DeepSeek tokens. The six MoE layers therefore produce
304,560 actual top-6 routed-expert selections over 8,460 valid tokens.
The output includes both token-weighted and prompt-balanced distributions. Its
95% intervals use 2,000 prompt-level bootstrap resamples within each domain,
rather than treating correlated tokens as independent observations.
```bash
PYTHONPATH=/path/to/transformers-4.41.2-deps \
python -B experiments/deepseek/v2_lite_routing_corpus.py \
--artifact-dir /path/to/deepseek-v2-lite \
--human-eval /path/to/HumanEval.jsonl.gz \
--gsm8k /path/to/gsm8k/test.jsonl \
--tnews /path/to/tnews/test.json \
--tnews-archive /path/to/tnews_public.zip \
--wikitext /path/to/wikitext-validation.parquet \
--output src/data/deepseek-v2-lite-routing-corpus.json \
--per-domain 32 \
--max-tokens 96 \
--batch-size 16 \
--bootstrap 2000 \
--seed 20260729 \
--captured-at 2026-07-29T07:45:00+00:00
```
The committed independent rerun is byte-exact. Both JSON files have SHA-256:
```text
4678a1d15395de93ffba757598cc3642bf35e9e07f71d82ddd87c27fc38a09e4
```
See `research/DEEPSEEK_ROUTING_CORPUS_AUDIT.md` for corpus revisions and hashes,
metric definitions, interval semantics, results, and claim boundaries.
## Paired length-control cohort
The corpus runner can also select one fixed cohort by untruncated source length
and execute nested prefixes. The committed 16-token and 24-token traces use the
same 128 source prompts, all selected from records with at least 24 DeepSeek
tokens:
```bash
common_args=(
--per-domain 32
--batch-size 16
--bootstrap 2000
--sample-salt llm-atlas-deepseek-routing-length-control-v1
--eligibility-min-tokens 24
)
python -B experiments/deepseek/v2_lite_routing_corpus.py \
...source arguments... \
"${common_args[@]}" \
--max-tokens 16 \
--output src/data/deepseek-v2-lite-routing-matched16.json
python -B experiments/deepseek/v2_lite_routing_corpus.py \
...source arguments... \
"${common_args[@]}" \
--max-tokens 24 \
--output src/data/deepseek-v2-lite-routing-matched24.json
```
The two real traces add 184,320 top-6 route selections. Their independent
reruns are byte-exact:
```text
matched-16 f8d437d5379ffb41ac8dca5a8e97c0f44ba10ce7b63f95d7be0b7c88ac0baebd
matched-24 bed54835ad243ca2ab46bf9574e53137e6c2c0e19267f581719b5f2f65546436
```
Use the paired comparison runner to resample identical prompt indices in the
short and long traces:
```bash
python -B experiments/deepseek/compare_routing_length_control.py \
--short src/data/deepseek-v2-lite-routing-matched16.json \
--long src/data/deepseek-v2-lite-routing-matched24.json \
--output src/data/deepseek-v2-lite-routing-length-sensitivity.json \
--bootstrap 2000 \
--seed 20260729
```
See `research/DEEPSEEK_ROUTING_LENGTH_CONTROL_AUDIT.md` for the sampling bias
audit, paired CV/JSD deltas, total-variation accounting, and interpretation
boundaries.
## Official chat-template sensitivity
`v2_lite_routing_template_probe.py` renders three variants of one fixed
23-content-token prefix:
```text
raw BOS + content
user BOS + "User: " + content + "\n\n"
generation user prefix + "Assistant:"
```
The latter two use the pinned official `chat_template` through
`apply_chat_template`. All three variants of one source prompt execute in the
same padded batch. Statistics are split between:
- the full operational input, including wrapper tokens;
- the exact intersection of `(relative character span, token ID)` inside the
source content across all three variants.
The `user → generation` comparison is a causal negative control: the appended
suffix must not change routes on their shared prefix.
```bash
PYTHONPATH=/path/to/transformers-4.41.2-deps:/usr/lib/python3/dist-packages \
python -B experiments/deepseek/v2_lite_routing_template_probe.py \
--artifact-dir /path/to/deepseek-v2-lite \
--human-eval /path/to/HumanEval.jsonl.gz \
--gsm8k /path/to/gsm8k/test.jsonl \
--tnews /path/to/tnews/test.json \
--tnews-archive /path/to/tnews_public.zip \
--wikitext /path/to/wikitext-validation.parquet \
--output src/data/deepseek-v2-lite-routing-template.json \
--per-domain 32 \
--content-tokens 23 \
--batch-prompts 8 \
--layers 7 \
--bootstrap 2000 \
--seed 20260729 \
--captured-at 2026-07-29T08:30:00+00:00
```
The three variants add 382,032 real top-6 route selections. Across six MoE
layers, all 21,852 `user → generation` shared-prefix token routes are
ordered-top-6 exact. The committed run and independent rerun are byte-exact:
```text
da1f10333b2fa269e64f9716a0ca6c1a656d23d3e70b8a25d6e59f8a5b3bc1b9
```
See `research/DEEPSEEK_ROUTING_TEMPLATE_AUDIT.md` for the aligned-content
contract, paired intervals, per-token route stability, and claim boundaries.
## System × one-shot message-history factorial
`v2_lite_routing_history_factorial_probe.py` reuses the exact 128-source,
23-content-token cohort from the official-template probe and renders four
official chat histories:
```text
S0F0 target user
S1F0 fixed system + target user
S0F1 fixed demo user/assistant + target user
S1F1 fixed system + fixed demo user/assistant + target user
```
Every condition uses `add_generation_prompt=True`. The fixed system treatment
adds 16 tokens per source in both `F0` and `F1`; the fixed one-shot treatment
adds 17 tokens per source in both `S0` and `S1`. All four variants of one
source execute in the same padded batch.
The output contains:
- full-input and exact target-content scopes;
- token-weighted and prompt-balanced aggregation;
- four paired factor edges;
- system and one-shot main effects;
- difference-in-differences interaction;
- per-target-token ordered/set top-6 stability and Jaccard;
- 2,000 source-prompt bootstrap replicates shared by all four cells.
```bash
PYTHONPATH=/path/to/transformers-4.41.2-deps:/usr/lib/python3/dist-packages \
python -B experiments/deepseek/v2_lite_routing_history_factorial_probe.py \
--artifact-dir /path/to/deepseek-v2-lite \
--human-eval /path/to/HumanEval.jsonl.gz \
--gsm8k /path/to/gsm8k/test.jsonl \
--tnews /path/to/tnews/test.json \
--tnews-archive /path/to/tnews_public.zip \
--wikitext /path/to/wikitext-validation.parquet \
--output src/data/deepseek-v2-lite-routing-history-factorial.json \
--per-domain 32 \
--content-tokens 23 \
--batch-prompts 8 \
--layers 7 \
--bootstrap 2000 \
--seed 20260729 \
--captured-at 2026-07-29T09:36:00+00:00
```
The four cells add 865,440 real top-6 route selections. The committed run and
independent rerun are byte-exact:
```text
5765fbf85fa6fd948cca90f11e2237254fd58c32f373670f44530cff0a70c1fb
```
See `research/DEEPSEEK_ROUTING_HISTORY_FACTORIAL_AUDIT.md` for the factorial
definition, all CV intervals, the 24/24 system-edge TV attenuation pattern,
per-token stability, and the strict boundary between a message-history effect
and a role-semantic claim.
## Equal-length filler history control
`v2_lite_routing_history_distance_control.py` extends the preceding 2×2 probe
to six cells:
```text
system off/on × no history / repeated-token filler / fixed one-shot
```
The filler turn is:
```text
User: x x x x x x x x x
Assistant: x
```
Under the pinned official template, both filler and one-shot add exactly 17
tokens per source on both system levels. They therefore share the same
user/assistant roles, assistant EOS, target position, generation prompt, and
batch shape. The filler is deliberately called low-information rather than
semantics-free: repeated `x` tokens remain learned inputs.
```bash
PYTHONPATH=/path/to/transformers-4.41.2-deps:/usr/lib/python3/dist-packages \
python -B experiments/deepseek/v2_lite_routing_history_distance_control.py \
--artifact-dir /path/to/deepseek-v2-lite \
--human-eval /path/to/HumanEval.jsonl.gz \
--gsm8k /path/to/gsm8k/test.jsonl \
--tnews /path/to/tnews/test.json \
--tnews-archive /path/to/tnews_public.zip \
--wikitext /path/to/wikitext-validation.parquet \
--output src/data/deepseek-v2-lite-routing-history-distance-control.json \
--per-domain 32 \
--content-tokens 23 \
--batch-prompts 5 \
--layers 7 \
--bootstrap 2000 \
--seed 20260729 \
--captured-at 2026-07-29T10:17:00+00:00
```
The six cells add 1,376,496 real top-6 route selections. The committed run and
independent rerun are byte-exact:
```text
423a095d92738d4bc6c2ace2afa9b35627e61811458c7883dca7b44aee4e648e
```
For exact target content under prompt-balanced aggregation, mean system-edge
TV is `0.0738 → 0.0378 → 0.0188` for none → filler → demo. Both steps decrease
in 24/24 layer×domain cells, with all 24 paired intervals below zero. See
`research/DEEPSEEK_ROUTING_HISTORY_DISTANCE_CONTROL_AUDIT.md` for the full
table, token-level route stability, BF16 batch-shape boundary, literature
context, and non-claims.
## History-boundary single-token control
`v2_lite_routing_history_boundary_token_control.py` keeps the repeated-token
history from the preceding probe and changes exactly one token ID at the
completed assistant boundary:
```text
system off/on × official EOS / x / period / newline
```
The official template places EOS between the filler assistant content and the
next `User:` marker. The three controls replace only that EOS ID after official
tokenization. They are explicit counterfactual token sequences, not valid
official chat serializations. All eight conditions preserve sequence length,
target position, role markers, attention mask, and within-run batch shape.
```bash
PYTHONPATH=/path/to/transformers-4.41.2-deps:/usr/lib/python3/dist-packages \
python -B experiments/deepseek/v2_lite_routing_history_boundary_token_control.py \
--artifact-dir /path/to/deepseek-v2-lite \
--human-eval /path/to/HumanEval.jsonl.gz \
--gsm8k /path/to/gsm8k/test.jsonl \
--tnews /path/to/tnews/test.json \
--tnews-archive /path/to/tnews_public.zip \
--wikitext /path/to/wikitext-validation.parquet \
--output src/data/deepseek-v2-lite-routing-history-boundary-token-control.json \
--per-domain 32 \
--content-tokens 23 \
--batch-prompts 4 \
--layers 7 \
--bootstrap 2000 \
--seed 20260729 \
--captured-at 2026-07-29T11:12:00+00:00
```
The eight cells add 2,044,224 real top-6 route selections. All 256
source×system groups are equal-length and equal-position; all 768
counterfactual cells differ from their official sequence at exactly one input
ID. The committed run and independent rerun are byte-exact:
```text
9bb93834ffd8536aeebe4325e45d2179ba590554c6ff6b6fcceaba2f499b9c37
```
For exact target content under prompt-balanced aggregation, mean system-edge
TV is `.0374 / .0539 / .0553 / .0492` for EOS / x / period / newline. X and
period exceed EOS in 24/24 layer×domain cells, while newline does so in 23/24.
See `research/DEEPSEEK_ROUTING_HISTORY_BOUNDARY_TOKEN_AUDIT.md` for all paired
intervals, per-token route alignment, scope split, BF16 batch-shape audit,
primary literature, and the boundary between an input-ID intervention and a
chat-turn semantic claim.
## Role-marker-head single-token control
`v2_lite_routing_role_marker_head_control.py` keeps the official filler
history, EOS, colon, target span, generation prompt, mask, length, and 32-row
batch shape, while changing one ordinary token ID:
```text
system off/on × official / pre-target User→Assistant /
pre-target User→x / post-target Assistant→User
```
The pinned tokenizer maps `User`, `Assistant`, `:`, and `x` to IDs `5726`,
`77398`, `25`, and `87`. The pre-target controls identify the effect of the
first role-marker token only; the colon remains. The post-target replacement
is a causal suffix negative control. None of the three counterfactuals is a
valid official chat serialization.
```bash
PYTHONPATH=/path/to/transformers-4.41.2-deps:/usr/lib/python3/dist-packages \
python -B experiments/deepseek/v2_lite_routing_role_marker_head_control.py \
--artifact-dir /path/to/deepseek-v2-lite \
--human-eval /path/to/HumanEval.jsonl.gz \
--gsm8k /path/to/gsm8k/test.jsonl \
--tnews /path/to/tnews/test.json \
--tnews-archive /path/to/tnews_public.zip \
--wikitext /path/to/wikitext-validation.parquet \
--output src/data/deepseek-v2-lite-routing-role-marker-head-control.json \
--per-domain 32 \
--content-tokens 23 \
--batch-prompts 4 \
--layers 7 \
--bootstrap 2000 \
--seed 20260729 \
--captured-at 2026-07-29T12:30:00+00:00
```
The eight cells add 2,044,224 real top-6 route selections. All 256
source×system groups preserve length and target position; every one of the
768 counterfactuals differs from its official sequence at exactly one ID. The
committed run and independent rerun are byte-exact:
```text
9dc0e37fbce6581269428dcfcb84c7a17b466c5239c8171e6741f66d5eeb8caf
```
For exact target content, direct replacement TV is about `.02–.03`, but the
system-edge contrast has mixed direction: `12↑12↓` for User→Assistant and
`13↑11↓` for User→x. The post-target control is exact for all 34,488 aligned
target-token ordered top-6 routes, with zero target TV, JSD, and ΔCV. See
`research/DEEPSEEK_ROUTING_ROLE_MARKER_HEAD_AUDIT.md` for all paired intervals,
depth maps, token-level alignment, cross-experiment BF16 batch-content audit,
primary sources, and the boundary against full role semantics or Chat-model
behavior.
## Special-token family control
`v2_lite_routing_special_token_family_control.py` keeps the same repeated-token
history and changes the completed assistant boundary to four single IDs:
```text
system off/on × EOS / BOS / x / period
```
The pinned tokenizer has exactly two special-token IDs: BOS `100000` and EOS
`100001`; PAD aliases EOS. EOS/BOS therefore exhaust the special inventory,
while `x` and period are only two selected ordinary controls. The 2-vs-2
family summary is descriptive for these four IDs and is not a
population-level specialness claim.
```bash
PYTHONPATH=/path/to/transformers-4.41.2-deps:/usr/lib/python3/dist-packages \
python -B experiments/deepseek/v2_lite_routing_special_token_family_control.py \
--artifact-dir /path/to/deepseek-v2-lite \
--human-eval /path/to/HumanEval.jsonl.gz \
--gsm8k /path/to/gsm8k/test.jsonl \
--tnews /path/to/tnews/test.json \
--tnews-archive /path/to/tnews_public.zip \
--wikitext /path/to/wikitext-validation.parquet \
--output src/data/deepseek-v2-lite-routing-special-token-family-control.json \
--per-domain 32 \
--content-tokens 23 \
--batch-prompts 4 \
--layers 7 \
--bootstrap 2000 \
--seed 20260729 \
--captured-at 2026-07-29T12:57:00+00:00
```
The eight cells add 2,044,224 real top-6 route selections. All 256
source×system groups preserve length and target position, and all 768
counterfactuals change exactly one ID. The committed run and independent
rerun are byte-exact:
```text
c372c1b03a8b15f615b54ded5d9257a8fc2cdb7728001735d3d4c1d8534af5bf
```
For exact target content under prompt-balanced aggregation, mean system-edge
TV is `.037518 / .048018 / .053975 / .055289` for EOS / BOS / x / period.
BOS−EOS is positive in 22/24 layer×domain cells; x−EOS and period−EOS are
positive in 24/24. The selected ordinary-control mean exceeds the complete
special-inventory mean by `.011864` in these four IDs only. See
`research/DEEPSEEK_ROUTING_SPECIAL_TOKEN_FAMILY_AUDIT.md` for all paired
intervals, direct edges, scope split, alignment, batch-content audit,
literature context, and non-claims.
## Full two-token role-marker factorial
`v2_lite_routing_role_marker_block_factorial.py` treats the pre-target
two-token marker as two independent factors:
```text
system off/on × head User/Assistant × delimiter colon/x
User: [5726, 25] Assistant: [77398, 25]
User x [5726, 87] Assistant x [77398, 87]
```
The official assistant EOS and post-target generation suffix remain unchanged.
The four blocks have zero, one, one, and two edited IDs relative to official
`User:`; every edit is verified while length, target position, attention mask,
and 32-row batch shape stay fixed.
```bash
PYTHONPATH=/path/to/transformers-4.41.2-deps:/usr/lib/python3/dist-packages \
python -B experiments/deepseek/v2_lite_routing_role_marker_block_factorial.py \
--artifact-dir /path/to/deepseek-v2-lite \
--human-eval /path/to/HumanEval.jsonl.gz \
--gsm8k /path/to/gsm8k/test.jsonl \
--tnews /path/to/tnews/test.json \
--tnews-archive /path/to/tnews_public.zip \
--wikitext /path/to/wikitext-validation.parquet \
--output src/data/deepseek-v2-lite-routing-role-marker-block-factorial.json \
--per-domain 32 \
--content-tokens 23 \
--batch-prompts 4 \
--layers 7 \
--bootstrap 2000 \
--seed 20260729 \
--captured-at 2026-07-29T13:06:00+00:00
```
The eight cells add 2,044,224 real top-6 route selections. The 256 official,
512 one-ID, and 256 two-ID cells all pass their exact edit contracts. The
committed run and independent rerun are byte-exact:
```text
a703dddb6d04182b1a608c213bfd341cfc32e1801be42c417dca30a505087e82
```
For exact target content under prompt-balanced aggregation, the four mean
system-edge TVs are `.037304 / .037015 / .037239 / .036220`. The head,
delimiter, and head×delimiter effects are small and mixed across the 24
layer×domain cells. Direct head and delimiter edges remain nonzero; replacing
colon with the single `x` control modestly reduces User-vs-Assistant direct TV
in most cells, without a uniform system modulation. See
`research/DEEPSEEK_ROUTING_ROLE_MARKER_BLOCK_AUDIT.md` for the exact factor
coding, intervals, direct dependencies, full-input split, route alignment,
cross-batch audit, sources, and non-claims.
## Full Chat-checkpoint generation behavior
`v2_lite_chat_special_token_behavior_probe.py` carries the special-token grid
into the full official SFT Chat checkpoint:
```text
deepseek-ai/DeepSeek-V2-Lite-Chat
@85864749cd611b4353ce1decdb286193298f64c7
system off/on × EOS / BOS / x / period
```
It reuses the pinned Base-routing source IDs and ranks but generates from the
full source text. All eight variants of one source run in one left-padded
batch. Decoding is deterministic greedy; BOS / x / period remain invalid-chat
single-ID counterfactuals.
The official BF16 checkpoint contains `31,412,968,448` tensor bytes. Because
the local RTX 5090 has less than the official 40GB single-GPU boundary, the
formal run uses Accelerate `device_map=auto`: embeddings and layers 0–24 are
placed on CUDA, while layers 25–26, final norm, and LM head are CPU-offloaded.
```bash
PYTHONPATH=/path/to/transformers-4.41.2-deps \
python -B experiments/deepseek/v2_lite_chat_special_token_behavior_probe.py \
--artifact-dir /path/to/deepseek-v2-lite-chat \
--reference-routing-json \
src/data/deepseek-v2-lite-routing-special-token-family-control.json \
--human-eval /path/to/HumanEval.jsonl.gz \
--gsm8k /path/to/gsm8k/test.jsonl \
--tnews /path/to/tnews/test.json \
--tnews-archive /path/to/tnews_public.zip \
--wikitext /path/to/wikitext-validation.parquet \
--output src/data/deepseek-v2-lite-chat-behavior.json \
--per-domain 4 \
--max-new-tokens 128 \
--gpu-memory 29GiB \
--cpu-memory 80GiB
```
The formal artifact contains 16 sources and 128 outputs. Only 31 outputs hit
EOS; 97 reach the 128-token cap, so task-score fields are diagnostic and are
not used for capability ranking. A 4-source / 32-cell long-sequence rerun is
exact for prompt hashes, generated token IDs, decoded text, and EOS state.
```text
formal 54496955d0dd20a3116e6b758a55f56e2444c43dd99bfef94a2cb7b140e6196e
repro b74c31606c0fc70a8e52cc6c3d6135c4e916d698b4d5dae51863124d025b0c72
```
See `research/DEEPSEEK_V2_LITE_CHAT_BEHAVIOR_PROTOCOL.md` and
`research/DEEPSEEK_V2_LITE_CHAT_BEHAVIOR_AUDIT.md` for the preregistered
contract, output-divergence table, offload device map, truncation boundary,
rerun coverage, execution fixes, primary sources, and forbidden conclusions.
## Completion-aware 512-token evaluation
The completion run uses the same 16 sources and eight-condition batches but
reruns all 128 cells with one uniform 512-token budget. The successful formal
placement leaves embeddings and layers 0–23 on CUDA and offloads layers 24–26,
final norm, and LM head. `expandable_segments:True` is part of the recorded
runtime contract.
```bash
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
PYTHONPATH=/path/to/transformers-4.41.2-deps \
python -B experiments/deepseek/v2_lite_chat_special_token_behavior_probe.py \
--artifact-dir /path/to/deepseek-v2-lite-chat \
--reference-routing-json \
src/data/deepseek-v2-lite-routing-special-token-family-control.json \
--human-eval /path/to/HumanEval.jsonl.gz \
--gsm8k /path/to/gsm8k/test.jsonl \
--tnews /path/to/tnews/test.json \
--tnews-archive /path/to/tnews_public.zip \
--wikitext /path/to/wikitext-validation.parquet \
--output src/data/deepseek-v2-lite-chat-completion-512.json \
--per-domain 4 \
--max-new-tokens 512 \
--gpu-memory 28GiB \
--cpu-memory 80GiB
```
The evaluator separates stopping, semantic terminal state, evaluator coverage,
and correctness. Each HumanEval candidate runs in a fresh pinned, networkless,
read-only Docker container with no host mounts:
```bash
python -B experiments/deepseek/v2_lite_chat_completion_evaluator.py \
--behavior-json \
src/data/deepseek-v2-lite-chat-completion-512.json \
--baseline-json src/data/deepseek-v2-lite-chat-behavior.json \
--human-eval /path/to/HumanEval.jsonl.gz \
--gsm8k /path/to/gsm8k/test.jsonl \
--sandbox-image \
python:3.11-alpine@sha256:25976e9d34a0fab1f278cae931f34c8303d97bf0c0d7f85b6b4dcf641d7702a4 \
--output \
src/data/deepseek-v2-lite-chat-completion-512-eval.json
```
The formal run reaches natural EOS in 121/128 cells, up from 31/128 at the
128-token budget. Strict-complete GSM8K numeric exact is 23/32; HumanEval
official-test pass is 24/32. These are four source tasks per domain, not
benchmark estimates. A fresh-process one-source-per-domain rerun reproduces
all 32 complete generated token sequences.
## Full 27-layer Chat hidden-state and router trace
`v2_lite_chat_full_depth_trace.py` performs a prompt-only `use_cache=False`
forward on the same checkpoint and inputs. It captures hashes and statistics
for embedding, every decoder-layer output, final norm, and every MoE gate.
Boundary-crossing tokenizer tokens are excluded from the exact target-content
scope.
```bash
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
PYTHONPATH=/path/to/transformers-4.41.2-deps \
python -B experiments/deepseek/v2_lite_chat_full_depth_trace.py \
--artifact-dir /path/to/deepseek-v2-lite-chat \
--reference-routing-json \
src/data/deepseek-v2-lite-routing-special-token-family-control.json \
--human-eval /path/to/HumanEval.jsonl.gz \
--gsm8k /path/to/gsm8k/test.jsonl \
--tnews /path/to/tnews/test.json \
--tnews-archive /path/to/tnews_public.zip \
--wikitext /path/to/wikitext-validation.parquet \
--output src/data/deepseek-v2-lite-chat-full-depth.json \
--per-domain 4 \
--gpu-memory 28GiB \
--cpu-memory 80GiB
```
The formal trace covers 1,537 exact interior content tokens per condition,
29 hidden stages, 26 gates, and 1,918,176 top-6 route decisions. A fresh
four-source rerun reproduces 1,856 hidden tensor hashes, 1,664 ordered route
hashes, 1,664 route-weight hashes, and every derived comparison.
See `research/DEEPSEEK_V2_LITE_CHAT_COMPLETION_PROTOCOL.md` and
`research/DEEPSEEK_V2_LITE_CHAT_COMPLETION_DEPTH_AUDIT.md` for the OOM
amendment, task evaluators, completion table, depth curves, reproduction
audit, artifact hashes, and claim boundaries.