Files
llm-atlas/experiments/deepseek

DeepSeek real-weight execution probes

The probes in this directory use official DeepSeek artifacts and keep their scope deliberately narrower than a full-model benchmark.

DeepSeek-V2-Lite truncated trace

v2_lite_trace.py executes layers 0–6 from the official BF16 checkpoint. Those seven layers are fully contained in shard 1; layer 7 is split across shards 1 and 2 and is therefore outside the default evidence boundary.

Pinned model:

deepseek-ai/DeepSeek-V2-Lite@604d5664dddd88a0433dbae533b7fe9472482de0

Required Python stack:

torch==2.11.0+cu128
transformers==4.41.2
safetensors==0.8.0

The 2024 remote code does not import under Transformers 5.5 because is_torch_fx_available was removed. The probe imports the official files as a read-only local package; it does not patch the model source.

Download the metadata, tokenizer, remote code, index, and first shard with the Hugging Face CLI, then run:

python experiments/deepseek/v2_lite_trace.py \
  --artifact-dir /path/to/deepseek-v2-lite \
  --output src/data/deepseek-v2-lite-trace.json

The result contains:

  • real tokenizer pieces and model-derived hidden states;
  • the actual [B,T,576] MLA compressed projection at each executed layer;
  • the expanded key/value tensors stored by the Hugging Face eager cache;
  • token-level top-6 routed expert IDs and weights for six MoE layers;
  • per-layer and per-prompt expert-load summaries;
  • explicit boundaries against global load, expert semantics, training traces, full-model generation, and production serving claims.

Run the probe twice and compare deterministic evidence while excluding timing:

python experiments/deepseek/compare_v2_lite_traces.py \
  --first /path/to/trace-1.json \
  --second /path/to/trace-2.json \
  --output src/data/deepseek-v2-lite-trace-repro.json

Fixed public routing corpus

v2_lite_routing_corpus.py keeps the same official layer 0–6 execution boundary but replaces the four authored prompts with 128 source-addressable public prompts:

  • 32 WikiText-2 raw validation passages;
  • 32 CLUE TNEWS public-test sentences;
  • 32 OpenAI HumanEval prompts, without solutions/tests or code execution;
  • 32 OpenAI GSM8K test questions, without answers.

Selection is the ascending SHA-256 rank of a fixed salt, domain, and source ID. Inputs are truncated to 96 DeepSeek tokens. The six MoE layers therefore produce 304,560 actual top-6 routed-expert selections over 8,460 valid tokens.

The output includes both token-weighted and prompt-balanced distributions. Its 95% intervals use 2,000 prompt-level bootstrap resamples within each domain, rather than treating correlated tokens as independent observations.

PYTHONPATH=/path/to/transformers-4.41.2-deps \
python -B experiments/deepseek/v2_lite_routing_corpus.py \
  --artifact-dir /path/to/deepseek-v2-lite \
  --human-eval /path/to/HumanEval.jsonl.gz \
  --gsm8k /path/to/gsm8k/test.jsonl \
  --tnews /path/to/tnews/test.json \
  --tnews-archive /path/to/tnews_public.zip \
  --wikitext /path/to/wikitext-validation.parquet \
  --output src/data/deepseek-v2-lite-routing-corpus.json \
  --per-domain 32 \
  --max-tokens 96 \
  --batch-size 16 \
  --bootstrap 2000 \
  --seed 20260729 \
  --captured-at 2026-07-29T07:45:00+00:00

The committed independent rerun is byte-exact. Both JSON files have SHA-256:

4678a1d15395de93ffba757598cc3642bf35e9e07f71d82ddd87c27fc38a09e4

See research/DEEPSEEK_ROUTING_CORPUS_AUDIT.md for corpus revisions and hashes, metric definitions, interval semantics, results, and claim boundaries.