Files

1.5 KiB

Attention Residuals spike-path diagnostics

This directory implements preregistered protocol llm-atlas-k3-attnres-spike-path-v1.

It is a targeted follow-up to Round 05. It replays the exact depth-32 Block training contract and adds diagnostic-only activation positions, gradient reductions, and same-forward backward-rule interventions. It is not a Kimi K3 checkpoint run and does not recover the paper's unpublished Figure 5 telemetry.

Frozen environment

Python /home/wuyang/.pyenv/versions/3.10.14/envs/navi-router-cu128/bin/python
PyTorch 2.11.0+cu128
GPU NVIDIA GeForce RTX 5090
CUBLAS_WORKSPACE_CONFIG=:4096:8

Step-0 smoke

CUBLAS_WORKSPACE_CONFIG=:4096:8 \
/home/wuyang/.pyenv/versions/3.10.14/envs/navi-router-cu128/bin/python \
  experiments/k3/attnres_spike/train.py \
  --run-kind smoke \
  --seed 2026073001 \
  --cache-dir /home/wuyang/.cache/llm-atlas/k3-attnres-gradient-scale-v1 \
  --parent-manifest experiments/k3/attnres_gradient/manifest.json \
  --manifest experiments/k3/attnres_spike/manifest.json \
  --output /home/wuyang/.cache/llm-atlas/k3-attnres-spike-path-v1/smoke/seed-2026073001.json

Formal cells use --run-kind formal and all three preregistered seeds. The independent replay uses --run-kind replay --seed 2026073001. Formal and replay runs are fixed to 8,000 steps; smoke performs the complete step-0 diagnostic gate without an optimizer step.

Raw outputs are copied into results/raw/ only after training equivalence, forward identity, loss-scale, reduction, and replay gates pass.