From f7670efcddf663b0bd52642d60c9e19eebeb3a53 Mon Sep 17 00:00:00 2001 From: wuyang <5700876+banisherwy@user.noreply.gitee.com> Date: Thu, 30 Jul 2026 10:24:42 +0800 Subject: [PATCH] feat: add AttnRes gradient scale lab --- PROGRESS.md | 16 +- README.md | 15 +- package.json | 2 + scripts/check-k3-attnres-gradient-browser.mjs | 231 +++++ scripts/check-k3-attnres-gradient-data.mjs | 68 ++ scripts/check-k3-browser.mjs | 10 +- src/components/K3AttnResGradientLab.astro | 813 ++++++++++++++++++ src/pages/k3/index.astro | 40 +- src/pages/progress/index.astro | 13 +- 9 files changed, 1191 insertions(+), 17 deletions(-) create mode 100644 scripts/check-k3-attnres-gradient-browser.mjs create mode 100644 scripts/check-k3-attnres-gradient-data.mjs create mode 100644 src/components/K3AttnResGradientLab.astro diff --git a/PROGRESS.md b/PROGRESS.md index 755b659..ecc800f 100644 --- a/PROGRESS.md +++ b/PROGRESS.md @@ -41,7 +41,7 @@ - [x] 完成 486 篇关键论文索引,覆盖 16 个标签专题与 Kimi/DeepSeek 聚光主线。 - [x] 完成可检索、可按专题筛选的论文库页面。 - [x] 完成 K3、语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全十七篇首版长文。 -- [x] 完成 K3 三轴架构、八联报告实验、四联开放工件实验与五联 AttnRes 独立实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 二十二联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等九十四个原创交互视图。 +- [x] 完成 K3 三轴架构、八联报告实验、四联开放工件实验、Round 04 / 05 各五联 AttnRes 独立实验、语言模型前史四联实验、Transformer 四联实验、表示深度四联实验、DeepSeek 二十二联实验、长上下文、MoE 路由、推理三页签,以及训练系统、推理服务、Scaling、数据工程、数值、Alignment、Agent、原生多模态与评测安全专题各四页签等九十九个原创交互视图。 - [x] 完成长上下文首版:五张成本账、26 篇一手论文、10+ 机制图与 8 策略交互实验室。 - [x] 核验 FlashAttention、DeepSeek-V2/V3.2/V4、Kimi Linear/K3 等六份论文原文,并建立长上下文研究账本。 - [x] 核验 Switch、ST-MoE、DeepSeekMoE、Loss-Free、V3、LatentMoE 与 K3 原文,并建立 MoE 研究账本。 @@ -279,10 +279,18 @@ - [x] K3 Round 04 五视图实验室完成:三 seed BPC 曲线、Residual RMS / Block 锯齿、Full / Block depth-weight heatmap、梯度反证、成本/哈希/claim boundary 分开展示;完整 9-run JSON、compact 数据、复现清单、协议、审计、训练与聚合代码进入公开仓库。 - [x] Round 04 本地闸门通过:91 个受检文件零诊断/提示,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;冻结数据、AttnRes 专项与 K3 全量真实 Chrome 回归通过,桌面/390px 移动端零文档级溢出、零运行时异常。 - [x] K3 Round 04 以功能源提交 `4ce780d`、不可变镜像 `20260729T233142Z-4ce780d` 发布;OCI index digest `sha256:6e89f802…25f582`,复用 NAS `12010→8080`、NPM host 31 / cert 41 与门户 `LLM ATLAS / projects / 180`。容器 healthy、0 次重启,21/21 公网页面、HTTPS/2、gzip / immutable assets、AttnRes 专项与 K3 全量生产 Chrome 回归通过;保留 `20260729T221654Z-975ed3d` 回滚。 +- [x] K3 Round 05 一手定义审计确认 Figure 5(c) 未公开 gradient tensor、norm、reduction、diagnostic batch、AMP / clipping 时点或统计代码;Round 04 参数梯度与 Round 05 post-MLP output activation gradient 永久分对象记账,不把本站 operationalization 冒充作者实现。 +- [x] 在任何 formal 输出前冻结 16 / 32 blocks、Baseline / Block、三个 seed、8,000 steps、六个诊断点、CV + 首尾四分位 imbalance 联合判据、FP32 residual accumulator、20-step 双 smoke 与 depth-32 Block 完整 replay;768,000-window schedule SHA-256 为 `5041e09b…f4e`。 +- [x] 12 个 formal 格全部完成,共 786,432,000 target bytes;公共主干与三个 gate input tensor hashes 在 paired 架构间 exact。Block 的验证 BPC 在 6 / 6 配对中更低,depth-16 / 32 mean delta 为 `−.008938 / −.009872`,但不追加事后 BPC support 阈值。 +- [x] activation-gradient 结果分裂:首/末四分位 imbalance 在 6 / 6 配对改善,depth-16 / 32 均值为 `+61.0% / +72.0%`;全层 CV 却在 6 / 6 配对恶化,均值相对 reduction 为 `−10.3% / −60.0%`。两个 depth 都按预注册规则判为 `mixed / inconclusive`,总判定 `depth-dependent or inconclusive`。 +- [x] 绝对 gradient mean 仅为 Baseline 的 `57.4% / 54.4%`;参数 gradient CV 从 `0.416→0.683 / 0.397→0.772`,继续保留反结果。Output RMS 最后/第一层比则由 Baseline `4.59× / 6.08×` 降至 Block `1.17× / 1.89×`。 +- [x] 指定 depth-32 / Block / seed-2026073001 从初始化完整重训 8,000 steps;排除 run-kind / timing 后冻结字段 compare SHA-256 同为 `46300a45…4817`,model / optimizer state hashes exact。正式/compact/reproduction 物理 SHA-256 为 `ad461cbe…a8d / 5377a5e7…68e3 / aedcde6a…dea6`。 +- [x] K3 Round 05 五视图实验室完成:论文定义已知/未定义、绝对/归一化深度谱、六 checkpoint 时间轨迹、Output RMS 组节律、activation/parameter/BPC/成本/重放联合账全部可切换;21 个 raw JSON、完整 aggregate、compact、runner、analyzer、协议与审计进入公开树。 +- [x] Round 05 本地闸门通过:94 个 Astro 文件零诊断/提示,21 个页面、1,151 个站内引用、12 个跨页锚点零失败;冻结数据、新专项、Round 04 与 K3 全量真实 Chrome 回归通过,桌面/390px 移动端零文档级溢出、零运行时异常。 ## 正在进行 -- [ ] K3 四轮下一闸门:对齐 AttnRes 论文的 activation / residual-output gradient 定义,增加模型深度与训练预算,检验本轮梯度反结果是否随尺度翻转;真实 K3 forward 继续等待 `A_log [128]→[96]` 官方转换或权重修订。 +- [ ] K3 五轮下一闸门:对齐 layer 21–25 的 activation-gradient 尖峰、pre-attention / pre-MLP 位置与 mixer source weights,并做公开 reduction sensitivity;真实 K3 forward 继续等待 `A_log [128]→[96]` 官方转换或权重修订。 - [ ] DeepSeek 八轮下一闸门:推进干预式 mediation、SM90 FlashMLA、FP8 / pipeline traces 与 R1-like RL 小模型复现。 - [ ] 表示、位置与残差二轮:真实 hidden-state / norm traces、长上下文位置外推复现与 mHC / AttnRes 深层稳定性消融。 - [ ] 评测安全二轮:真实 cross-harness / pass@k 复跑、Judge 元评测、动态污染与过拒案例。 @@ -476,6 +484,10 @@ | 2026-07-30 | 主结果与机制反结果同时发布 | Full / Block BPC 方向支持;核心参数 gradient RMS CV 却高于 Baseline,明确写成未复现论文梯度叙述 | | 2026-07-30 | 独立重放按数值合同而非计时合同验收 | Block / seed-1 的八组冻结字段 2,000 steps exact;wall time 受调度影响,不要求或声称 bit-exact | | 2026-07-30 | K3 Round 04 缩小 AttnRes 里程碑发布 | 功能源 `4ce780d`、镜像 `20260729T233142Z-4ce780d`、OCI `sha256:6e89f802…25f582`;21/21 公网页面与生产专项/全量 Chrome 通过,保留 Round 08 回滚点 | +| 2026-07-30 | AttnRes 的“梯度”先按公开证据拆对象 | 论文 Figure 5 没有公开唯一 telemetry 合同;参数梯度与 post-MLP activation gradient 不再互相代称 | +| 2026-07-30 | “更均匀”拆成首尾失衡与全层 CV | Block 6/6 改善 first/last,却 6/6 恶化 CV;局部尖峰与系统性早层隆起必须分开解释 | +| 2026-07-30 | 绝对尺度与归一化形状永久同报 | Block activation-gradient mean 约为 Baseline 54%–57%;不能把更接近 1 的首尾比自动解释为各层信号更强 | +| 2026-07-30 | Round 05 完整重放过闸 | depth-32 Block seed-1 从零重训 8,000 steps;全部冻结字段与 model/optimizer state hashes exact,timing 仍单独报告 | ## 未决问题 diff --git a/README.md b/README.md index 28b4029..4388f46 100644 --- a/README.md +++ b/README.md @@ -19,7 +19,7 @@ 当前里程碑包含 17 专题学习地图、486 篇关键论文索引、Kimi K3 完整导读, 语言模型前史、Transformer 基础、表示/位置/残差、DeepSeek 技术谱系、Scaling Laws、数据工程、长上下文、MoE、指令微调与人类偏好、推理、工具使用与长程 Agent、原生多模态、训练系统、推理服务、数值优化,以及评测与安全深度专题, -以及 94 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、 +以及 99 个覆盖核心机制的原创交互视图。K3 二轮导读以 32 张问题账、16 图 / 5 表审计、 8 个交互实验和 100 个一手/官方节点,完整覆盖架构、预训练、后训练、系统、评测、案例与附录。 第三轮已完成开放工件与首个真实 kernel 里程碑:固定官方模型与 FlashKDA revisions,审计 96 个 checkpoint shards、 497,220 个 tensor entries、真实 KDA / MLA / MoE / MoonViT shapes 与小范围参数统计,并用 4 个新视图 @@ -39,6 +39,19 @@ [K3_ATTNRES_REDUCED_PROTOCOL.md](./research/K3_ATTNRES_REDUCED_PROTOCOL.md)、 [K3_ATTNRES_REDUCED_AUDIT.md](./research/K3_ATTNRES_REDUCED_AUDIT.md) 与 [AttnRes experiment](./experiments/k3/attnres/)。 +第五轮先审计 Attention Residuals Figure 5 的公开定义边界,再冻结 +`llm-atlas-k3-attnres-gradient-scale-v1`:以 post-MLP block output activation gradient +为公开 operationalization,把深度扩为 16 / 32 blocks、预算扩为每格 8,000 steps, +完成 Baseline / Block × 三 seed 共 12 格、786,432,000 formal target bytes。结果把 +“更均匀”拆成两个相反方向:Block 在 6 / 6 配对中把首/末四分位失衡改善 56%–81%, +却因中后段局部尖峰让全层 CV 在 6 / 6 配对中恶化;两个深度都按预注册联合规则判为 +mixed / inconclusive。验证 BPC 仍在 6 / 6 配对中更低,但实际 step time 约 2.6×、 +peak allocated memory 约 2.2×,不冒充同算力优势。指定 depth-32 / Block / seed-1 +从零重训完整 8,000 steps,全部冻结字段以及 model / optimizer state hashes exact。 +详见 +[K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md](./research/K3_ATTNRES_GRADIENT_DEFINITION_AUDIT.md)、 +[K3_ATTNRES_GRADIENT_SCALE_AUDIT.md](./research/K3_ATTNRES_GRADIENT_SCALE_AUDIT.md) 与 +[gradient experiment](./experiments/k3/attnres_gradient/)。 DeepSeek 八轮专题以 24 张问题账、10 次技术转向、 22 个交互实验和 60 个一手/官方节点,串起 Dense、MoE、MLA、V3 协同、R1 与 V4; 并固定官方 V2-Lite revision,在 RTX 5090 上连续执行 7/27 层,记录 3,240 次真实专家选择、 diff --git a/package.json b/package.json index d3ad81f..142a2f1 100644 --- a/package.json +++ b/package.json @@ -20,6 +20,7 @@ "build:data:deepseek-chat-task-bootstrap": "node scripts/build-deepseek-chat-task-bootstrap-crn-compact.mjs", "check:data:deepseek-chat-task-bootstrap": "node scripts/check-deepseek-chat-task-bootstrap-crn-data.mjs", "check:data:k3-attnres": "node scripts/check-k3-attnres-data.mjs", + "check:data:k3-attnres-gradient": "node scripts/check-k3-attnres-gradient-data.mjs", "check:site": "node scripts/check-site.mjs", "check:moe-browser": "node scripts/check-moe-browser.mjs", "check:reasoning-browser": "node scripts/check-reasoning-browser.mjs", @@ -40,6 +41,7 @@ "check:deepseek-cross-source-sampling-browser": "node scripts/check-deepseek-cross-source-sampling-browser.mjs", "check:deepseek-task-bootstrap-browser": "node scripts/check-deepseek-task-bootstrap-browser.mjs", "check:k3-attnres-browser": "node scripts/check-k3-attnres-browser.mjs", + "check:k3-attnres-gradient-browser": "node scripts/check-k3-attnres-gradient-browser.mjs", "check:k3-browser": "node scripts/check-k3-browser.mjs" }, "dependencies": { diff --git a/scripts/check-k3-attnres-gradient-browser.mjs b/scripts/check-k3-attnres-gradient-browser.mjs new file mode 100644 index 0000000..a317a82 --- /dev/null +++ b/scripts/check-k3-attnres-gradient-browser.mjs @@ -0,0 +1,231 @@ +import { writeFileSync } from "node:fs"; + +const cdpPort = process.env.CDP_PORT ?? "9228"; +const baseUrl = process.env.SITE_URL ?? "http://127.0.0.1:4328"; +const pages = await fetch(`http://127.0.0.1:${cdpPort}/json/list`).then((response) => response.json()); +const page = pages.find((entry) => entry.type === "page"); +if (!page) throw new Error(`CDP ${cdpPort} 没有可用页面`); + +const socket = new WebSocket(page.webSocketDebuggerUrl); +await new Promise((resolve, reject) => { + socket.addEventListener("open", resolve, { once: true }); + socket.addEventListener("error", reject, { once: true }); +}); + +let nextId = 0; +const pending = new Map(); +const exceptions = []; +socket.addEventListener("message", (event) => { + const message = JSON.parse(event.data); + if (message.id && pending.has(message.id)) { + const { resolve, reject } = pending.get(message.id); + pending.delete(message.id); + if (message.error) reject(new Error(message.error.message)); + else resolve(message.result); + } + if (message.method === "Runtime.exceptionThrown") { + exceptions.push(message.params.exceptionDetails.exception?.description ?? message.params.exceptionDetails.text); + } +}); + +const command = (method, params = {}) => new Promise((resolve, reject) => { + const id = ++nextId; + pending.set(id, { resolve, reject }); + socket.send(JSON.stringify({ id, method, params })); +}); +const pause = (milliseconds) => new Promise((resolve) => setTimeout(resolve, milliseconds)); +const evaluate = async (expression) => { + const result = await command("Runtime.evaluate", { expression, returnByValue: true, awaitPromise: true }); + if (result.exceptionDetails) throw new Error(result.exceptionDetails.exception?.description ?? result.exceptionDetails.text); + return result.result.value; +}; +const navigate = async (path) => { + await command("Page.navigate", { url: `${baseUrl}${path}` }); + for (let attempt = 0; attempt < 100; attempt += 1) { + await pause(100); + if (await evaluate("document.readyState === 'complete'")) return; + } + throw new Error(`${path} 加载超时`); +}; +const screenshot = async (path) => { + const result = await command("Page.captureScreenshot", { format: "png", captureBeyondViewport: false }); + writeFileSync(path, Buffer.from(result.data, "base64")); +}; + +await command("Page.enable"); +await command("Runtime.enable"); +await command("Emulation.setDeviceMetricsOverride", { + width: 1440, + height: 1100, + deviceScaleFactor: 1, + mobile: false, +}); +await navigate("/k3/"); + +const desktop = await evaluate(`(() => { + const root = document.querySelector("[data-gradient-lab]"); + root.scrollIntoView({ block: "start", behavior: "instant" }); + window.scrollBy(0, -78); + const text = (selector) => root.querySelector(selector)?.textContent.trim(); + const panel = () => root.querySelector("[data-gradient-panel]:not([hidden])")?.dataset.gradientPanel; + const points = (selector) => root.querySelector(selector)?.getAttribute("points"); + const setSelect = (selector, value) => { + const node = root.querySelector(selector); + node.value = value; + node.dispatchEvent(new Event("change", { bubbles: true })); + }; + + const initial = { + panel: panel(), + tabs: root.querySelectorAll("[data-gradient-tab]").length, + panels: root.querySelectorAll("[data-gradient-panel]").length, + ledger: root.querySelectorAll(".gradient-ledger article").length, + definition: root.textContent.includes("论文作者就是这样算的") && + root.textContent.includes("activation gradient") && + root.textContent.includes("parameter gradient"), + }; + + root.querySelector('[data-gradient-tab="spectrum"]').click(); + const spectrumInitial = { + panel: panel(), + baseCv: text("[data-spectrum-base-cv]"), + blockCv: text("[data-spectrum-block-cv]"), + baseRatio: text("[data-spectrum-base-ratio]"), + blockRatio: text("[data-spectrum-block-ratio]"), + baseLine: points('[data-chart-line="baseline"]'), + blockLine: points('[data-chart-line="block"]'), + boundaries: root.querySelectorAll("[data-chart-groups] .group-boundary").length, + }; + root.querySelector('[data-spectrum-scale="normalized"]').click(); + const normalized = { + title: text("[data-spectrum-title]"), + baseLine: points('[data-chart-line="baseline"]'), + }; + root.querySelector('[data-spectrum-depth="16"]').click(); + const depth16 = { + baseCv: text("[data-spectrum-base-cv]"), + blockCv: text("[data-spectrum-block-cv]"), + pointCount: points('[data-chart-line="baseline"]').split(" ").length, + boundaries: root.querySelectorAll("[data-chart-groups] .group-boundary").length, + }; + setSelect("[data-spectrum-seed]", "2026073001"); + setSelect("[data-spectrum-step]", "2000"); + const seedStep = { + state: text("[data-spectrum-state]"), + baseCv: text("[data-spectrum-base-cv]"), + blockCv: text("[data-spectrum-block-cv]"), + }; + + root.querySelector('[data-gradient-tab="timeline"]').click(); + const timelineInitial = { + panel: panel(), + title: text("[data-time-title]"), + line: points('[data-time-line="block"]'), + pointCount: root.querySelectorAll('[data-time-points="block"] circle').length, + }; + root.querySelector('[data-time-metric="imbalance"]').click(); + const timelineChanged = { + title: text("[data-time-title]"), + line: points('[data-time-line="block"]'), + }; + + root.querySelector('[data-gradient-tab="output"]').click(); + const outputInitial = { + panel: panel(), + pointCount: points('[data-output-line="block"]').split(" ").length, + bars: root.querySelectorAll("[data-output-bars] i").length, + boundaries: root.querySelectorAll("[data-output-groups] .group-boundary").length, + copy: text("[data-output-copy]"), + }; + root.querySelector('[data-output-depth="16"]').click(); + const outputDepth16 = { + pointCount: points('[data-output-line="block"]').split(" ").length, + bars: root.querySelectorAll("[data-output-bars] i").length, + copy: text("[data-output-copy]"), + }; + + root.querySelector('[data-gradient-tab="verdict"]').click(); + const verdict = { + panel: panel(), + rows: root.querySelectorAll(".verdict-table tbody tr").length, + metrics: root.querySelectorAll(".metric-pairs article").length, + costs: root.querySelectorAll(".cost-compare article").length, + hashes: root.querySelectorAll(".hash-ledger code").length, + exact: root.textContent.includes("model + optimizer exact") && + root.textContent.includes("all frozen fields exact"), + mixed: root.textContent.includes("depth-dependent or inconclusive"), + }; + + const first = root.querySelector('[data-gradient-tab="definition"]'); + first.focus(); + first.dispatchEvent(new KeyboardEvent("keydown", { key: "ArrowRight", bubbles: true })); + const keyboard = { + selected: root.querySelector('[data-gradient-tab][aria-selected="true"]').dataset.gradientTab, + panel: panel(), + }; + + return { + initial, spectrumInitial, normalized, depth16, seedStep, + timelineInitial, timelineChanged, outputInitial, outputDepth16, + verdict, keyboard, + documentOverflow: document.documentElement.scrollWidth - document.documentElement.clientWidth, + rootOverflow: root.scrollWidth - root.clientWidth, + }; +})()`); +await pause(180); +await screenshot("/tmp/llm-atlas-k3-attnres-gradient-desktop.png"); + +await command("Emulation.setDeviceMetricsOverride", { + width: 390, + height: 844, + deviceScaleFactor: 1, + mobile: true, +}); +await navigate("/k3/"); +const mobile = await evaluate(`(() => { + const root = document.querySelector("[data-gradient-lab]"); + root.scrollIntoView({ block: "start", behavior: "instant" }); + window.scrollBy(0, -64); + root.querySelector('[data-gradient-tab="spectrum"]').click(); + root.querySelector('[data-spectrum-depth="16"]').click(); + root.querySelector('[data-gradient-tab="output"]').click(); + root.querySelector('[data-output-depth="16"]').click(); + return { + tabs: root.querySelectorAll("[data-gradient-tab]").length, + ledger: root.querySelectorAll(".gradient-ledger article").length, + outputBars: root.querySelectorAll("[data-output-bars] i").length, + visiblePanel: root.querySelector("[data-gradient-panel]:not([hidden])")?.dataset.gradientPanel, + documentOverflow: document.documentElement.scrollWidth - document.documentElement.clientWidth, + rootOverflow: root.scrollWidth - root.clientWidth, + }; +})()`); +await pause(180); +await screenshot("/tmp/llm-atlas-k3-attnres-gradient-mobile.png"); + +const report = { desktop, mobile, exceptions }; +console.log(JSON.stringify(report, null, 2)); + +const numeric = (value) => Number.parseFloat(value.replace("−", "-")); +const failures = []; +if (desktop.initial.panel !== "definition" || desktop.initial.tabs !== 5 || desktop.initial.panels !== 5 || desktop.initial.ledger !== 6) failures.push("五视图初始结构异常"); +if (!desktop.initial.definition) failures.push("定义与对象边界缺失"); +if (desktop.spectrumInitial.panel !== "spectrum" || Math.abs(numeric(desktop.spectrumInitial.baseCv) - 0.3786) > 1e-4 || Math.abs(numeric(desktop.spectrumInitial.blockCv) - 0.5996) > 1e-4) failures.push("depth-32 final spectrum 读数异常"); +if (desktop.spectrumInitial.boundaries !== 7 || desktop.spectrumInitial.baseLine === desktop.normalized.baseLine || !desktop.normalized.title.includes("NORMALIZED")) failures.push("绝对/归一化谱或组边界异常"); +if (desktop.depth16.pointCount !== 16 || desktop.depth16.boundaries !== 7 || Math.abs(numeric(desktop.depth16.baseCv) - 0.4246) > 1e-4 || Math.abs(numeric(desktop.depth16.blockCv) - 0.4678) > 1e-4) failures.push("depth-16 spectrum 切换异常"); +if (!desktop.seedStep.state.includes("2026073001") || !desktop.seedStep.state.includes("2,000") || Math.abs(numeric(desktop.seedStep.baseCv) - 0.4101) > 1e-4 || Math.abs(numeric(desktop.seedStep.blockCv) - 0.3656) > 1e-4) failures.push("seed/checkpoint spectrum 切换异常"); +if (desktop.timelineInitial.panel !== "timeline" || desktop.timelineInitial.pointCount !== 6 || desktop.timelineInitial.line === desktop.timelineChanged.line || !desktop.timelineChanged.title.includes("FIRST / LAST")) failures.push("六时点轨迹指标切换异常"); +if (desktop.outputInitial.panel !== "output" || desktop.outputInitial.pointCount !== 32 || desktop.outputInitial.bars !== 32 || desktop.outputInitial.boundaries !== 7 || desktop.outputDepth16.pointCount !== 16 || desktop.outputDepth16.bars !== 16 || !desktop.outputDepth16.copy.includes("DEPTH 16")) failures.push("Output RMS 深度/组节律切换异常"); +if (desktop.verdict.panel !== "verdict" || desktop.verdict.rows !== 2 || desktop.verdict.metrics !== 4 || desktop.verdict.costs !== 3 || desktop.verdict.hashes !== 4 || !desktop.verdict.exact || !desktop.verdict.mixed) failures.push("联合判定、成本或重放视图异常"); +if (desktop.keyboard.selected !== "spectrum" || desktop.keyboard.panel !== "spectrum") failures.push("键盘 tab 导航异常"); +if (desktop.documentOverflow > 1 || desktop.rootOverflow > 1 || mobile.documentOverflow > 1 || mobile.rootOverflow > 1) failures.push("桌面或移动端出现文档级横向溢出"); +if (mobile.tabs !== 5 || mobile.ledger !== 6 || mobile.outputBars !== 16 || mobile.visiblePanel !== "output") failures.push("移动端交互结构异常"); +if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`); + +if (failures.length) { + console.error(`\nFAIL\n- ${failures.join("\n- ")}`); + process.exitCode = 1; +} else { + console.log("\nPASS K3 AttnRes gradient browser regression"); +} + +socket.close(); diff --git a/scripts/check-k3-attnres-gradient-data.mjs b/scripts/check-k3-attnres-gradient-data.mjs new file mode 100644 index 0000000..b3395fe --- /dev/null +++ b/scripts/check-k3-attnres-gradient-data.mjs @@ -0,0 +1,68 @@ +import { createHash } from "node:crypto"; +import { readdirSync, readFileSync } from "node:fs"; + +const read = (path) => { + const bytes = readFileSync(new URL(path, import.meta.url)); + return { + bytes, + json: JSON.parse(bytes), + sha256: createHash("sha256").update(bytes).digest("hex"), + }; +}; + +const raw = read("../src/data/k3-attnres-gradient.json"); +const compact = read("../src/data/k3-attnres-gradient-compact.json"); +const reproduction = read("../experiments/k3/attnres_gradient/reproduction.json"); +const rawDirectory = new URL("../experiments/k3/attnres_gradient/results/raw/", import.meta.url); +const failures = []; +const expect = (condition, message) => { + if (!condition) failures.push(message); +}; + +expect(raw.sha256 === "ad461cbe74fc671f356288d37c9b628618814915116e4fbe03aaaa48367e6a8d", "raw aggregate SHA-256 changed"); +expect(compact.sha256 === "5377a5e731db2fdc85a0327f05d43f1cd067d3fc34633619d3004f3fc31968e3", "compact payload SHA-256 changed"); +expect(reproduction.sha256 === "aedcde6accc6eb1c24b122ffb55706ef620062e848d9fe4c5f622b084a4fdea6", "reproduction payload SHA-256 changed"); +expect(compact.json.protocol_id === "llm-atlas-k3-attnres-gradient-scale-v1", "protocol identity mismatch"); +expect(compact.json.study.formal_runs === 12, "formal grid is not 12 cells"); +expect(compact.json.study.formal_target_bytes === 786_432_000, "formal target-byte budget mismatch"); +expect(compact.json.study.replay_target_bytes === 65_536_000, "replay byte budget mismatch"); +expect(compact.json.cells.length === 12, "compact cell count mismatch"); +expect(Object.keys(raw.json.runs).length === 12, "raw formal cell count mismatch"); +expect(readdirSync(rawDirectory).filter((name) => name.endsWith(".json")).length === 21, "public raw output count is not 21"); +expect(reproduction.json.replay_exact.exact, "full formal replay is not exact"); +expect(Object.values(reproduction.json.smoke_exact).every((row) => row.exact && row.gradient_gate.passed), "paired smoke or loss-scale gate failed"); + +for (const depth of ["16", "32"]) { + const summary = compact.json.depth_summaries[depth]; + expect(summary.verdict.label === "mixed / inconclusive at this depth", `depth ${depth} verdict changed`); + expect(summary.by_seed.length === 3, `depth ${depth} seed count changed`); + expect(summary.by_seed.every((row) => row.block_minus_baseline_bpc < 0), `depth ${depth} BPC pairing changed`); + expect(summary.by_seed.every((row) => row.relative_cv_reduction < 0), `depth ${depth} CV counterevidence changed`); + expect(summary.by_seed.every((row) => row.relative_imbalance_reduction > 0), `depth ${depth} first/last improvement changed`); +} + +expect(Math.abs(compact.json.depth_summaries["16"].means.relative_cv_reduction - (-0.10264135379685868)) < 1e-15, "depth-16 CV contrast changed"); +expect(Math.abs(compact.json.depth_summaries["32"].means.relative_cv_reduction - (-0.600330100169428)) < 1e-15, "depth-32 CV contrast changed"); +expect(Math.abs(compact.json.depth_summaries["16"].means.relative_imbalance_reduction - 0.6099652484299795) < 1e-15, "depth-16 imbalance contrast changed"); +expect(Math.abs(compact.json.depth_summaries["32"].means.relative_imbalance_reduction - 0.7200517407719272) < 1e-15, "depth-32 imbalance contrast changed"); +expect(compact.json.overall_verdict === "depth-dependent or inconclusive", "overall preregistered verdict changed"); + +if (failures.length) { + console.error(`FAIL K3 AttnRes gradient data\n- ${failures.join("\n- ")}`); + process.exit(1); +} + +console.log(JSON.stringify({ + protocol: compact.json.protocol_id, + formalRuns: compact.json.study.formal_runs, + formalTargetBytes: compact.json.study.formal_target_bytes, + depth16: compact.json.depth_summaries["16"].verdict.label, + depth32: compact.json.depth_summaries["32"].verdict.label, + replayExact: reproduction.json.replay_exact.exact, + hashes: { + raw: raw.sha256, + compact: compact.sha256, + reproduction: reproduction.sha256, + }, +}, null, 2)); +console.log("PASS K3 AttnRes gradient frozen data"); diff --git a/scripts/check-k3-browser.mjs b/scripts/check-k3-browser.mjs index 380e64c..4d76a9b 100644 --- a/scripts/check-k3-browser.mjs +++ b/scripts/check-k3-browser.mjs @@ -81,6 +81,8 @@ const overview = await evaluate(`(() => ({ artifactMismatch: document.querySelector("#artifacts")?.textContent.includes("A_log [128] ≠ expected [96]"), attnresTabs: document.querySelectorAll("[data-attnres-tab]").length, attnresPanels: document.querySelectorAll("[data-attnres-panel]").length, + gradientTabs: document.querySelectorAll("[data-gradient-tab]").length, + gradientPanels: document.querySelectorAll("[data-gradient-panel]").length, nativeVisionCorrected: document.body.textContent.includes("MoonViT‑V2 从头训练") && document.body.textContent.includes("同一个 next-token prediction objective"), staleVisionClaim: document.body.textContent.includes("先固定语言模型训练视觉组件"), @@ -290,8 +292,9 @@ const mobile = await evaluate(`(() => { artifactTabs: document.querySelectorAll("[data-artifact-tab]").length, artifactLayers: document.querySelectorAll("[data-layer-cell]").length, attnresTabs: document.querySelectorAll("[data-attnres-tab]").length, + gradientTabs: document.querySelectorAll("[data-gradient-tab]").length, offenders: [...document.querySelectorAll("body *")] - .filter((node) => !node.closest(".paper-chain, .spec-table-wrap, .cache-strip, .architecture-explorer, [data-k3-lab], [data-k3-artifact-lab], [data-attnres-lab]")) + .filter((node) => !node.closest(".paper-chain, .spec-table-wrap, .cache-strip, .architecture-explorer, [data-k3-lab], [data-k3-artifact-lab], [data-attnres-lab], [data-gradient-lab]")) .filter((node) => node.getBoundingClientRect().right > document.documentElement.clientWidth + 1) .slice(0, 15) .map((node) => ({ @@ -322,12 +325,13 @@ console.log(JSON.stringify(report, null, 2)); const numeric = (text) => Number.parseFloat(text.replaceAll(",", "").replace("−", "-")); const failures = []; if (!overview.title.includes("因果环节")) failures.push("K3 二轮标题异常"); -if (overview.sections !== 33 || overview.tocLinks !== 33) failures.push("32 个编号专题加阅读链的目录结构异常"); +if (overview.sections !== 34 || overview.tocLinks !== 34) failures.push("33 个编号专题加阅读链的目录结构异常"); if (overview.ledgers !== 32 || overview.reportMap !== 9) failures.push("32 张问题账或报告地图异常"); if (overview.figureAtlas !== 21 || overview.paperLinks !== 100 || overview.paperGroups < 12) failures.push("图表审计或 100 节点阅读链异常"); if (overview.labTabs !== 8 || overview.labPanels !== 8) failures.push("八联实验结构异常"); if (overview.artifactTabs !== 4 || overview.artifactPanels !== 4 || overview.artifactLayers !== 93 || !overview.artifactMismatch) failures.push("开放工件四视图、93 层条带或形状冲突异常"); if (overview.attnresTabs !== 5 || overview.attnresPanels !== 5) failures.push("AttnRes 独立实验五视图异常"); +if (overview.gradientTabs !== 5 || overview.gradientPanels !== 5) failures.push("AttnRes 梯度定义扩展五视图异常"); if (!overview.nativeVisionCorrected || overview.staleVisionClaim) failures.push("原生多模态纠错未生效或旧错误残留"); if (overview.documentOverflow > 1 || mobile.documentOverflow > 1) failures.push("桌面或移动端存在文档级横向溢出"); if (labs.memoryInitial.panel !== "memory" || numeric(labs.memoryInitial.additiveError) <= numeric(labs.memoryInitial.deltaError)) failures.push("Delta memory 初始递推异常"); @@ -352,7 +356,7 @@ if (artifacts.parameterChanged.shape !== "[96,128] F32" || !artifacts.parameterC if (artifacts.reproductionInitial.panel !== "reproduction" || numeric(artifacts.reproductionInitial.speedup) !== 1.85 || numeric(artifacts.reproductionInitial.localMean) < 2.6 || !artifacts.reproductionInitial.exactSuite || numeric(artifacts.reproductionInitial.cv) < 2) failures.push("FlashKDA H20、本机 exact suite 或 router 初始探针异常"); if (numeric(artifacts.reproductionChanged.speedup) !== 3.27 || numeric(artifacts.reproductionChanged.flash) !== 0.7064 || numeric(artifacts.reproductionChanged.localMean) >= numeric(artifacts.reproductionInitial.localMean) || !artifacts.reproductionChanged.localMode.includes("FP32 state") || numeric(artifacts.reproductionChanged.cv) <= numeric(artifacts.reproductionInitial.cv) || numeric(artifacts.reproductionChanged.zero) <= numeric(artifacts.reproductionInitial.zero)) failures.push("GB200 benchmark、本机 varlen/state 或 synthetic router counterexample 未更新"); if (artifacts.keyboardSelected !== "tensors" || artifacts.keyboardVisible !== "tensors") failures.push("开放工件键盘 tab 导航异常"); -if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 8 || mobile.artifactTabs !== 4 || mobile.artifactLayers !== 93 || mobile.attnresTabs !== 5) failures.push("移动端导航或实验异常"); +if (!mobile.menuVisible || mobile.menuOpen !== "true" || mobile.tabs !== 8 || mobile.artifactTabs !== 4 || mobile.artifactLayers !== 93 || mobile.attnresTabs !== 5 || mobile.gradientTabs !== 5) failures.push("移动端导航或实验异常"); if (mobile.offenders.length) failures.push(`移动端越界元素:${JSON.stringify(mobile.offenders)}`); if (exceptions.length) failures.push(`浏览器异常:${exceptions.join(" | ")}`); diff --git a/src/components/K3AttnResGradientLab.astro b/src/components/K3AttnResGradientLab.astro new file mode 100644 index 0000000..98eaf5d --- /dev/null +++ b/src/components/K3AttnResGradientLab.astro @@ -0,0 +1,813 @@ +--- +import rawLab from "@/data/k3-attnres-gradient-compact.json"; + +const lab = rawLab as any; +const json = JSON.stringify(lab).replaceAll("<", "\\u003c"); +const depth16 = lab.depth_summaries["16"]; +const depth32 = lab.depth_summaries["32"]; +const formatSigned = (value: number, digits = 4) => + `${value < 0 ? "−" : value > 0 ? "+" : ""}${Math.abs(value).toFixed(digits)}`; +const pct = (value: number, digits = 1) => `${(value * 100).toFixed(digits)}%`; +const shortHash = (value: string) => `${value.slice(0, 10)}…${value.slice(-8)}`; +--- + +
+
+
+

ROUND 05 / GRADIENT DEFINITION × DEPTH SCALE

+

“早层不再过大”与“整条谱更均匀”不是同一件事

+
+

+ 16 / 32 Transformer blocks · Baseline / Block · 3 seeds · 8,000 steps。 + 被测对象是 post-MLP output activation gradient,不冒充论文未公开的 Figure 5 实现。 +

+
+ +
+
FORMAL GRID2 × 2 × 3

12 个独立训练格

+
TARGET BYTES786.432M

每格 65,536,000

+
DEPTH16 → 32

宽度固定 192

+
CV6 / 6 更差

中后段局部尖峰

+
FIRST ↔ LAST6 / 6 更近

失衡改善 56%–81%

+
FULL REPLAYexact

model + optimizer state

+
+ +
+ + + + + +
+ +
+
+
I / DEFINITION AUDIT

论文画出一条“梯度曲线”,却没有给出足以唯一重算的测量合同

+

+ Figure 5(c) 只写 “Each transformer block’s gradient magnitude”。 + 官方仓库没有训练代码、checkpoint、统计脚本或原始数组。 +

+
+ +
+
+ OFFICIAL / KNOWN +
图与文字能确认
+
    +
  • 横轴是 Transformer block index。
  • +
  • Figure 5(b) 同时画 block output magnitude。
  • +
  • 正文说 Baseline 早层梯度过大,Block 更均匀。
  • +
  • 最终模型约 27 blocks / 54 residual layers。
  • +
+
+
+ OFFICIAL / UNDERDEFINED +
无法从公开工件唯一恢复
+
    +
  • activation、branch 还是 parameter gradient?
  • +
  • L2、RMS、mean absolute 还是别的 norm?
  • +
  • batch / token / channel 怎样 reduction?
  • +
  • 哪个 checkpoint、AMP / clip 前还是后?
  • +
+
+
+ +
+
FIXED INPUT16 × 256 bytes

同一 diagnostic tensor

+ → +
POST-MLP OUTPUTh₁ … hL

每个 Transformer block 一个

+ → +
TOKEN-MEAN CEℒ

FP32 cross entropy

+ → +
MEASURERMS(∂ℒ/∂hl)

B × T × C 联合 RMS

+
+ +
+
ROUND 04∇θlℒ

核心参数梯度:权重收到多大更新信号。

+ ≠ +
ROUND 05∂ℒ/∂hl

activation gradient:损失对这一深度表示有多敏感。

+

两种对象都公开;新指标不会覆盖上一轮反结果。

+
+ +
+ 能说

“这是与 Figure 5 叙述对齐的一种公开 operationalization。”

+ 不能说

“论文作者就是这样算的,或本站复画了 Figure 5(c)。”

+
+
+ + + + + + + + + + +
+ + + + diff --git a/src/pages/k3/index.astro b/src/pages/k3/index.astro index 4544025..75165b6 100644 --- a/src/pages/k3/index.astro +++ b/src/pages/k3/index.astro @@ -2,6 +2,7 @@ import BaseLayout from "@/layouts/BaseLayout.astro"; import ArchitectureExplorer from "@/components/ArchitectureExplorer.astro"; import K3ArtifactLab from "@/components/K3ArtifactLab.astro"; +import K3AttnResGradientLab from "@/components/K3AttnResGradientLab.astro"; import K3AttnResTraceLab from "@/components/K3AttnResTraceLab.astro"; import K3ReportLab from "@/components/K3ReportLab.astro"; import { k3FigureAtlas, k3Ledgers, k3PaperChain, k3ReportMap } from "@/data/k3"; @@ -38,7 +39,8 @@ const toc = [ ["28", "lab", "八联交互实验"], ["29", "artifacts", "开放权重工件审计"], ["30", "attnres-reduced", "AttnRes 缩小机制实验"], - ["31", "audit", "21 张图表审计"], + ["31", "attnres-gradient", "梯度定义与深度扩展"], + ["32", "audit", "21 张图表审计"], ["↳", "papers", "100 节点阅读链"], ]; @@ -107,13 +109,13 @@ const paperGroups = [
-

ANCHOR REPORT / ROUND 04 KIMI K3 · REPORT → ARTIFACTS → INDEPENDENT PROBE

+

ANCHOR REPORT / ROUND 05 KIMI K3 · REPORT → ARTIFACTS → INDEPENDENT PROBE

不把报告压成摘要
把每个因果环节
重新展开

K3 同时扩展序列、深度、宽度、视觉与 Agent 轨迹。真正值得读的不是 2.8T 这个最大数字, @@ -123,11 +125,11 @@ const paperGroups = [

QUESTIONS
32 张问题账
REPORT
16 Figures · 5 Tables
-
LABS
8 + 4 + 5 个交互视图
+
LABS
8 + 4 + 5 + 5 个交互视图
READING
100 个一手 / 官方节点
MODEL
2.78T total / 104.2B active
ARTIFACTS
96 shards · 497,220 tensors
-
STATUS
K3 四轮 · AttnRes 实验
+
STATUS
K3 五轮 · 梯度定义闭环
@@ -892,12 +894,36 @@ const paperGroups = [ 阅读完整研究审计 复跑公开实验代码 Attention Residuals 原论文 - 官方实现 + 官方论文工件 + + + +
+

31 GRADIENT DEFINITION × DEPTH SCALE

+

“论文说梯度更均匀”,和上一轮参数梯度反结果,测的是同一件事吗?

+

+ 第五轮先审计 Attention Residuals 官方论文与仓库:Figure 5(c) 没有公开 gradient tensor、 + norm、reduction、diagnostic batch、AMP / clipping 时点或统计代码。本站因此冻结一个可复现的 + post-MLP output activation-gradient 定义,把深度扩到 16 / 32 blocks、预算扩到 8,000 steps, + 再用三 seed 检查“首尾平衡”和“全层离散度”是否真的同方向。 +

+
+
F / FROZEN12 × 8,000 steps

786,432,000 formal target bytes;两深度、两结构、三 seed。

+
X / OBSERVEDfirst/last 6 / 6 改善

depth-16 平均 61.0%;depth-32 平均 72.0%。

+
X / COUNTEREVIDENCECV 6 / 6 恶化

局部尖峰让 depth-16 / 32 平均相对恶化 10.3% / 60.0%。

+
R / REPLAYmodel + optimizer exact

指定 32-layer Block 格从零重训 8,000 steps,冻结字段逐项一致。

+
+ +
-

31 FIGURE & TABLE AUDIT

+

32 FIGURE & TABLE AUDIT

Figure 1–16、Table 1–5:每张图究竟支持什么,不能支持什么

{k3FigureAtlas.map(([id, report, title, contract]) => ( diff --git a/src/pages/progress/index.astro b/src/pages/progress/index.astro index e940a42..3b6b743 100644 --- a/src/pages/progress/index.astro +++ b/src/pages/progress/index.astro @@ -9,7 +9,7 @@ const researching = chapters.filter((chapter) => ["researching", "drafting"].inc const workstreams = [ { label: "研究框架与规范", value: 83, next: "给 Scaling 与推理专题补逐篇图表/实验精读层级" }, { label: "网站设计系统", value: 89, next: "打印样式与更多通用可视化组件" }, - { label: "Kimi K3 深读", value: 96, next: "对齐 AttnRes 梯度定义并扩展深度/预算;等待 A_log 官方转换合同" }, + { label: "Kimi K3 深读", value: 98, next: "对齐 layer 21–25 梯度尖峰与 mixer weights;等待 A_log 官方转换合同" }, { label: "语言模型前史", value: 78, next: "逐图精读 Kneser–Ney、LSTM 与 Bahdanau,并加入真实小语料复现" }, { label: "Transformer 基础", value: 79, next: "逐图精读多头电路、Pre/Post-LN 与真实 kernel / KV 配置" }, { label: "表示、位置与残差高速公路", value: 81, next: "加入真实 hidden-state / norm traces、长上下文位置外推复现与更多深层稳定性消融" }, @@ -50,7 +50,7 @@ const workstreams = [
OVERALL
专题平均 {average}%
READABLE
{published} 个首版可读专题
ACTIVE
{researching} 个研究/写作中
-
UPDATED
2026-07-30 07:30 CST
+
UPDATED
2026-07-30 10:05 CST
MODE
持续迭代,不锁死版本
@@ -97,7 +97,7 @@ const workstreams = [
✓

K3 报告已结构化拆解

47 页报告目录、151 条参考来源和架构/后训练/系统主线已经提取。

✓

17 专题知识图

从语言模型基础到评测安全,包含先修依赖和三条贯穿案例。

✓

编辑式网站系统

响应式导航、章节模板、侧栏、进度、论文链和证据提示组件。

-
✓

九十四个原创交互视图

K3 三轴图、八联报告实验、四联开放工件实验与五联 AttnRes 独立实验,DeepSeek 四联公式实验、十三联 Base 工件实验、Chat 行为、completion/full-depth、multi-seed、cross-source 与 task-bootstrap CRN 五轮实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。

+
✓

九十九个原创交互视图

K3 三轴图、八联报告实验、四联开放工件实验与两轮十联 AttnRes 独立实验,DeepSeek 四联公式实验、十三联 Base 工件实验、Chat 行为、completion/full-depth、multi-seed、cross-source 与 task-bootstrap CRN 五轮实验,以及语言模型前史、Transformer、表示深度、长上下文、MoE、推理、Agent、多模态、训练系统、推理服务、Scaling、数据工程、数值、Alignment 与评测安全专题。

✓

十七篇首版长文

K3、语言模型前史、Transformer、表示/位置/残差、DeepSeek、Scaling、数据工程、长上下文、MoE、后训练、推理、Agent、原生多模态、训练系统、推理服务、数值优化与评测安全专题。

✓

语言模型前史深度专题

八张独立问题账、33 个正式节点、20 段长文与概率—向量—记忆—对齐四联实验。

✓

Transformer 深度专题

十张独立问题账、40 个正式节点、21 段正文与 QKV—Mask—多头位置—Block 成本四联实验。

@@ -106,6 +106,7 @@ const workstreams = [
✓

Kimi K3 技术报告二轮深读

三十二张问题账、Figure 1–16 / Table 1–5 审计、100 节点阅读链,以及 Delta—Decay—AttnRes—LatentMoE—SiTU—QB—MOPD—Cache 八联实验。

✓

Kimi K3 三轮开放工件里程碑

固定官方 revisions,审计 96 个 shards、497,220 个 tensor entries 与真实 KDA / MLA / MoE / MoonViT shapes;四联实验分开显示层型、tensor anatomy、参数范围和复现边界。

✓

Kimi K3 四轮 AttnRes 独立实验

冻结三结构 × 三 seed 的 9 个 2,000-step 格;Full / Block 相对 Baseline 的平均 paired delta 为 −0.01457 / −0.04247 BPC,但核心参数梯度 CV 没有复现论文叙述。指定正式格全新进程八字段 exact,五视图同时展示结果、反证、成本与 claim boundary。

+
✓

Kimi K3 五轮梯度定义与深度扩展

先确认 Figure 5 没有公开唯一 gradient telemetry 合同,再冻结 16/32 blocks × Baseline/Block × 3 seeds 的 12 个 8,000-step 格。Block 的首尾失衡 6/6 改善但全层 CV 6/6 恶化,两个深度都判为 mixed;指定 32 层格完整重训的模型、优化器与全部冻结字段 exact。

✓

FlashKDA RTX 5090 执行闸门

隔离 CUDA 13.0 / glibc 2.39 编译 sm_120a wheel;6/6 官方参考逐元素相等,并完成 fixed / varlen、三种 state mode 的 1,800 个 CUDA Event samples。

✓

Scaling Laws 深度专题

九张账、29 个一手节点、DeepSeek/Kimi 双谱系与曲面—部署—复用—涌现四联实验。

✓

数据工程深度专题

十二张账、31 个一手节点、DeepSeek/Kimi 双谱系与流水线—去重—混合—改写四联实验。

@@ -134,7 +135,7 @@ const workstreams = [
优先级专题本轮交付完成闸门
-
P0K3 四轮后续

对齐论文梯度定义 → 增加 depth / budget → 等待 A_log 官方合同后进入真实 checkpoint forward

尺度复查 + 工件边界
+
P0K3 五轮后续

对齐 layer 21–25 尖峰、pre-attention / pre-MLP 与 mixer source weights → 等待 A_log 官方合同后进入真实 checkpoint forward

局部机制 + 工件边界
P0DeepSeek 八轮后续

干预式 mediation → SM90 FlashMLA / FP8 / pipeline traces → R1-like RL 小模型复现

运行证据 + 独立复现
P0Transformer 二轮

多头电路逐图 → Pre/Post-LN 真实 traces → Flash/KV 配置与 kernel 对照

逐图笔记 + 实测边界
P0表示、位置与残差二轮

真实 hidden-state / norm traces → 长上下文位置外推 → mHC / AttnRes 深层稳定性消融

可复现实验 + 逐图笔记
@@ -220,6 +221,10 @@ const workstreams = [
AttnRes 缩小实验先冻结、后运行

三结构共享公共主干、初始化、窗口与优化器;只按三个 paired seed 和预注册 −0.010 BPC 阈值给出本协议内方向判断。

支持结果与梯度反结果同时进入主视区

Full / Block 的最终 BPC 同向改善;核心参数 gradient RMS CV 却高于 Baseline,不换指标掩盖。

正式重放不把 wall time 纳入 exact

Block / seed-1 的模型、优化器、曲线、历史、诊断和环境八字段 exact;计时受调度影响,单独报告。

+
Figure 5 的“梯度”不再靠猜测补合同

官方未公开 gradient tensor、norm、reduction 与统计代码;本站 activation-gradient 定义只叫 operationalization,不叫论文复画。

+
首尾平衡与全层 CV 永久分账

Block 在 6/6 配对中改善 first/last,却因中后段局部尖峰让 CV 在 6/6 配对中恶化;联合判定保持 mixed。

+
绝对梯度尺度必须与归一化谱同屏

Block mean gradient 约为 Baseline 的 54%–57%;更接近 1 的首尾比不能偷换成各层信号更强。

+
32 层完整重放扩到状态哈希

8,000-step fresh replay 的全部冻结字段以及 model / optimizer state hashes exact;额外 replay bytes 单列,不混入 formal 预算。

32-token 对照改为同源 16→24

TNEWS 只有 105/10,000 条达到 32 tokens,强行统一会落入约 1% 极端长尾;24-token eligibility 仍保留 1,609 条中文候选。

长度敏感性必须成对重采样

16-token 输入严格是 24-token 输入前缀,2,000 次 bootstrap 共用 prompt indices;结果只描述固定 cohort 的长度敏感性。

三类 cohort 永久分身份

自然长度回答本批样本如何路由;matched-16 / 24 回答同一 prompt 多看 8 tokens 后如何变化,不把二者混成内容因果。