Files
llm-atlas/research/ALIGNMENT_GROK_LEADS.md
T
2026-07-29 06:07:42 +08:00

5.6 KiB
Raw Blame History

AlignmentGrok 候选论文线索

状态:未核验候选账
生成方式:本机 Grok CLI Headless 模式,2026-07-29
使用约束:这里只记录检索线索,不能直接作为正文事实、公式、数字或历史结论的依据。正式内容必须回到论文 PDF、会议页面或研究机构原文逐项核验。

这次线索检索的用途

正式研究账本已经覆盖 SFT、偏好数据、奖励模型、PPO/RLHF、RLAIF、DPO、RLVR、DeepSeek 与 Kimi/K3。Grok 的任务不是替正文写作,而是做一次“漏项扫描”:寻找可能缺失的论文分支、失败模式和边界案例。

它给出 45 个候选条目,按下列九组组织:

候选组 条目数 对正式研究最有价值的提醒
SFT 6 FLAN、T0、Natural Instructions、SuperNI、Self-Instruct、LIMA 需要区分任务泛化、数据生成与少量高质量示范
偏好 / 奖励模型 5 不只看标注规模,也要检查偏好分歧、长度相关性和奖励泛化
PPO / RLHF 6 Christiano → Ziegler → Stiennon → InstructGPT 的主线应与 PPO 算法本体分开
RLAIF 4 Constitutional AI、自我批评修订与 AI preference 不是同一层概念
DPO 分支 8 RRHF、SLiC-HF、IPO、KTO、ORPO、SimPO 等不能笼统称作“DPO 小改版”
失败模式 5 reward hacking、overoptimization、sycophancy、verbosity bias 应拆成不同故障
RLVR 4 PRM 是前驱但不等于程序化 verifierDeepSeekMath 的 GRPO 与具体奖励来源要分开
DeepSeek 4 LLM、Math、V3、R1 的角色不同,不能把 R1-Zero 的纯 RL 结论外推到完整 R1
Kimi 3 k1.5、K2、K2.5 的推理、agentic、多模态主线不同

候选清单

以下名称保留为后续查漏入口;存在重复跨组复用,不代表 45 篇互不重复的唯一论文。

SFT 与指令数据

  • Finetuned Language Models Are Zero-Shot LearnersFLAN
  • Multitask Prompted Training Enables Zero-Shot Task GeneralizationT0
  • Natural Instructions
  • Super-NaturalInstructions
  • Scaling Instruction-Finetuned Language Models
  • Self-Instruct
  • LIMA: Less Is More for Alignment

偏好、奖励模型与 RLHF

  • Deep Reinforcement Learning from Human Preferences
  • Fine-Tuning Language Models from Human Preferences
  • Learning to Summarize from Human Feedback
  • Training Language Models to Follow Instructions with Human Feedback
  • Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
  • Proximal Policy Optimization Algorithms
  • Scaling Laws for Reward Model Overoptimization
  • AlpacaFarm
  • OpenAssistant Conversations
  • Llama 2: Open Foundation and Fine-Tuned Chat Models
  • Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMs

RLAIF 与宪法式约束

  • Constitutional AI: Harmlessness from AI Feedback
  • RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
  • UltraFeedback
  • Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

直接偏好优化分支

  • Direct Preference Optimization: Your Language Model is Secretly a Reward Model
  • RRHF: Rank Responses to Align Language Models with Human Feedback
  • SLiC-HF: Sequence Likelihood Calibration with Human Feedback
  • A General Theoretical Paradigm to Understand Learning from Human PreferencesIPO
  • KTO: Model Alignment as Prospect Theoretic Optimization
  • ORPO: Monolithic Preference Optimization without Reference Model
  • SimPO: Simple Preference Optimization with a Reference-Free Reward
  • Zephyr: Direct Distillation of LM Alignment
  • Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

失败模式与评测

  • Sycophancy / 用户信念迎合相关研究
  • A Long Way to Go: Investigating Length Correlations in RLHF
  • RewardBench
  • reward overoptimization / Goodhart 效应相关研究
  • 偏好标注分歧、both-bad pair 与 verbosity bias 相关研究

RLVR、DeepSeek 与 Kimi

  • Lets Verify Step by Step
  • Tülu 3: Pushing Frontiers in Open Language Model Post-Training
  • DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
  • Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
  • DeepSeek LLM
  • DeepSeek-V3 Technical Report
  • DeepSeek-R1
  • Kimi k1.5
  • Kimi K2
  • Kimi K2.5

必须保留的边界提醒

  1. “偏好更长”不能直接写成人类真正想要冗长答案;奖励模型可能只学到了长度代理特征。
  2. sycophancy 不完全等于撒谎,它可能来自 helpfulness、agreeableness 与偏好标签的纠缠。
  3. 人类标注的步骤正确性不等于程序可验证奖励,PRM 只能作为 RLVR 前驱来讲。
  4. DeepSeekMath 同时讨论模型奖励与规则奖励;“用了 GRPO”不能反推出“奖励一定可验证”。
  5. R1-Zero 的少监督 / 纯 RL 路径不能外推为完整 DeepSeek-R1 不需要冷启动与监督数据。
  6. Kimi K2、K2.5 的 agentic / multimodal 叙事不能替代 k1.5 的推理 RL 算法细节。
  7. “DPO 不需要奖励模型”仅指不单独训练显式 RM 并运行 PPO;偏好、参考策略与数据分布仍然存在。

当前裁决

  • 已进入正式账本:有一手 PDF 或官方页面支撑、且边界已经逐项复核的节点。
  • 暂留候选:Crossing the Reward Bridge、具体 sycophancy 论文、长度相关论文等,需要在对应专题扩写时重新核对标题、版本、实验设定和数字。
  • 明确不采用:任何仅凭二手叙事得出的“某算法彻底取代另一算法”“纯 RL 不再需要数据”“AI feedback 等于没有人类价值输入”等强结论。