feat: publish alignment deep dive
This commit is contained in:
@@ -21,3 +21,13 @@
|
||||
|
||||
其余已核节点通过 DOI、ACL Anthology、PMLR、ISCA、JMLR、NeurIPS 或官方 arXiv 页面定位,
|
||||
完整证据边界见 `../LANGUAGE_MODEL_HISTORY_RESEARCH.md`。
|
||||
|
||||
Alignment 首轮缓存位于 `alignment/`(不提交 PDF/TXT):
|
||||
|
||||
- Christiano 2017、Ziegler 2019、Stiennon 2020 与 InstructGPT;
|
||||
- Natural Instructions、FLAN、T0、HH-RLHF、Constitutional AI 与 Self-Instruct;
|
||||
- DPO、RRHF、SLiC-HF、LIMA、RLAIF、UltraFeedback、IPO、KTO、ORPO、RewardBench 与 SimPO。
|
||||
|
||||
DeepSeek LLM/V2/V3/R1/V3.2/V4、Kimi K2/K2.5/K3 与 MOPD 复用既有本地缓存。
|
||||
正式机制、公式、版本边界与 44 节点论文链见 `../ALIGNMENT_RESEARCH.md`;Grok 候选线索单独保存在
|
||||
`../ALIGNMENT_GROK_LEADS.md`,不能直接作为正文证据。
|
||||
|
||||
Reference in New Issue
Block a user