Refresh papers and define Agent evaluation

This commit is contained in:
wuyang
2026-07-27 16:28:23 +08:00
parent 8e4ff3779b
commit df475e8d90
180 changed files with 31313 additions and 160 deletions
+13 -7
View File
@@ -6,13 +6,16 @@ status: seed
## Current Snapshot
- date: 2026-07-08
- scope: arXiv Agent-related papers from 2025-07-08 to 2026-07-08
- candidates seen: 1506
- high-relevance candidates selected into paper items: 970
- reserve candidates retained in manifest: 536
- total paper items: 974
- detailed summary: [Paper Corpus Summary](corpus-summary-2026-07-08.md)
- date: 2026-07-27
- baseline scope: arXiv Agent-related papers from 2025-07-08 to 2026-07-08
- incremental scope: 2026-07-09 to 2026-07-27
- incremental candidates seen: 249
- incremental high-relevance candidates: 157
- incremental new paper items: 151
- incremental reserve candidates: 92
- total paper items after refresh: 1125
- baseline summary: [2026-07-08 corpus](corpus-summary-2026-07-08.md)
- incremental summary: [2026-07-27 corpus increment](corpus-summary-2026-07-27.md)
- caveat: high-recall automated sweep; most items are `queued` and need human skim.
## Strong Signals
@@ -23,6 +26,9 @@ status: seed
- JD 高频词和论文趋势高度重合:memory、eval、planning/tool use、coding agent。
- 最近一年 Agent 论文的重心明显落在 evaluation、tool use、memory、safety、coding agent、computer-use/GUI agent。
- 2026-05 到 2026-07 的论文密度明显上升,说明 Agent research 正在快速扩张。
- 最新增量继续把评测推向 protocol validity、动态工具环境、trajectory 诊断、确定性
verification、成本和部署条件。
- 组件评测需要同时报告 gains 与 regressions;平均提升可能掩盖原本成功任务被新组件破坏。
## Weak or Unproven Claims