Refresh papers and define Agent evaluation
This commit is contained in:
@@ -10,6 +10,12 @@ status: seed
|
||||
| P0 | SWE-EVO | coding agent benchmark | 长程软件演进任务比 SWE-Bench 更接近真实工程 | skimmed | https://arxiv.org/html/2512.18470v5 |
|
||||
| P1 | Dialogue-SWEBench | human-in-the-loop coding agent | 评估 coding agent 澄清需求能力,补足全自动 benchmark 盲区 | skimmed | https://arxiv.org/html/2606.13995v1 |
|
||||
| P1 | Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents | safety benchmark taxonomy | 用于选择和比较 Agent safety benchmark | queued | https://arxiv.org/html/2605.16282v1 |
|
||||
| P0 | The Regression Tax | component evaluation | 需要验证 gains/regressions 分解、三类 skill regression 和跨 harness 设置 | queued | https://arxiv.org/abs/2607.22520 |
|
||||
| P0 | Do Agent Benchmarks Measure Capability? | evaluation validity | 直接检查 benchmark protocol 是否真的测到目标能力 | queued | https://arxiv.org/abs/2607.22368 |
|
||||
| P0 | AgentLens | coding-agent trajectory | 生产轨迹评审和失败类型可能直接支持 K1412 eval task/failure taxonomy | queued | https://arxiv.org/abs/2607.06624 |
|
||||
| P0 | DynamicMCPBench | dynamic tool evaluation | 动态 MCP server、trace grounding 和 effect scoring 对工具 Agent 评测直接相关 | queued | https://arxiv.org/abs/2607.20531 |
|
||||
| P1 | AgentDebugX | observability / recovery | 评估失败可观测、归因和恢复工具的实现与证据 | queued | https://arxiv.org/abs/2607.18754 |
|
||||
| P1 | Reason Less, Verify More | deterministic verification | 核验确定性 gate 是否真的恢复策略违规失败 | queued | https://arxiv.org/abs/2607.07405 |
|
||||
|
||||
## Priority Rules
|
||||
|
||||
|
||||
Reference in New Issue
Block a user