Refresh papers and define Agent evaluation
This commit is contained in:
@@ -0,0 +1,57 @@
|
||||
# Paper Corpus Increment: 2026-07-09 to 2026-07-27
|
||||
|
||||
status: generated-summary
|
||||
|
||||
## Scope
|
||||
|
||||
- source: arXiv official search
|
||||
- window: 2026-07-09 to 2026-07-27
|
||||
- unique candidates: 249
|
||||
- high-relevance candidates: 157
|
||||
- newly written paper items: 151
|
||||
- selected items already in the repository: 6
|
||||
- reserve candidates: 92
|
||||
|
||||
## Data
|
||||
|
||||
- candidates: `data/arxiv-agent-candidates-2026-07-09-to-2026-07-27.json`
|
||||
- selected: `data/arxiv-agent-papers-2026-07-09-to-2026-07-27.json`
|
||||
- run record: `collection-runs/2026-07-27-arxiv-incremental-refresh.md`
|
||||
|
||||
## What Changed
|
||||
|
||||
Evaluation remains the strongest high-recall signal: 137 of 157 selected candidates received the broad
|
||||
`agent-evaluation` tag, and 119 received `tool-use`. These tags overlap and are not quality judgments.
|
||||
|
||||
The useful new evaluation clusters are:
|
||||
|
||||
1. **Protocol validity and dynamic environments**: benchmarks increasingly vary tools, MCP servers,
|
||||
interaction protocols and real deployment conditions.
|
||||
2. **Trajectory diagnosis**: new work targets failure observability, attribution, runtime intervention
|
||||
and recovery rather than only final scores.
|
||||
3. **Independent verification**: several candidates emphasize deterministic gates, effect-scored traces,
|
||||
provenance and reward-hacking-resistant validators.
|
||||
4. **Regression and cost accounting**: average improvement is insufficient when a component fixes some
|
||||
tasks while breaking previously successful ones.
|
||||
5. **Interactive coding work**: coding-agent evaluation continues moving toward project building,
|
||||
production trajectories, payment/backend integration and repository-level optimization.
|
||||
|
||||
## Priority Candidates
|
||||
|
||||
- [The Regression Tax](https://arxiv.org/abs/2607.22520)
|
||||
- [Do Agent Benchmarks Measure Capability?](https://arxiv.org/abs/2607.22368)
|
||||
- [AgentLens](https://arxiv.org/abs/2607.06624)
|
||||
- [DynamicMCPBench](https://arxiv.org/abs/2607.20531)
|
||||
- [AgentDebugX](https://arxiv.org/abs/2607.18754)
|
||||
- [Reason Less, Verify More](https://arxiv.org/abs/2607.07405)
|
||||
- [Baselines Before Architecture](https://arxiv.org/abs/2607.13085)
|
||||
|
||||
All remain `queued`; the descriptions above are based on title/abstract triage.
|
||||
|
||||
## Caveats
|
||||
|
||||
- The official export API was unavailable during this run, so the official arXiv search interface was
|
||||
used through the new collector backend.
|
||||
- Search queries are overlapping and broad.
|
||||
- Automated relevance scoring can admit domain-specific uses of Agents with limited architectural value.
|
||||
- No new paper from this increment has been promoted to `skimmed` or evidence-backed status.
|
||||
Reference in New Issue
Block a user