58 lines
2.5 KiB
Markdown
58 lines
2.5 KiB
Markdown
# Paper Corpus Increment: 2026-07-09 to 2026-07-27
|
|
|
|
status: generated-summary
|
|
|
|
## Scope
|
|
|
|
- source: arXiv official search
|
|
- window: 2026-07-09 to 2026-07-27
|
|
- unique candidates: 249
|
|
- high-relevance candidates: 157
|
|
- newly written paper items: 151
|
|
- selected items already in the repository: 6
|
|
- reserve candidates: 92
|
|
|
|
## Data
|
|
|
|
- candidates: `data/arxiv-agent-candidates-2026-07-09-to-2026-07-27.json`
|
|
- selected: `data/arxiv-agent-papers-2026-07-09-to-2026-07-27.json`
|
|
- run record: `collection-runs/2026-07-27-arxiv-incremental-refresh.md`
|
|
|
|
## What Changed
|
|
|
|
Evaluation remains the strongest high-recall signal: 137 of 157 selected candidates received the broad
|
|
`agent-evaluation` tag, and 119 received `tool-use`. These tags overlap and are not quality judgments.
|
|
|
|
The useful new evaluation clusters are:
|
|
|
|
1. **Protocol validity and dynamic environments**: benchmarks increasingly vary tools, MCP servers,
|
|
interaction protocols and real deployment conditions.
|
|
2. **Trajectory diagnosis**: new work targets failure observability, attribution, runtime intervention
|
|
and recovery rather than only final scores.
|
|
3. **Independent verification**: several candidates emphasize deterministic gates, effect-scored traces,
|
|
provenance and reward-hacking-resistant validators.
|
|
4. **Regression and cost accounting**: average improvement is insufficient when a component fixes some
|
|
tasks while breaking previously successful ones.
|
|
5. **Interactive coding work**: coding-agent evaluation continues moving toward project building,
|
|
production trajectories, payment/backend integration and repository-level optimization.
|
|
|
|
## Priority Candidates
|
|
|
|
- [The Regression Tax](https://arxiv.org/abs/2607.22520)
|
|
- [Do Agent Benchmarks Measure Capability?](https://arxiv.org/abs/2607.22368)
|
|
- [AgentLens](https://arxiv.org/abs/2607.06624)
|
|
- [DynamicMCPBench](https://arxiv.org/abs/2607.20531)
|
|
- [AgentDebugX](https://arxiv.org/abs/2607.18754)
|
|
- [Reason Less, Verify More](https://arxiv.org/abs/2607.07405)
|
|
- [Baselines Before Architecture](https://arxiv.org/abs/2607.13085)
|
|
|
|
All remain `queued`; the descriptions above are based on title/abstract triage.
|
|
|
|
## Caveats
|
|
|
|
- The official export API was unavailable during this run, so the official arXiv search interface was
|
|
used through the new collector backend.
|
|
- Search queries are overlapping and broad.
|
|
- Automated relevance scoring can admit domain-specific uses of Agents with limited architectural value.
|
|
- No new paper from this increment has been promoted to `skimmed` or evidence-backed status.
|