Files
agent/papers/corpus-summary-2026-07-27.md
T
2026-07-27 16:28:23 +08:00

58 lines
2.5 KiB
Markdown

# Paper Corpus Increment: 2026-07-09 to 2026-07-27
status: generated-summary
## Scope
- source: arXiv official search
- window: 2026-07-09 to 2026-07-27
- unique candidates: 249
- high-relevance candidates: 157
- newly written paper items: 151
- selected items already in the repository: 6
- reserve candidates: 92
## Data
- candidates: `data/arxiv-agent-candidates-2026-07-09-to-2026-07-27.json`
- selected: `data/arxiv-agent-papers-2026-07-09-to-2026-07-27.json`
- run record: `collection-runs/2026-07-27-arxiv-incremental-refresh.md`
## What Changed
Evaluation remains the strongest high-recall signal: 137 of 157 selected candidates received the broad
`agent-evaluation` tag, and 119 received `tool-use`. These tags overlap and are not quality judgments.
The useful new evaluation clusters are:
1. **Protocol validity and dynamic environments**: benchmarks increasingly vary tools, MCP servers,
interaction protocols and real deployment conditions.
2. **Trajectory diagnosis**: new work targets failure observability, attribution, runtime intervention
and recovery rather than only final scores.
3. **Independent verification**: several candidates emphasize deterministic gates, effect-scored traces,
provenance and reward-hacking-resistant validators.
4. **Regression and cost accounting**: average improvement is insufficient when a component fixes some
tasks while breaking previously successful ones.
5. **Interactive coding work**: coding-agent evaluation continues moving toward project building,
production trajectories, payment/backend integration and repository-level optimization.
## Priority Candidates
- [The Regression Tax](https://arxiv.org/abs/2607.22520)
- [Do Agent Benchmarks Measure Capability?](https://arxiv.org/abs/2607.22368)
- [AgentLens](https://arxiv.org/abs/2607.06624)
- [DynamicMCPBench](https://arxiv.org/abs/2607.20531)
- [AgentDebugX](https://arxiv.org/abs/2607.18754)
- [Reason Less, Verify More](https://arxiv.org/abs/2607.07405)
- [Baselines Before Architecture](https://arxiv.org/abs/2607.13085)
All remain `queued`; the descriptions above are based on title/abstract triage.
## Caveats
- The official export API was unavailable during this run, so the official arXiv search interface was
used through the new collector backend.
- Search queries are overlapping and broad.
- Automated relevance scoring can admit domain-specific uses of Agents with limited architectural value.
- No new paper from this increment has been promoted to `skimmed` or evidence-backed status.