# Paper Corpus Increment: 2026-07-09 to 2026-07-27 status: generated-summary ## Scope - source: arXiv official search - window: 2026-07-09 to 2026-07-27 - unique candidates: 249 - high-relevance candidates: 157 - newly written paper items: 151 - selected items already in the repository: 6 - reserve candidates: 92 ## Data - candidates: `data/arxiv-agent-candidates-2026-07-09-to-2026-07-27.json` - selected: `data/arxiv-agent-papers-2026-07-09-to-2026-07-27.json` - run record: `collection-runs/2026-07-27-arxiv-incremental-refresh.md` ## What Changed Evaluation remains the strongest high-recall signal: 137 of 157 selected candidates received the broad `agent-evaluation` tag, and 119 received `tool-use`. These tags overlap and are not quality judgments. The useful new evaluation clusters are: 1. **Protocol validity and dynamic environments**: benchmarks increasingly vary tools, MCP servers, interaction protocols and real deployment conditions. 2. **Trajectory diagnosis**: new work targets failure observability, attribution, runtime intervention and recovery rather than only final scores. 3. **Independent verification**: several candidates emphasize deterministic gates, effect-scored traces, provenance and reward-hacking-resistant validators. 4. **Regression and cost accounting**: average improvement is insufficient when a component fixes some tasks while breaking previously successful ones. 5. **Interactive coding work**: coding-agent evaluation continues moving toward project building, production trajectories, payment/backend integration and repository-level optimization. ## Priority Candidates - [The Regression Tax](https://arxiv.org/abs/2607.22520) - [Do Agent Benchmarks Measure Capability?](https://arxiv.org/abs/2607.22368) - [AgentLens](https://arxiv.org/abs/2607.06624) - [DynamicMCPBench](https://arxiv.org/abs/2607.20531) - [AgentDebugX](https://arxiv.org/abs/2607.18754) - [Reason Less, Verify More](https://arxiv.org/abs/2607.07405) - [Baselines Before Architecture](https://arxiv.org/abs/2607.13085) All remain `queued`; the descriptions above are based on title/abstract triage. ## Caveats - The official export API was unavailable during this run, so the official arXiv search interface was used through the new collector backend. - Search queries are overlapping and broad. - Automated relevance scoring can admit domain-specific uses of Agents with limited architectural value. - No new paper from this increment has been promoted to `skimmed` or evidence-backed status.