2.5 KiB
2.5 KiB
Paper Corpus Increment: 2026-07-09 to 2026-07-27
status: generated-summary
Scope
- source: arXiv official search
- window: 2026-07-09 to 2026-07-27
- unique candidates: 249
- high-relevance candidates: 157
- newly written paper items: 151
- selected items already in the repository: 6
- reserve candidates: 92
Data
- candidates:
data/arxiv-agent-candidates-2026-07-09-to-2026-07-27.json - selected:
data/arxiv-agent-papers-2026-07-09-to-2026-07-27.json - run record:
collection-runs/2026-07-27-arxiv-incremental-refresh.md
What Changed
Evaluation remains the strongest high-recall signal: 137 of 157 selected candidates received the broad
agent-evaluation tag, and 119 received tool-use. These tags overlap and are not quality judgments.
The useful new evaluation clusters are:
- Protocol validity and dynamic environments: benchmarks increasingly vary tools, MCP servers, interaction protocols and real deployment conditions.
- Trajectory diagnosis: new work targets failure observability, attribution, runtime intervention and recovery rather than only final scores.
- Independent verification: several candidates emphasize deterministic gates, effect-scored traces, provenance and reward-hacking-resistant validators.
- Regression and cost accounting: average improvement is insufficient when a component fixes some tasks while breaking previously successful ones.
- Interactive coding work: coding-agent evaluation continues moving toward project building, production trajectories, payment/backend integration and repository-level optimization.
Priority Candidates
- The Regression Tax
- Do Agent Benchmarks Measure Capability?
- AgentLens
- DynamicMCPBench
- AgentDebugX
- Reason Less, Verify More
- Baselines Before Architecture
All remain queued; the descriptions above are based on title/abstract triage.
Caveats
- The official export API was unavailable during this run, so the official arXiv search interface was used through the new collector backend.
- Search queries are overlapping and broad.
- Automated relevance scoring can admit domain-specific uses of Agents with limited architectural value.
- No new paper from this increment has been promoted to
skimmedor evidence-backed status.