Files
agent/papers/corpus-summary-2026-07-27.md
2026-07-27 16:28:23 +08:00

2.5 KiB

Paper Corpus Increment: 2026-07-09 to 2026-07-27

status: generated-summary

Scope

  • source: arXiv official search
  • window: 2026-07-09 to 2026-07-27
  • unique candidates: 249
  • high-relevance candidates: 157
  • newly written paper items: 151
  • selected items already in the repository: 6
  • reserve candidates: 92

Data

  • candidates: data/arxiv-agent-candidates-2026-07-09-to-2026-07-27.json
  • selected: data/arxiv-agent-papers-2026-07-09-to-2026-07-27.json
  • run record: collection-runs/2026-07-27-arxiv-incremental-refresh.md

What Changed

Evaluation remains the strongest high-recall signal: 137 of 157 selected candidates received the broad agent-evaluation tag, and 119 received tool-use. These tags overlap and are not quality judgments.

The useful new evaluation clusters are:

  1. Protocol validity and dynamic environments: benchmarks increasingly vary tools, MCP servers, interaction protocols and real deployment conditions.
  2. Trajectory diagnosis: new work targets failure observability, attribution, runtime intervention and recovery rather than only final scores.
  3. Independent verification: several candidates emphasize deterministic gates, effect-scored traces, provenance and reward-hacking-resistant validators.
  4. Regression and cost accounting: average improvement is insufficient when a component fixes some tasks while breaking previously successful ones.
  5. Interactive coding work: coding-agent evaluation continues moving toward project building, production trajectories, payment/backend integration and repository-level optimization.

Priority Candidates

All remain queued; the descriptions above are based on title/abstract triage.

Caveats

  • The official export API was unavailable during this run, so the official arXiv search interface was used through the new collector backend.
  • Search queries are overlapping and broad.
  • Automated relevance scoring can admit domain-specific uses of Agents with limited architectural value.
  • No new paper from this increment has been promoted to skimmed or evidence-backed status.