Refresh papers and define Agent evaluation

This commit is contained in:
wuyang
2026-07-27 16:28:23 +08:00
parent 8e4ff3779b
commit df475e8d90
180 changed files with 31313 additions and 160 deletions
@@ -0,0 +1,61 @@
# Paper: AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
---
type: paper
title: "AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation"
authors: Andrey Podivilov, Vadim Lomshakov, Sergey Savin, Matvei Startsev, Roman Pozharskiy, Maksim Parshin, Sergey Nikolenko
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.06624
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-14
updated_at: 2026-07-14
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- planning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 19
collection_queries: agent-evaluation, coding-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-evaluation, coding-agent
- inferred topics: agent-evaluation, coding-agent, planning, tool-use
- arXiv categories: cs.AI
- collection score: 19
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.06624
@@ -0,0 +1,60 @@
# Paper: Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents
---
type: paper
title: "Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents"
authors: Vikas Reddy, Sumanth Reddy Challaram, Abhishek Basu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.07405
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-11
updated_at: 2026-07-11
status: queued
relevance: high
topics:
- agent-evaluation
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: tool-use
- inferred topics: agent-evaluation, reasoning, tool-use
- arXiv categories: cs.AI
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.07405
@@ -0,0 +1,61 @@
# Paper: Out of Sight: Compression-Aware Content Protection against Agentic Crawlers
---
type: paper
title: "Out of Sight: Compression-Aware Content Protection against Agentic Crawlers"
authors: Xuefei Wang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.08180
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-09
updated_at: 2026-07-09
status: queued
relevance: high
topics:
- computer-use
- memory
- reasoning
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: agentic-ai
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agentic-ai
- inferred topics: computer-use, memory, reasoning, workflow-agent
- arXiv categories: cs.CR
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.08180
@@ -0,0 +1,63 @@
# Paper: Multi-Agent Firewall Architecture for Privacy Protection of Sensitive Data in Interactions with Language Models
---
type: paper
title: Multi-Agent Firewall Architecture for Privacy Protection of Sensitive Data in Interactions with Language Models
authors: Hugo García Cuesta, Pablo Mateo Torrejón, Alfonso Sánchez-Macián
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.08282
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-09
updated_at: 2026-07-09
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- computer-use
- multi-agent
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: multi-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: multi-agent-llm
- inferred topics: agent-evaluation, agent-safety, computer-use, multi-agent, tool-use, workflow-agent
- arXiv categories: cs.CR
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.08282
@@ -0,0 +1,65 @@
# Paper: Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents
---
type: paper
title: "Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents"
authors: Yixian Zhang, Huanming Zhang, Feng Gao, Xiao Li, Zhihao Liu, Chunyang Zhu, Jiaxing Qiu, Yuchen Yan, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.08448
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-15
updated_at: 2026-07-15
status: queued
relevance: high
topics:
- coding-agent
- computer-use
- embodied-agent
- memory
- planning
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.RO
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 20
collection_queries: coding-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: coding-agent
- inferred topics: coding-agent, computer-use, embodied-agent, memory, planning, rag, reasoning, tool-use
- arXiv categories: cs.RO
- collection score: 20
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.08448
@@ -0,0 +1,61 @@
# Paper: SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets
---
type: paper
title: "SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets"
authors: Shilin Ou, Yifan Xu, Luyao Zhang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.08681
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-09
updated_at: 2026-07-09
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- planning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: agent-evaluation, agentic-ai, autonomous-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-evaluation, agentic-ai, autonomous-agent-llm
- inferred topics: agent-evaluation, agent-safety, planning, tool-use
- arXiv categories: cs.AI
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.08681
@@ -0,0 +1,63 @@
# Paper: Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
---
type: paper
title: "Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents"
authors: Yifan Wu, Lizhu Zhang, Yuhang Zhou, Mingyi Wang, Bo Peng, Serena Li, Xiangjun Fan, Zhuokai Zhao
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.08716
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-09
updated_at: 2026-07-09
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- memory
- planning
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: agent-memory
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-memory
- inferred topics: agent-evaluation, computer-use, memory, planning, rag, tool-use
- arXiv categories: cs.AI
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.08716
@@ -0,0 +1,64 @@
# Paper: GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning
---
type: paper
title: "GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning"
authors: Maureese Williams, Dymitr Nowicki
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.08894
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-09
updated_at: 2026-07-09
status: queued
relevance: high
topics:
- coding-agent
- computer-use
- embodied-agent
- planning
- tool-use
- workflow-agent
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: language-agent, planning-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent, planning-agent
- inferred topics: coding-agent, computer-use, embodied-agent, planning, tool-use, workflow-agent, world-model
- arXiv categories: cs.AI
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.08894
@@ -0,0 +1,62 @@
# Paper: Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution
---
type: paper
title: "Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution"
authors: Ning Liu, Kalle Kujanpää, Zhaoxuan Zhu, P Aditya Sreekar, Kaiwen Liu, Chuanneng Sun, Jorge Marchena Menendez, Matthew Bales, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.08960
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-09
updated_at: 2026-07-09
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- memory
- multi-agent
- reasoning
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 19
collection_queries: multi-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: multi-agent-llm
- inferred topics: agent-evaluation, computer-use, memory, multi-agent, reasoning
- arXiv categories: cs.LG
- collection score: 19
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.08960
@@ -0,0 +1,61 @@
# Paper: SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation
---
type: paper
title: "SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation"
authors: Sijia Gu, Noor Nashid, Ali Mesbah
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.08983
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-09
updated_at: 2026-07-09
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: coding-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: coding-agent
- inferred topics: agent-evaluation, coding-agent, rag, tool-use
- arXiv categories: cs.SE
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.08983
@@ -0,0 +1,62 @@
# Paper: Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls
---
type: paper
title: "Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls"
authors: Saroj Gopali, Bipin Chhetri, Deepika Giri, Sima Siami-Namini, Akbar Siami Namin
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.09076
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-09
updated_at: 2026-07-09
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- planning
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: agentic-ai
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agentic-ai
- inferred topics: agent-evaluation, agent-safety, planning, reasoning, tool-use
- arXiv categories: cs.AI
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.09076
@@ -0,0 +1,61 @@
# Paper: AgentKGV: Agentic LLM-RAG Framework with Two-Stage Training for the Fact Verification of Knowledge Graphs
---
type: paper
title: "AgentKGV: Agentic LLM-RAG Framework with Two-Stage Training for the Fact Verification of Knowledge Graphs"
authors: Yumin Heo, Hyeon-gu Lee, Sumin Seo, Youngjoong Ko
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.09092
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-10
updated_at: 2026-07-10
status: queued
relevance: high
topics:
- agent-evaluation
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: rag-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: rag-agent
- inferred topics: agent-evaluation, rag, reasoning, tool-use
- arXiv categories: cs.CL
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.09092
@@ -0,0 +1,65 @@
# Paper: Multi-Agent LLM Collaboration for Unit Test Generation via Human-Testing-Inspired Workflows
---
type: paper
title: Multi-Agent LLM Collaboration for Unit Test Generation via Human-Testing-Inspired Workflows
authors: Quanjun Zhang, Ye Shang, Siqi Gu, Jianyi Zhou, Chunrong Fang, Zhenyu Chen, Liang Xiao
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.09101
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-10
updated_at: 2026-07-10
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- multi-agent
- planning
- rag
- reasoning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: multi-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: multi-agent-llm
- inferred topics: agent-evaluation, coding-agent, multi-agent, planning, rag, reasoning, tool-use, workflow-agent
- arXiv categories: cs.SE
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.09101
@@ -0,0 +1,62 @@
# Paper: KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling
---
type: paper
title: "KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling"
authors: Peng Kuang, Haibo Jin, Xiaoyu Han, Yanli Wang, Xiaopeng Yuan, Ye Yu, Kaidi Xu, Haohan Wang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.09153
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-10
updated_at: 2026-07-10
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- computer-use
- memory
- multi-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: multi-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: multi-agent-llm
- inferred topics: agent-evaluation, coding-agent, computer-use, memory, multi-agent
- arXiv categories: cs.AI
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.09153
@@ -0,0 +1,63 @@
# Paper: Malaika: Understanding Malware through Tri-Grounded Agentic Reasoning
---
type: paper
title: "Malaika: Understanding Malware through Tri-Grounded Agentic Reasoning"
authors: Xingzhi Qian, Xinran Zheng, Yiling He, Lorenzo Cavallaro
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.09179
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-10
updated_at: 2026-07-10
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- multi-agent
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: multi-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: multi-agent-llm
- inferred topics: agent-evaluation, agent-safety, multi-agent, rag, reasoning, tool-use
- arXiv categories: cs.CR
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.09179
@@ -0,0 +1,61 @@
# Paper: Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents
---
type: paper
title: "Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents"
authors: Izumi Takahara, Teruyasu Mizoguchi
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.09195
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-10
updated_at: 2026-07-10
status: queued
relevance: high
topics:
- agent-evaluation
- planning
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: planning-agent, tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: planning-agent, tool-use
- inferred topics: agent-evaluation, planning, reasoning, tool-use
- arXiv categories: cs.AI
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.09195
@@ -0,0 +1,60 @@
# Paper: Writing Bug Reports for Software Repair Agents: What Information Matters Most?
---
type: paper
title: "Writing Bug Reports for Software Repair Agents: What Information Matters Most?"
authors: Vincenzo Luigi Bruno, Alessandro Giagnorio, Daniele Bifolco, Leon Wienges, Massimiliano Di Penta, Gabriele Bavota
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.09553
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-10
updated_at: 2026-07-10
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: coding-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: coding-agent
- inferred topics: agent-evaluation, coding-agent, workflow-agent
- arXiv categories: cs.SE
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.09553
@@ -0,0 +1,65 @@
# Paper: VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents
---
type: paper
title: "VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents"
authors: Katherine Swinea, Kshitiz Aryal, Lopamudra Praharaj, Maanak Gupta
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.09653
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-10
updated_at: 2026-07-10
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- multi-agent
- planning
- rag
- reasoning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 20
collection_queries: multi-agent-llm, planning-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: multi-agent-llm, planning-agent
- inferred topics: agent-evaluation, agent-safety, multi-agent, planning, rag, reasoning, tool-use, workflow-agent
- arXiv categories: cs.CR
- collection score: 20
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.09653
@@ -0,0 +1,62 @@
# Paper: Do These Violent Delights Have Violent Ends? Measuring the Post-Merge Fate of Agentic Code
---
type: paper
title: Do These Violent Delights Have Violent Ends? Measuring the Post-Merge Fate of Agentic Code
authors: Chunqiu Steven Xia, Courtney Miller
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.09902
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-10
updated_at: 2026-07-10
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- coding-agent
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: coding-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: coding-agent
- inferred topics: agent-evaluation, agent-safety, coding-agent, rag, tool-use
- arXiv categories: cs.SE
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.09902
@@ -0,0 +1,59 @@
# Paper: Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?
---
type: paper
title: "Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?"
authors: Jiale Liu, Huajun Xi, Shaokun Zhang, Yifan Zeng, Tianwei Yue, Chi Wang, Jian Kang, Qingyun Wu, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.09996
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-10
updated_at: 2026-07-10
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: ai-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: ai-agent
- inferred topics: agent-evaluation, computer-use
- arXiv categories: cs.AI
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.09996
@@ -0,0 +1,59 @@
# Paper: Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation
---
type: paper
title: "Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation"
authors: Dongping Liu, Aoyu Zhang, Luyao Zhang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.10057
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-10
updated_at: 2026-07-10
status: queued
relevance: high
topics:
- agent-evaluation
- reasoning
methods:
-
benchmarks:
-
models:
-
datasets:
- quant-ph
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: ai-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: ai-agent
- inferred topics: agent-evaluation, reasoning
- arXiv categories: quant-ph
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.10057
@@ -0,0 +1,62 @@
# Paper: TGMS: An Agent-Native Bi-Temporal Graph Management System
---
type: paper
title: "TGMS: An Agent-Native Bi-Temporal Graph Management System"
authors: Xiaofei Zhang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.10265
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-11
updated_at: 2026-07-11
status: queued
relevance: high
topics:
- agent-evaluation
- planning
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.DB
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: planning-agent, rag-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: planning-agent, rag-agent
- inferred topics: agent-evaluation, planning, rag, reasoning, tool-use
- arXiv categories: cs.DB
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.10265
@@ -0,0 +1,60 @@
# Paper: Can Agentic Trading Systems Pay for Their Own Intelligence?
---
type: paper
title: Can Agentic Trading Systems Pay for Their Own Intelligence?
authors: Qiqi Duan, Changlun Li, Chen Wang, Fan Zhang, Mengxiang Wang, Dayi Miao, Peixian Ma, Jiangpeng Yan, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.10286
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-11
updated_at: 2026-07-11
status: queued
relevance: high
topics:
- agent-evaluation
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: tool-use
- inferred topics: agent-evaluation, reasoning, tool-use
- arXiv categories: cs.AI
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.10286
@@ -0,0 +1,61 @@
# Paper: GRASP: GRanularity-Aware Search Policy for Agentic RAG
---
type: paper
title: "GRASP: GRanularity-Aware Search Policy for Agentic RAG"
authors: Varun Gandhi, Jaewook Lee, Shantanu Todmal, Franck Dernoncourt, Ryan Rossi, Zichao Wang, Andrew Lan
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.10463
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-11
updated_at: 2026-07-11
status: queued
relevance: high
topics:
- agent-evaluation
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: rag-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: rag-agent
- inferred topics: agent-evaluation, rag, reasoning, tool-use
- arXiv categories: cs.AI
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.10463
@@ -0,0 +1,60 @@
# Paper: NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations
---
type: paper
title: "NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations"
authors: Ruksat Khan Shayoni, Muhammad Faraz Shoaib, S M Asif Hossain, M. F. Mridha
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.10490
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-11
updated_at: 2026-07-11
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: tool-use
- inferred topics: agent-evaluation, agent-safety, tool-use
- arXiv categories: cs.CR
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.10490
@@ -0,0 +1,63 @@
# Paper: MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference
---
type: paper
title: "MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference"
authors: Venkatesha Matam, Keon Kim
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.10582
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-12
updated_at: 2026-07-12
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- memory
- planning
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: planning-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: planning-agent
- inferred topics: agent-evaluation, coding-agent, memory, planning, reasoning, tool-use
- arXiv categories: cs.LG
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.10582
@@ -0,0 +1,63 @@
# Paper: The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory
---
type: paper
title: "The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory"
authors: Yixiong Chen, Xinyi Bai, Alan Yuille
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.10608
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-12
updated_at: 2026-07-12
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- memory
- planning
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: ai-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: ai-agent
- inferred topics: agent-evaluation, computer-use, memory, planning, rag, tool-use
- arXiv categories: cs.AI
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.10608
@@ -0,0 +1,60 @@
# Paper: Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging
---
type: paper
title: "Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging"
authors: Siyi Chen, Jiahe Ying, Yixuan Jia, Yuxuan Gu, Enze Ye, Weimin Bai, Zhijun Zeng, Shaochi Ren, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.10789
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-12
updated_at: 2026-07-12
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- planning
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: coding-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: coding-agent
- inferred topics: agent-evaluation, coding-agent, planning
- arXiv categories: cs.AI
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.10789
@@ -0,0 +1,66 @@
# Paper: A Multi-Agent Framework for Zero-Dimensional Reduced-Order Model Planning
---
type: paper
title: A Multi-Agent Framework for Zero-Dimensional Reduced-Order Model Planning
authors: Bingteng Sun, Hao Yin, Yiling Chen, Renjie Xiao, Lei Xie, Shanyou Wang, Ruonan Wang, Shubao Chen, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.10994
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-12
updated_at: 2026-07-12
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- embodied-agent
- multi-agent
- planning
- rag
- reasoning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 22
collection_queries: multi-agent-llm, planning-agent, rag-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: multi-agent-llm, planning-agent, rag-agent
- inferred topics: agent-evaluation, computer-use, embodied-agent, multi-agent, planning, rag, reasoning, tool-use, workflow-agent
- arXiv categories: cs.LG
- collection score: 22
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.10994
@@ -0,0 +1,60 @@
# Paper: BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services
---
type: paper
title: "BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services"
authors: Yuzhe Guo, Mengzhou Wu, Yuan Cao, Jialei Wei, Dezhi Ran, Wei Yang, Tao Xie
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.11042
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-12
updated_at: 2026-07-12
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: coding-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: coding-agent
- inferred topics: agent-evaluation, coding-agent, tool-use
- arXiv categories: cs.SE
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.11042
@@ -0,0 +1,61 @@
# Paper: Retrieval-Oriented Code Representations in Agentic Bug Localization
---
type: paper
title: Retrieval-Oriented Code Representations in Agentic Bug Localization
authors: Genevieve Caumartin, Tse-Hsun, Chen, Diego Elias Costa
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.11046
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-12
updated_at: 2026-07-12
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- computer-use
- rag
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: coding-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: coding-agent
- inferred topics: agent-evaluation, coding-agent, computer-use, rag
- arXiv categories: cs.SE
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.11046
@@ -0,0 +1,64 @@
# Paper: VIA: Visual Interface Agent for Robot Control
---
type: paper
title: "VIA: Visual Interface Agent for Robot Control"
authors: Hengyuan Hu, Priya Sundaresan, Jensen Gao, Dorsa Sadigh
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.11119
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- coding-agent
- computer-use
- embodied-agent
- planning
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.RO
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: web-gui-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: web-gui-agent
- inferred topics: coding-agent, computer-use, embodied-agent, planning, rag, reasoning, tool-use
- arXiv categories: cs.RO
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.11119
@@ -0,0 +1,61 @@
# Paper: ToolAtlas: Learning Once, Reusing Everywhere with Tool-Side Memory
---
type: paper
title: "ToolAtlas: Learning Once, Reusing Everywhere with Tool-Side Memory"
authors: Yue Fang, Zhibang Yang, Fangkai Yang, Xiaoting Qin, Liqun Li, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.11126
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- memory
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: tool-use
- inferred topics: agent-evaluation, computer-use, memory, tool-use
- arXiv categories: cs.LG
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.11126
@@ -0,0 +1,61 @@
# Paper: NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management
---
type: paper
title: "NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management"
authors: Changlun Li, Peixian Ma, Qiqi Duan, Zhenyu Lin, Peineng Wu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.11141
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- multi-agent
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: multi-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: multi-agent-llm
- inferred topics: agent-evaluation, agent-safety, multi-agent, tool-use
- arXiv categories: cs.AI
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.11141
@@ -0,0 +1,61 @@
# Paper: SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL
---
type: paper
title: "SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL"
authors: Bowen Lv, Xiao Liu, Yanyu Ren, Hanyu Lai, Bohao Jing, Hanchen Zhang, Yanxiao Zhao, Shuntian Yao, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.11185
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- computer-use
- multi-agent
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: web-gui-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: web-gui-agent
- inferred topics: computer-use, multi-agent, tool-use, workflow-agent
- arXiv categories: cs.AI
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.11185
@@ -0,0 +1,61 @@
# Paper: Multi-Agent LLMs Fail to Explore Each Other
---
type: paper
title: Multi-Agent LLMs Fail to Explore Each Other
authors: Hyeong Kyu Choi, Jiatong Li, Wendi Li, Xin Eric Wang, Sharon Li
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.11250
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- multi-agent
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.MA
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: multi-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: multi-agent-llm
- inferred topics: agent-evaluation, computer-use, multi-agent, tool-use
- arXiv categories: cs.MA
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.11250
@@ -0,0 +1,62 @@
# Paper: StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure
---
type: paper
title: "StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure"
authors: Wenyi Wu, Sibo Zhu, Kun Zhou, Aayush Salvi, Zixuan Song, Biwei Huang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.11388
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- computer-use
- planning
- reasoning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: web-gui-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: web-gui-agent
- inferred topics: computer-use, planning, reasoning, tool-use, workflow-agent
- arXiv categories: cs.AI
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.11388
@@ -0,0 +1,61 @@
# Paper: ToFu: A White-Box, Token-Efficient Agent Harness for Researchers
---
type: paper
title: "ToFu: A White-Box, Token-Efficient Agent Harness for Researchers"
authors: Junhao Ruan, Yuan Ge, Bei Li, Yongjing Yin, Yuchun Fan, Xin Chen, Jingang Wang, Chenglong Wang, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.11423
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: tool-use
- inferred topics: agent-evaluation, coding-agent, tool-use, workflow-agent
- arXiv categories: cs.CL
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.11423
@@ -0,0 +1,61 @@
# Paper: UMoE:Unlocking Every Expert in Domain-Specific Training
---
type: paper
title: "UMoE:Unlocking Every Expert in Domain-Specific Training"
authors: Xuefeng Li, Pengfei Liu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.11444
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: coding-agent, tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: coding-agent, tool-use
- inferred topics: agent-evaluation, coding-agent, rag, tool-use
- arXiv categories: cs.CL
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.11444
@@ -0,0 +1,61 @@
# Paper: Heuristic Learning for Active Flow Control Using Coding Agents
---
type: paper
title: Heuristic Learning for Active Flow Control Using Coding Agents
authors: Paul Garnier, Jonathan Viquerat, Elie Hachem
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.11565
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- tool-use
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: coding-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: coding-agent
- inferred topics: agent-evaluation, coding-agent, tool-use, world-model
- arXiv categories: cs.LG
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.11565
@@ -0,0 +1,63 @@
# Paper: When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems
---
type: paper
title: "When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems"
authors: Yibo Hu, Ren Wang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.11751
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- coding-agent
- multi-agent
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: multi-agent-llm, tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: multi-agent-llm, tool-use
- inferred topics: agent-evaluation, agent-safety, coding-agent, multi-agent, rag, tool-use
- arXiv categories: cs.CR
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.11751
@@ -0,0 +1,62 @@
# Paper: Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability
---
type: paper
title: "Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability"
authors: Said Elnaffar, Farzad Rashidi
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.12056
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- rag
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: ai-agent, web-gui-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: ai-agent, web-gui-agent
- inferred topics: agent-evaluation, computer-use, rag, tool-use, workflow-agent
- arXiv categories: cs.AI
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.12056
@@ -0,0 +1,61 @@
# Paper: Beyond Test Presence: Assessing the Quality and Robustness of Agent-Generated Tests in Open-Source Projects
---
type: paper
title: "Beyond Test Presence: Assessing the Quality and Robustness of Agent-Generated Tests in Open-Source Projects"
authors: Preet Jhanglani, Zeel Kaushal Desai, Vidhi Kansara, Eman Abdullah AlOmar
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.12068
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- coding-agent
- rag
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: ai-agent, coding-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: ai-agent, coding-agent
- inferred topics: agent-evaluation, agent-safety, coding-agent, rag
- arXiv categories: cs.SE
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.12068
@@ -0,0 +1,60 @@
# Paper: An Agentic AI Scientific Community for Automated Neural Operator Discovery
---
type: paper
title: An Agentic AI Scientific Community for Automated Neural Operator Discovery
authors: Luis Loo, Ulisses Braga-Neto
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.12122
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- agent-evaluation
- planning
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: agentic-ai
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agentic-ai
- inferred topics: agent-evaluation, planning, world-model
- arXiv categories: cs.LG
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.12122
@@ -0,0 +1,60 @@
# Paper: Fine-Tuned Multi-Agent Framework for Detecting OCEAN in Life Narratives
---
type: paper
title: Fine-Tuned Multi-Agent Framework for Detecting OCEAN in Life Narratives
authors: Rasiq Hussain, Darshil Italiya, Joshua Oltmanns, Mehak Gupta
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.12215
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- agent-evaluation
- multi-agent
- reasoning
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: multi-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: multi-agent-llm
- inferred topics: agent-evaluation, multi-agent, reasoning
- arXiv categories: cs.CL
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.12215
@@ -0,0 +1,62 @@
# Paper: Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents
---
type: paper
title: "Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents"
authors: Ning Liu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.12267
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- agent-evaluation
- memory
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 19
collection_queries: language-agent, tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent, tool-use
- inferred topics: agent-evaluation, memory, rag, reasoning, tool-use
- arXiv categories: cs.LG
- collection score: 19
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.12267
@@ -0,0 +1,60 @@
# Paper: Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents
---
type: paper
title: "Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents"
authors: Yaopei Zeng, Congchao Wang, JianHang Chen, Nan Wang, Yurui Chang, Lu Lin
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.12397
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-14
updated_at: 2026-07-14
status: queued
relevance: high
topics:
- agent-evaluation
- memory
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: agent-evaluation
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-evaluation
- inferred topics: agent-evaluation, memory, tool-use
- arXiv categories: cs.AI
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.12397
@@ -0,0 +1,63 @@
# Paper: Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions
---
type: paper
title: "Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions"
authors: Huihao Jing, Wenbin Hu, Shaojin Chen, Haochen Shi, Sirui Zhang, Hanyu Yang, Changxuan Fan, Zhongwei Xie, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.12406
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-14
updated_at: 2026-07-14
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- memory
- rag
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: agentic-ai
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agentic-ai
- inferred topics: agent-evaluation, agent-safety, memory, rag, tool-use, workflow-agent
- arXiv categories: cs.AI
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.12406
@@ -0,0 +1,61 @@
# Paper: Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
---
type: paper
title: Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
authors: Yubo Wang, Jiarong Liang, Yuxuan Zhang, Xuye Liu, Cong Wei, Yuyu Zhang, Ping Nie, Wenhu Chen
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.12463
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-19
updated_at: 2026-07-19
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: coding-agent, function-calling, tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: coding-agent, function-calling, tool-use
- inferred topics: agent-evaluation, coding-agent, reasoning, tool-use
- arXiv categories: cs.AI
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.12463
@@ -0,0 +1,61 @@
# Paper: PalmClaw: A Native On-Device Agent Framework for Mobile Phones
---
type: paper
title: "PalmClaw: A Native On-Device Agent Framework for Mobile Phones"
authors: Hongru Cai, Yongqi Li, Ran Wei, Wenjie Li
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.13027
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-14
updated_at: 2026-07-14
status: queued
relevance: high
topics:
- computer-use
- memory
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: tool-use
- inferred topics: computer-use, memory, tool-use, workflow-agent
- arXiv categories: cs.CL
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.13027
@@ -0,0 +1,62 @@
# Paper: SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification
---
type: paper
title: "SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification"
authors: SingGuard Team
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.13081
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- computer-use
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: agent-safety, agentic-ai
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety, agentic-ai
- inferred topics: agent-evaluation, agent-safety, computer-use, reasoning, tool-use
- arXiv categories: cs.CR
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.13081
@@ -0,0 +1,61 @@
# Paper: Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testing
---
type: paper
title: "Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testing"
authors: Ananda Dhakal, Krish Neupane, Aarjan Chaudhary
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.13085
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-13
updated_at: 2026-07-13
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- coding-agent
- rag
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: coding-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: coding-agent
- inferred topics: agent-evaluation, agent-safety, coding-agent, rag
- arXiv categories: cs.CR
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.13085
@@ -0,0 +1,61 @@
# Paper: From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality
---
type: paper
title: "From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality"
authors: Suzhen Zhong, Shayan Noei, Bram Adams, Ying Zou
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.13196
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-14
updated_at: 2026-07-14
status: queued
relevance: high
topics:
- computer-use
- multi-agent
- planning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: ai-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: ai-agent
- inferred topics: computer-use, multi-agent, planning, tool-use
- arXiv categories: cs.SE
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.13196
@@ -0,0 +1,61 @@
# Paper: Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents
---
type: paper
title: "Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents"
authors: Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu, Levina Li, Dong Liu, Xiao Liang, Rui Sun, Yubei Li, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.13591
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-15
updated_at: 2026-07-15
status: queued
relevance: high
topics:
- agent-evaluation
- memory
- planning
- rag
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 19
collection_queries: planning-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: planning-agent
- inferred topics: agent-evaluation, memory, planning, rag
- arXiv categories: cs.CL
- collection score: 19
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.13591
@@ -0,0 +1,59 @@
# Paper: STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle
---
type: paper
title: "STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle"
authors: Sagar Deb, Ashwanth Krishnan
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.13618
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-15
updated_at: 2026-07-15
status: queued
relevance: high
topics:
- agent-evaluation
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: llm-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: llm-agent
- inferred topics: agent-evaluation, tool-use
- arXiv categories: cs.AI
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.13618
@@ -0,0 +1,60 @@
# Paper: AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
---
type: paper
title: "AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities"
authors: Kai Chen, Zichen Ding, Jiaye Ge, Shufan Jiang, Mo Li, Qingqiu Li, Zehao Li, Zonglin Li, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.13705
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-20
updated_at: 2026-07-20
status: queued
relevance: high
topics:
- agent-evaluation
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: autonomous-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: autonomous-agent-llm
- inferred topics: agent-evaluation, rag, tool-use
- arXiv categories: cs.AI
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.13705
@@ -0,0 +1,64 @@
# Paper: SPyCE: Skill-Policy Co-evolution for Multimodal Agents
---
type: paper
title: "SPyCE: Skill-Policy Co-evolution for Multimodal Agents"
authors: Ru Zhang, Weijie Qiu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.13854
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-15
updated_at: 2026-07-15
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- memory
- rag
- reasoning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: tool-use
- inferred topics: agent-evaluation, computer-use, memory, rag, reasoning, tool-use, workflow-agent
- arXiv categories: cs.CL
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.13854
@@ -0,0 +1,62 @@
# Paper: Agent Skill Security: Threat Models, Attacks, Defenses, and Evaluation
---
type: paper
title: "Agent Skill Security: Threat Models, Attacks, Defenses, and Evaluation"
authors: Sanket Badhe, Priyanka Tiwari
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.13987
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-15
updated_at: 2026-07-15
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- coding-agent
- planning
- rag
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: llm-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: llm-agent
- inferred topics: agent-evaluation, agent-safety, coding-agent, planning, rag
- arXiv categories: cs.CR
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.13987
@@ -0,0 +1,61 @@
# Paper: ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability
---
type: paper
title: "ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability"
authors: Weiting Liu, Jieyi Bi, Wanqi Zhou, Jianfeng Feng, Yining Ma, Ai Han, Wenlian Lu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.14145
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-14
updated_at: 2026-07-14
status: queued
relevance: high
topics:
- agent-evaluation
- planning
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: tool-use
- inferred topics: agent-evaluation, planning, reasoning, tool-use
- arXiv categories: cs.AI
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.14145
@@ -0,0 +1,61 @@
# Paper: Towards Reliable AI-Assisted Analog Design: Template-Constrained LLM Agents for SAR ADC Generation
---
type: paper
title: "Towards Reliable AI-Assisted Analog Design: Template-Constrained LLM Agents for SAR ADC Generation"
authors: Dimple Vijay Kochar, Hae-Seung Lee, Anantha P. Chandrakasan
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.14165
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-15
updated_at: 2026-07-15
status: queued
relevance: high
topics:
- planning
- rag
- workflow-agent
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: llm-agent, planning-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: llm-agent, planning-agent
- inferred topics: planning, rag, workflow-agent, world-model
- arXiv categories: cs.SE
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.14165
@@ -0,0 +1,61 @@
# Paper: ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System
---
type: paper
title: "ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System"
authors: Yutong He, Daibo Li, Guohong Li, Jiahe Geng, Zhengyang Huang, Can Ren, Zekun Zhang, Yifan Liu, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.14178
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-15
updated_at: 2026-07-15
status: queued
relevance: high
topics:
- agent-evaluation
- multi-agent
- rag
- reasoning
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: ai-agent, autonomous-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: ai-agent, autonomous-agent-llm
- inferred topics: agent-evaluation, multi-agent, rag, reasoning
- arXiv categories: cs.AI
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.14178
@@ -0,0 +1,61 @@
# Paper: MonteRET: AI Agent Enhancing Multimodal LLMs with Multi-granularity Knowledge Retrieval for Chest CT Report Generation
---
type: paper
title: "MonteRET: AI Agent Enhancing Multimodal LLMs with Multi-granularity Knowledge Retrieval for Chest CT Report Generation"
authors: Yi Lin, Yihao Ding, Elana Benishay, Elefterios Trikantzopoulos, David Nauheim, Hanley Ong, Jiang Bian, Hua Xu, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.14264
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-15
updated_at: 2026-07-15
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- computer-use
- rag
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CV
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: ai-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: ai-agent
- inferred topics: agent-evaluation, agent-safety, computer-use, rag
- arXiv categories: cs.CV
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.14264
@@ -0,0 +1,61 @@
# Paper: Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making
---
type: paper
title: "Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making"
authors: Amirhosein Ghasemabadi, Ruichen Chen, Bahador Rashidi, Di Niu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.14277
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-15
updated_at: 2026-07-15
status: queued
relevance: high
topics:
- agent-evaluation
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: llm-agent, tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: llm-agent, tool-use
- inferred topics: agent-evaluation, rag, reasoning, tool-use
- arXiv categories: cs.CL
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.14277
@@ -0,0 +1,61 @@
# Paper: CIPHER: A Decoupled Exploration-Selection Framework for Test-Time Scaling of Data Science Agents
---
type: paper
title: "CIPHER: A Decoupled Exploration-Selection Framework for Test-Time Scaling of Data Science Agents"
authors: Maxime Heuillet, Sharadind Peddiraju
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.14386
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-15
updated_at: 2026-07-15
status: queued
relevance: high
topics:
- agent-evaluation
- rag
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: ai-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: ai-agent
- inferred topics: agent-evaluation, rag, tool-use, workflow-agent
- arXiv categories: cs.AI
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.14386
@@ -0,0 +1,59 @@
# Paper: Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution
---
type: paper
title: "Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution"
authors: Harris Borman, Herman Wandabwa, Fusun Yu, Sandeepa Kannangara, Justin Liu, Anna Leontjeva, Ritchie Ng
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.14456
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-15
updated_at: 2026-07-15
status: queued
relevance: high
topics:
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: agentic-ai, tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agentic-ai, tool-use
- inferred topics: tool-use, workflow-agent
- arXiv categories: cs.SE
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.14456
@@ -0,0 +1,62 @@
# Paper: HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents
---
type: paper
title: "HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents"
authors: Hy Vision Team, Huawen Shen, Zhengyang Tang, Shangpin Peng, Liang Wu, Anran Zhang, Weinong Wang, Yiduo Guo, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.14548
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-16
updated_at: 2026-07-16
status: queued
relevance: high
topics:
- computer-use
- embodied-agent
- planning
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CV
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: web-gui-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: web-gui-agent
- inferred topics: computer-use, embodied-agent, planning, reasoning, tool-use
- arXiv categories: cs.CV
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.14548
@@ -0,0 +1,63 @@
# Paper: Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents
---
type: paper
title: "Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents"
authors: Shiyu Ying, Xuejie Cao, Yingfan Ma, Yuanhao Dong, Wenyu Chen, Bowen Song, Lin Zhu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.14573
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-21
updated_at: 2026-07-21
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- coding-agent
- computer-use
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 19
collection_queries: coding-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: coding-agent
- inferred topics: agent-evaluation, agent-safety, coding-agent, computer-use, rag, tool-use
- arXiv categories: cs.AI
- collection score: 19
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.14573
@@ -0,0 +1,63 @@
# Paper: MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers
---
type: paper
title: "MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers"
authors: Huanxi Liu, Kun Hu, Jiaqi Liao, Qiang Wang, Pengfei Qian, YuanZhao Zhai, Dawei Feng, Bo Ding, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.14642
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-16
updated_at: 2026-07-16
status: queued
relevance: high
topics:
- agent-evaluation
- planning
- rag
- reasoning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 21
collection_queries: llm-agent, planning-agent, tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: llm-agent, planning-agent, tool-use
- inferred topics: agent-evaluation, planning, rag, reasoning, tool-use, workflow-agent
- arXiv categories: cs.AI
- collection score: 21
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.14642
@@ -0,0 +1,62 @@
# Paper: MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents
---
type: paper
title: "MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents"
authors: Jifeng Gao, Kang Xia, Yi Zhang, Xiaobin Hong, Mingkai Lin, Xingshen Wei, Wenzhong Li, Sanglu Lu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.14651
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-16
updated_at: 2026-07-16
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- memory
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 20
collection_queries: llm-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: llm-agent
- inferred topics: agent-evaluation, agent-safety, memory, rag, tool-use
- arXiv categories: cs.CR
- collection score: 20
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.14651
@@ -0,0 +1,61 @@
# Paper: SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
---
type: paper
title: "SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning"
authors: Jinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.14777
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-16
updated_at: 2026-07-16
status: queued
relevance: high
topics:
- computer-use
- planning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: tool-use
- inferred topics: computer-use, planning, tool-use, workflow-agent
- arXiv categories: cs.CL
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.14777
@@ -0,0 +1,61 @@
# Paper: OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
---
type: paper
title: "OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios"
authors: Chengyu Shen, Yujie Fu, Gangtao Xin, Yanheng Hou, Wenlong Fei, Guojie Zhu, Jiawei Li, Hongcheng Gao, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.14989
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-16
updated_at: 2026-07-16
status: queued
relevance: high
topics:
- agent-evaluation
- planning
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 19
collection_queries: agent-evaluation, ai-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-evaluation, ai-agent
- inferred topics: agent-evaluation, planning, rag, tool-use
- arXiv categories: cs.CL
- collection score: 19
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.14989
@@ -0,0 +1,63 @@
# Paper: BrainPilot: Automating Brain Discovery with Agentic Research
---
type: paper
title: "BrainPilot: Automating Brain Discovery with Agentic Research"
authors: Haoxuan Li, Tianci Gao, Jianhe Li, Yang Fan, Runze Shi, Weiran Wang, Tianxiang Zhao, Zezhao Wu, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.15079
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-17
updated_at: 2026-07-17
status: queued
relevance: high
topics:
- agent-evaluation
- multi-agent
- rag
- reasoning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 20
collection_queries: ai-agent, tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: ai-agent, tool-use
- inferred topics: agent-evaluation, multi-agent, rag, reasoning, tool-use, workflow-agent
- arXiv categories: cs.AI
- collection score: 20
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.15079
@@ -0,0 +1,62 @@
# Paper: Plover: Steering GUI Agents through Plan-Centric Interaction
---
type: paper
title: "Plover: Steering GUI Agents through Plan-Centric Interaction"
authors: Madhumitha Venkatesan, Shicheng Wen, Jiajing Guo, Jorge Piazentin Ono, Liu Ren, Dongyu Liu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.15193
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-16
updated_at: 2026-07-16
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- planning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: web-gui-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: web-gui-agent
- inferred topics: agent-evaluation, computer-use, planning, tool-use, workflow-agent
- arXiv categories: cs.AI
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.15193
@@ -0,0 +1,60 @@
# Paper: ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
---
type: paper
title: "ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors"
authors: Christos Korgialas, Gabriel Lee Jun Rong, Dion Jia Xu Ho, Pai Chet Ng, Xiaoxiao Miao, Konstantinos N. Plataniotis
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.15246
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-16
updated_at: 2026-07-16
status: queued
relevance: high
topics:
- agent-evaluation
- multi-agent
- rag
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CV
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: agent-evaluation, multi-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-evaluation, multi-agent-llm
- inferred topics: agent-evaluation, multi-agent, rag
- arXiv categories: cs.CV
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.15246
@@ -0,0 +1,62 @@
# Paper: Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
---
type: paper
title: "Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents"
authors: Paul Kassianik, Blaine Nelson, Yaron Singer
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.15263
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-17
updated_at: 2026-07-17
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- embodied-agent
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: agent-evaluation, tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-evaluation, tool-use
- inferred topics: agent-evaluation, agent-safety, embodied-agent, reasoning, tool-use
- arXiv categories: cs.CR
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.15263
@@ -0,0 +1,63 @@
# Paper: AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery
---
type: paper
title: "AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery"
authors: Raunak B Sinha
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.15367
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-16
updated_at: 2026-07-16
status: queued
relevance: high
topics:
- agent-safety
- computer-use
- multi-agent
- planning
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: multi-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: multi-agent-llm
- inferred topics: agent-safety, computer-use, multi-agent, planning, rag, tool-use
- arXiv categories: cs.AI
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.15367
@@ -0,0 +1,60 @@
# Paper: Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
---
type: paper
title: "Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation"
authors: Jasmine Brazilek, Maheep Chaudhary, Zoe Lu, Miles Tidmarsh
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.15434
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-22
updated_at: 2026-07-22
status: queued
relevance: high
topics:
- agent-evaluation
- multi-agent
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.MA
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 21
collection_queries: agent-evaluation, ai-agent, multi-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-evaluation, ai-agent, multi-agent-llm
- inferred topics: agent-evaluation, multi-agent, tool-use
- arXiv categories: cs.MA
- collection score: 21
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.15434
@@ -0,0 +1,61 @@
# Paper: A Tool-Invariant Framework for Teaching and Assessing Computational Methods in the Age of Agentic AI
---
type: paper
title: A Tool-Invariant Framework for Teaching and Assessing Computational Methods in the Age of Agentic AI
authors: Larry Engelhardt
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.15518
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-16
updated_at: 2026-07-16
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- tool-use
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- physics.ed-ph
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: agentic-ai, language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agentic-ai, language-agent
- inferred topics: agent-evaluation, coding-agent, tool-use, world-model
- arXiv categories: physics.ed-ph
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.15518
@@ -0,0 +1,62 @@
# Paper: Symbolic Predicate-Guided Language Agents for Inverse Design of Perovskite Oxides
---
type: paper
title: Symbolic Predicate-Guided Language Agents for Inverse Design of Perovskite Oxides
authors: Dong Hyeon Mok, Seoin Back, Victor Fung, Guoxiang Hu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.15535
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-16
updated_at: 2026-07-16
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- multi-agent
- rag
- reasoning
methods:
-
benchmarks:
-
models:
-
datasets:
- cond-mat.mtrl-sci
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 19
collection_queries: language-agent, llm-agent, multi-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent, llm-agent, multi-agent-llm
- inferred topics: agent-evaluation, computer-use, multi-agent, rag, reasoning
- arXiv categories: cond-mat.mtrl-sci
- collection score: 19
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.15535
@@ -0,0 +1,60 @@
# Paper: SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents
---
type: paper
title: "SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents"
authors: Yanze Wang, Pengfei Yao, Tianyi Sun, Chuanrui Hu, Yan Xiao, Yunyun Han, Yifan Chen, Jun Sun, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.15557
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-23
updated_at: 2026-07-23
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- rag
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: llm-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: llm-agent
- inferred topics: agent-evaluation, agent-safety, rag
- arXiv categories: cs.CL
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.15557
@@ -0,0 +1,62 @@
# Paper: Scalable LLM Agent Tool Access in the Cloud
---
type: paper
title: Scalable LLM Agent Tool Access in the Cloud
authors: Mingxin Li, Enge Song, Yueshang Zuo, Xiaodong Liu, Rong Wen, Qiang Fu, Gianni Antichi, Jian He, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.15593
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-16
updated_at: 2026-07-16
status: queued
relevance: high
topics:
- agent-evaluation
- memory
- planning
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.DC
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: llm-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: llm-agent
- inferred topics: agent-evaluation, memory, planning, rag, tool-use
- arXiv categories: cs.DC
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.15593
@@ -0,0 +1,62 @@
# Paper: ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning
---
type: paper
title: "ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning"
authors: Shuaiyu Zhou, Fengpeng Yue, Zengjie Hu, Yuanzhe Shen, Chenyang Zhang, feng hong, Cao Liu, Ke Zeng
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.15660
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-17
updated_at: 2026-07-17
status: queued
relevance: high
topics:
- agent-evaluation
- planning
- reasoning
- tool-use
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 20
collection_queries: agent-evaluation, llm-agent, tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-evaluation, llm-agent, tool-use
- inferred topics: agent-evaluation, planning, reasoning, tool-use, world-model
- arXiv categories: cs.AI
- collection score: 20
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.15660
@@ -0,0 +1,63 @@
# Paper: Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents
---
type: paper
title: "Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents"
authors: Lujia Zhang, Xingzhou Chen, Hongwei Feng
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.15715
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-17
updated_at: 2026-07-17
status: queued
relevance: high
topics:
- agent-evaluation
- memory
- rag
- reasoning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: llm-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: llm-agent
- inferred topics: agent-evaluation, memory, rag, reasoning, tool-use, workflow-agent
- arXiv categories: cs.AI
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.15715
@@ -0,0 +1,64 @@
# Paper: AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
---
type: paper
title: "AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets"
authors: Ming Chen, Pranav Pai
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.15781
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-24
updated_at: 2026-07-24
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- coding-agent
- multi-agent
- planning
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 21
collection_queries: multi-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: multi-agent-llm
- inferred topics: agent-evaluation, agent-safety, coding-agent, multi-agent, planning, rag, tool-use
- arXiv categories: cs.AI
- collection score: 21
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.15781
@@ -0,0 +1,61 @@
# Paper: Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
---
type: paper
title: "Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning"
authors: Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, Karim Lakhani
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.16057
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-20
updated_at: 2026-07-20
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: tool-use
- inferred topics: agent-evaluation, coding-agent, reasoning, tool-use
- arXiv categories: cs.CL
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.16057
@@ -0,0 +1,60 @@
# Paper: When Do Multi-Agent Systems Help? An Information Bottleneck Perspective
---
type: paper
title: When Do Multi-Agent Systems Help? An Information Bottleneck Perspective
authors: Wendi Yu, Lianhao Zhou, Xiangjue Dong, Sai Sudarshan Barath, Declan Staunton, Byung-Jun Yoon, Xiaoning Qian, James Caverlee, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.16133
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-17
updated_at: 2026-07-17
status: queued
relevance: high
topics:
- agent-evaluation
- multi-agent
- reasoning
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: multi-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: multi-agent-llm
- inferred topics: agent-evaluation, multi-agent, reasoning
- arXiv categories: cs.LG
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.16133
@@ -0,0 +1,63 @@
# Paper: Just A Rather Very Intelligent Spoken Agent
---
type: paper
title: Just A Rather Very Intelligent Spoken Agent
authors: Chen Chen, Zhehuai Chen
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.16610
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-17
updated_at: 2026-07-17
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- multi-agent
- planning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: agentic-ai, ai-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agentic-ai, ai-agent
- inferred topics: agent-evaluation, computer-use, multi-agent, planning, tool-use, workflow-agent
- arXiv categories: cs.AI
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.16610
@@ -0,0 +1,61 @@
# Paper: Model-Driven Discipline for Multi-Agent LLMs: Requirement-to-Verification Generation of Traceable System Models
---
type: paper
title: "Model-Driven Discipline for Multi-Agent LLMs: Requirement-to-Verification Generation of Traceable System Models"
authors: Ran Wei, Le Zhu, Haochi Wang, Ruizhe Yang, Jiapeng Guan, Siyuan Ji, Yuchen Hu, Zhe Jiang, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.16708
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-18
updated_at: 2026-07-18
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- multi-agent
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: multi-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: multi-agent-llm
- inferred topics: agent-evaluation, agent-safety, multi-agent, tool-use
- arXiv categories: cs.SE
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.16708
@@ -0,0 +1,62 @@
# Paper: AgentBrew: Lifelong Knowledge Brewing from Strong Teachers to Weak LLM Agents
---
type: paper
title: "AgentBrew: Lifelong Knowledge Brewing from Strong Teachers to Weak LLM Agents"
authors: Yangqin Jiang, Chao Huang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.16851
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-18
updated_at: 2026-07-18
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- computer-use
- memory
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: llm-agent, tool-use
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: llm-agent, tool-use
- inferred topics: agent-evaluation, coding-agent, computer-use, memory, tool-use
- arXiv categories: cs.AI
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.16851
@@ -0,0 +1,63 @@
# Paper: Environment-free Synthetic Data Generation for API-Calling Agents
---
type: paper
title: Environment-free Synthetic Data Generation for API-Calling Agents
authors: Seanie Lee, Sanjoy Chowdhury, Chao Jiang, Cheng-Yu Hsieh, Ting-Yao Hu, Alexander T Toshev, Oncel Tuzel, Raviteja Vemulapalli
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.16900
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-21
updated_at: 2026-07-21
status: queued
relevance: high
topics:
- agent-evaluation
- memory
- rag
- tool-use
- workflow-agent
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: llm-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: llm-agent
- inferred topics: agent-evaluation, memory, rag, tool-use, workflow-agent, world-model
- arXiv categories: cs.AI
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.16900
@@ -0,0 +1,62 @@
# Paper: EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding
---
type: paper
title: "EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding"
authors: Yaohan Yang, Minglei Shi, Borui Zhang, Jie Zhou, Jiwen Lu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.17050
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-18
updated_at: 2026-07-18
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- planning
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CV
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: agent-evaluation, web-gui-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-evaluation, web-gui-agent
- inferred topics: agent-evaluation, computer-use, planning, rag, tool-use
- arXiv categories: cs.CV
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.17050
@@ -0,0 +1,60 @@
# Paper: A Diagnostic Framework for AI Agent Behavior
---
type: paper
title: A Diagnostic Framework for AI Agent Behavior
authors: Xichen Zhang, Yingjie Zhang, Tianshu Sun
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.17149
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-19
updated_at: 2026-07-19
status: queued
relevance: high
topics:
- agent-evaluation
- memory
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: ai-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: ai-agent
- inferred topics: agent-evaluation, memory, tool-use
- arXiv categories: cs.AI
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.17149
@@ -0,0 +1,60 @@
# Paper: SAGA: Synthetic Agentic Graph Architecture for Temporal Benchmark Generation
---
type: paper
title: "SAGA: Synthetic Agentic Graph Architecture for Temporal Benchmark Generation"
authors: Jiacheng Ding, Xiaofei Zhang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.17288
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-19
updated_at: 2026-07-19
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- rag
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.DB
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: llm-agent, rag-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: llm-agent, rag-agent
- inferred topics: agent-evaluation, agent-safety, rag
- arXiv categories: cs.DB
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.17288
@@ -0,0 +1,64 @@
# Paper: Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning
---
type: paper
title: "Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning"
authors: Zhihao Liu, Tianyu Wang, Xi Vincent Wang, Lihui Wang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.17331
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-19
updated_at: 2026-07-19
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- multi-agent
- planning
- tool-use
- workflow-agent
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 24
collection_queries: llm-agent, multi-agent-llm, planning-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: llm-agent, multi-agent-llm, planning-agent
- inferred topics: agent-evaluation, agent-safety, multi-agent, planning, tool-use, workflow-agent, world-model
- arXiv categories: cs.AI
- collection score: 24
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.17331
@@ -0,0 +1,63 @@
# Paper: Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions
---
type: paper
title: Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions
authors: Chen Xia, Zexi Kuang, Yuqing Hu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.17437
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-19
updated_at: 2026-07-19
status: queued
relevance: high
topics:
- agent-evaluation
- memory
- planning
- rag
- reasoning
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 19
collection_queries: llm-agent, planning-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: llm-agent, planning-agent
- inferred topics: agent-evaluation, memory, planning, rag, reasoning, world-model
- arXiv categories: cs.AI
- collection score: 19
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.17437
@@ -0,0 +1,62 @@
# Paper: Sidekick: Designing Communication for Effective Multitasking with Computer Use Agents
---
type: paper
title: "Sidekick: Designing Communication for Effective Multitasking with Computer Use Agents"
authors: Ruei-Che Chang, Wenqian Xu, Dingzeyu Li, Bryan Wang, Anhong Guo
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.17527
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-20
updated_at: 2026-07-20
status: queued
relevance: high
topics:
- computer-use
- multi-agent
- planning
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.HC
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: web-gui-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: web-gui-agent
- inferred topics: computer-use, multi-agent, planning, reasoning, tool-use
- arXiv categories: cs.HC
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.17527
@@ -0,0 +1,62 @@
# Paper: Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows
---
type: paper
title: Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows
authors: Jinyuan Deng, Zhengrui Chen, Xufeng Wei, Tianyu Xing, Chenyi Wen, Qi Sun, Cheng Zhuo
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.17528
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-22
updated_at: 2026-07-22
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- planning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 23
collection_queries: ai-agent, coding-agent, llm-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: ai-agent, coding-agent, llm-agent
- inferred topics: agent-evaluation, coding-agent, planning, tool-use, workflow-agent
- arXiv categories: cs.AI
- collection score: 23
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.17528
@@ -0,0 +1,62 @@
# Paper: Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory
---
type: paper
title: Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory
authors: Qingcan Kang, Mingyang Liu, Shixiong Kai, Kaichao Liang, Zhentao Tang, Yuqi Cui, Tao Zhong, Mingxuan Yuan
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.17545
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-20
updated_at: 2026-07-20
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- memory
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 20
collection_queries: agent-memory, language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-memory, language-agent
- inferred topics: agent-evaluation, agent-safety, memory, rag, tool-use
- arXiv categories: cs.AI
- collection score: 20
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.17545
@@ -0,0 +1,62 @@
# Paper: Mechanistic Attention Guidance for Agent Memory Refinement
---
type: paper
title: Mechanistic Attention Guidance for Agent Memory Refinement
authors: Yechao Hong, Haiquan Qiu, Yaqing Wang, Quanming Yao
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.17621
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-20
updated_at: 2026-07-20
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- memory
- rag
- reasoning
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: agent-memory
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-memory
- inferred topics: agent-evaluation, computer-use, memory, rag, reasoning
- arXiv categories: cs.AI
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.17621
@@ -0,0 +1,62 @@
# Paper: Exploratory and Assimilating Reflection: Reflective Recall Cycle for Long-term Memory
---
type: paper
title: "Exploratory and Assimilating Reflection: Reflective Recall Cycle for Long-term Memory"
authors: Ganesh Senrayan, Moyuru Yamada, Ishan Jindal, Kiran Purohit
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2607.17879
code_url:
source: arxiv
collected_at: 2026-07-27
published_at: 2026-07-20
updated_at: 2026-07-20
status: queued
relevance: high
topics:
- agent-evaluation
- memory
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: autonomous-agent-llm
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: autonomous-agent-llm
- inferred topics: agent-evaluation, memory, rag, reasoning, tool-use
- arXiv categories: cs.AI
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2607.17879

Some files were not shown because too many files have changed in this diff Show More