Expand arXiv paper corpus

This commit is contained in:
wuyang
2026-07-08 12:25:30 +08:00
parent b4af5e1fbc
commit 391b3a4732
980 changed files with 189464 additions and 58 deletions
@@ -0,0 +1,61 @@
# Paper: MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models
---
type: paper
title: "MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models"
authors: Hafsteinn Einarsson
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2507.20395
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-07-27
updated_at: 2025-07-27
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- embodied-agent
- reasoning
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, computer-use, embodied-agent, reasoning
- arXiv categories: cs.AI
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2507.20395
@@ -0,0 +1,63 @@
# Paper: MIMII-Agent: Leveraging LLMs with Function Calling for Relative Evaluation of Anomalous Sound Detection
---
type: paper
title: "MIMII-Agent: Leveraging LLMs with Function Calling for Relative Evaluation of Anomalous Sound Detection"
authors: Harsh Purohit, Tomoya Nishida, Kota Dohi, Takashi Endo, Yohei Kawaguchi
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2507.20666
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-07-28
updated_at: 2025-07-28
status: queued
relevance: high
topics:
- agent-evaluation
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- eess.AS
- cs.AI
- cs.LG
- cs.SD
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, rag, tool-use
- arXiv categories: eess.AS, cs.AI, cs.LG, cs.SD
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2507.20666
@@ -0,0 +1,60 @@
# Paper: MCPToolBench++: A Large Scale AI Agent Model Context Protocol MCP Tool Use Benchmark
---
type: paper
title: "MCPToolBench++: A Large Scale AI Agent Model Context Protocol MCP Tool Use Benchmark"
authors: Shiqing Fan, Xichen Ding, Liang Zhang, Linjian Mo
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2508.07575
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-08-11
updated_at: 2025-08-11
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 21
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, computer-use, tool-use
- arXiv categories: cs.AI
- collection score: 21
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2508.07575
@@ -0,0 +1,61 @@
# Paper: Hell or High Water: Evaluating Agentic Recovery from External Failures
---
type: paper
title: "Hell or High Water: Evaluating Agentic Recovery from External Failures"
authors: Andrew Wang, Sophia Hager, Adi Asija, Daniel Khashabi, Nicholas Andrews
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2508.11027
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-08-14
updated_at: 2025-08-14
status: queued
relevance: high
topics:
- agent-evaluation
- planning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, planning, tool-use, workflow-agent
- arXiv categories: cs.CL
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2508.11027
@@ -0,0 +1,61 @@
# Paper: ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
---
type: paper
title: "ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction"
authors: Xingshan Zeng, Weiwen Liu, Lingzhi Wang, Liangyou Li, Fei Mi, Yasheng Wang, Lifeng Shang, Xin Jiang, et al.
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2508.12685
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-08-18
updated_at: 2026-02-13
status: queued
relevance: high
topics:
- tool-use
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
- cs.AI
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: tool-use, world-model
- arXiv categories: cs.CL, cs.AI, cs.LG
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2508.12685
@@ -0,0 +1,63 @@
# Paper: PowerChain: A Verifiable Agentic AI System for Automating Distribution Grid Analyses
---
type: paper
title: "PowerChain: A Verifiable Agentic AI System for Automating Distribution Grid Analyses"
authors: Emmanuel O. Badmus, Peng Sang, Dimitrios Stamoulis, Amritanshu Pandey
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2508.17094
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-08-23
updated_at: 2025-10-21
status: queued
relevance: high
topics:
- planning
- rag
- reasoning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- eess.SY
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: planning, rag, reasoning, tool-use, workflow-agent
- arXiv categories: cs.AI, eess.SY
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2508.17094
@@ -0,0 +1,66 @@
# Paper: AppCopilot: Toward General, Accurate, Long-Horizon, and Efficient Mobile Agent
---
type: paper
title: "AppCopilot: Toward General, Accurate, Long-Horizon, and Efficient Mobile Agent"
authors: Jingru Fan, Yufan Dang, Jingyao Wu, Huatao Li, Runde Yang, Xiyuan Yang, Yuheng Wang, Chen Qian
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2509.02444
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-09-02
updated_at: 2025-10-17
status: queued
relevance: high
topics:
- computer-use
- memory
- multi-agent
- planning
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.CL
- cs.CV
- cs.HC
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: computer-use, memory, multi-agent, planning, reasoning, tool-use
- arXiv categories: cs.AI, cs.CL, cs.CV, cs.HC
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2509.02444
@@ -0,0 +1,60 @@
# Paper: GridMind: LLMs-Powered Agents for Power System Analysis and Operations
---
type: paper
title: "GridMind: LLMs-Powered Agents for Power System Analysis and Operations"
authors: Hongwei Jin, Kibaek Kim, Jonghwan Kwon
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2509.02494
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-09-02
updated_at: 2025-09-02
status: queued
relevance: high
topics:
- agent-evaluation
- multi-agent
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, multi-agent, workflow-agent
- arXiv categories: cs.AI
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2509.02494
@@ -0,0 +1,62 @@
# Paper: GeoJSON Agents:A Multi-Agent LLM Architecture for Geospatial Analysis-Function Calling vs Code Generation
---
type: paper
title: "GeoJSON Agents:A Multi-Agent LLM Architecture for Geospatial Analysis-Function Calling vs Code Generation"
authors: Qianqian Luo, Qingming Lin, Liuchang Xu, Sensen Wu, Ruichen Mao, Chao Wang, Hailin Feng, Bo Huang, et al.
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2509.08863
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-09-10
updated_at: 2025-12-03
status: queued
relevance: high
topics:
- agent-evaluation
- multi-agent
- planning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, multi-agent, planning, tool-use, workflow-agent
- arXiv categories: cs.SE
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2509.08863
@@ -0,0 +1,64 @@
# Paper: AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
---
type: paper
title: "AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise"
authors: Tara Bogavelli, Roshnee Sharma, Hari Subramani
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2509.10769
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-09-13
updated_at: 2026-01-06
status: queued
relevance: high
topics:
- agent-evaluation
- memory
- multi-agent
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.CL
- cs.MA
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, memory, multi-agent, tool-use, workflow-agent
- arXiv categories: cs.AI, cs.CL, cs.MA
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2509.10769
@@ -0,0 +1,59 @@
# Paper: Towards General Agentic Intelligence via Environment Scaling
---
type: paper
title: Towards General Agentic Intelligence via Environment Scaling
authors: Runnan Fang, Shihao Cai, Baixuan Li, Jialong Wu, Guangyu Li, Wenbiao Yin, Xinyu Wang, Xiaobin Wang, et al.
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2509.13311
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-09-16
updated_at: 2025-09-16
status: queued
relevance: high
topics:
- agent-evaluation
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, tool-use
- arXiv categories: cs.CL
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2509.13311
@@ -0,0 +1,60 @@
# Paper: Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
---
type: paper
title: "Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation"
authors: Thales Sales Almeida, João Guilherme Alves Santos, Thiago Laitz, Giovana Kerche Bonás
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2509.14477
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-09-17
updated_at: 2025-09-17
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- reasoning
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, computer-use, reasoning
- arXiv categories: cs.CL
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2509.14477
@@ -0,0 +1,61 @@
# Paper: CORE: Full-Path Evaluation of LLM Agents Beyond Final State
---
type: paper
title: "CORE: Full-Path Evaluation of LLM Agents Beyond Final State"
authors: Panagiotis Michelakis, Yiannis Hadjiyiannis, Dimitrios Stamoulis
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2509.20998
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-09-25
updated_at: 2025-09-25
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- tool-use
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 20
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, agent-safety, tool-use, world-model
- arXiv categories: cs.AI
- collection score: 20
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2509.20998
@@ -0,0 +1,61 @@
# Paper: Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling
---
type: paper
title: "Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling"
authors: Seiji Maekawa, Jackson Hassell, Pouya Pezeshkpour, Tom Mitchell, Estevam Hruschka
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2509.26553
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-09-30
updated_at: 2026-02-06
status: queued
relevance: high
topics:
- agent-evaluation
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
- cs.PL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, reasoning, tool-use
- arXiv categories: cs.CL, cs.PL
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2509.26553
@@ -0,0 +1,65 @@
# Paper: Small Language Models for Agentic Systems: A Survey of Architectures, Capabilities, and Deployment Trade offs
---
type: paper
title: "Small Language Models for Agentic Systems: A Survey of Architectures, Capabilities, and Deployment Trade offs"
authors: Raghav Sharma, Manan Mehta
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2510.03847
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-10-04
updated_at: 2025-10-04
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- computer-use
- planning
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 19
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, coding-agent, computer-use, planning, rag, reasoning, tool-use
- arXiv categories: cs.AI, cs.LG
- collection score: 19
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2510.03847
@@ -0,0 +1,59 @@
# Paper: AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
---
type: paper
title: "AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework"
authors: Hanchen Zhang, Xiao Liu, Bowen Lv, Xueqiao Sun, Bohao Jing, Iat Long Iong, Zhenyu Hou, Zehan Qi, et al.
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2510.04206
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-10-05
updated_at: 2025-10-05
status: queued
relevance: high
topics:
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: rag, tool-use
- arXiv categories: cs.AI
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2510.04206
@@ -0,0 +1,61 @@
# Paper: LLM Agents Beyond Utility: An Open-Ended Perspective
---
type: paper
title: "LLM Agents Beyond Utility: An Open-Ended Perspective"
authors: Asen Nachkov, Xi Wang, Luc Van Gool
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2510.14548
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-10-16
updated_at: 2025-10-16
status: queued
relevance: high
topics:
- memory
- planning
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: memory, planning, reasoning, tool-use
- arXiv categories: cs.AI
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2510.14548
@@ -0,0 +1,61 @@
# Paper: TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
---
type: paper
title: "TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications"
authors: Zhuohang Bian, Feiyang Wu, Zhuoran Li, Teng Ma, Youwei Zhuo
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2510.18586
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-10-21
updated_at: 2026-05-20
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- memory
- multi-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.DC
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, computer-use, memory, multi-agent
- arXiv categories: cs.DC
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2510.18586
@@ -0,0 +1,61 @@
# Paper: EU-Agent-Bench: Measuring Illegal Behavior of LLM Agents Under EU Law
---
type: paper
title: "EU-Agent-Bench: Measuring Illegal Behavior of LLM Agents Under EU Law"
authors: Ilija Lichkovski, Alexander Müller, Mariam Ibrahim, Tiwai Mhundwa
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2510.21524
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-10-24
updated_at: 2025-10-24
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, agent-safety, rag, tool-use
- arXiv categories: cs.AI
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2510.21524
@@ -0,0 +1,62 @@
# Paper: Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion
---
type: paper
title: Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion
authors: Haoyi Qiu, Yilun Zhou, Pranav Narayanan Venkit, Kung-Hsiang Huang, Jiaxin Zhang, Nanyun Peng, Chien-Sheng Wu
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2510.22768
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-10-26
updated_at: 2026-06-02
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- multi-agent
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, agent-safety, multi-agent, rag, tool-use
- arXiv categories: cs.CL
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2510.22768
@@ -0,0 +1,61 @@
# Paper: FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use
---
type: paper
title: "FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use"
authors: Zengzhuang Xu, Bingguang Hao, Zechuan Wang, Yuntao Wen, Xinyi Xu, Yang Liu, Long Chen, Dong Wang, et al.
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2510.24645
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-10-28
updated_at: 2025-11-16
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- multi-agent
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, computer-use, multi-agent, tool-use
- arXiv categories: cs.AI
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2510.24645
@@ -0,0 +1,60 @@
# Paper: ToolRM: Towards Agentic Tool-Use Reward Modeling
---
type: paper
title: "ToolRM: Towards Agentic Tool-Use Reward Modeling"
authors: Renhao Li, Jianhong Tu, Yang Su, Yantao Liu, Fei Huang, Hamid Alinejad-Rokny, Derek F. Wong, Junyang Lin, et al.
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2510.26167
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-10-30
updated_at: 2026-01-13
status: queued
relevance: high
topics:
- agent-evaluation
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, tool-use
- arXiv categories: cs.AI, cs.CL
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2510.26167
@@ -0,0 +1,63 @@
# Paper: Test-Time Adaptation for LLM Agents via Environment Interaction
---
type: paper
title: Test-Time Adaptation for LLM Agents via Environment Interaction
authors: Arthur Chen, Zuxin Liu, Jianguo Zhang, Akshara Prabhakar, Zhiwei Liu, Shelby Heinecke, Silvio Savarese, Victor Zhong, et al.
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2511.04847
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-11-06
updated_at: 2026-02-22
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- embodied-agent
- rag
- tool-use
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, agent-safety, embodied-agent, rag, tool-use, world-model
- arXiv categories: cs.LG
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2511.04847
@@ -0,0 +1,63 @@
# Paper: Refine and Align: Confidence Calibration through Multi-Agent Interaction in VQA
---
type: paper
title: "Refine and Align: Confidence Calibration through Multi-Agent Interaction in VQA"
authors: Ayush Pandey, Jai Bardhan, Ishita Jain, Ramya S Hebbalaguppe, Rohan Raju Dhanakshirur, Lovekesh Vig
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2511.11169
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-11-14
updated_at: 2025-11-14
status: queued
relevance: high
topics:
- agent-evaluation
- embodied-agent
- multi-agent
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CV
- cs.AI
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, embodied-agent, multi-agent, tool-use
- arXiv categories: cs.CV, cs.AI, cs.LG
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2511.11169
@@ -0,0 +1,62 @@
# Paper: Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
---
type: paper
title: Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
authors: Zimo Ji, Xunguang Wang, Zongjie Li, Pingchuan Ma, Yudong Gao, Daoyuan Wu, Xincheng Yan, Tian Tian, et al.
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2511.15203
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-11-19
updated_at: 2025-11-19
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, agent-safety, rag, tool-use
- arXiv categories: cs.CR, cs.AI
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2511.15203
@@ -0,0 +1,60 @@
# Paper: TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge Devices
---
type: paper
title: "TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge Devices"
authors: Mohd Ariful Haque, Fahad Rahman, Kishor Datta Gupta, Khalil Shujaee, Roy George
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2511.22138
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-11-27
updated_at: 2025-11-27
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, agent-safety, tool-use
- arXiv categories: cs.LG
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2511.22138
@@ -0,0 +1,65 @@
# Paper: IACT: A Self-Organizing Recursive Model for General AI Agents: A Technical White Paper on the Architecture Behind kragent.ai
---
type: paper
title: "IACT: A Self-Organizing Recursive Model for General AI Agents: A Technical White Paper on the Architecture Behind kragent.ai"
authors: Pengju Lu
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2512.02605
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-12-02
updated_at: 2025-12-02
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- memory
- rag
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.MA
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, computer-use, memory, rag, tool-use, workflow-agent
- arXiv categories: cs.AI, cs.MA, cs.SE
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2512.02605
@@ -0,0 +1,65 @@
# Paper: MedAI: Evaluating TxAgent's Therapeutic Agentic Reasoning in the NeurIPS CURE-Bench Competition
---
type: paper
title: "MedAI: Evaluating TxAgent's Therapeutic Agentic Reasoning in the NeurIPS CURE-Bench Competition"
authors: Tim Cofala, Christian Kalfar, Jingge Xiao, Johanna Schrader, Michelle Tang, Wolfgang Nejdl
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2512.11682
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-12-12
updated_at: 2026-06-15
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- computer-use
- planning
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, agent-safety, computer-use, planning, rag, reasoning, tool-use
- arXiv categories: cs.AI, cs.LG
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2512.11682
@@ -0,0 +1,62 @@
# Paper: Close the Loop: Synthesizing Infinite Tool-Use Data via Multi-Agent Role-Playing
---
type: paper
title: "Close the Loop: Synthesizing Infinite Tool-Use Data via Multi-Agent Role-Playing"
authors: Yuwen Li, Wei Zhang, Zelong Huang, Mason Yang, Jiajun Wu, Shawn Guo, Huahao Hu, Lingyi Sun, et al.
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2512.23611
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-12-29
updated_at: 2025-12-29
status: queued
relevance: high
topics:
- agent-evaluation
- multi-agent
- rag
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, multi-agent, rag, tool-use, workflow-agent
- arXiv categories: cs.CL
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2512.23611
@@ -0,0 +1,65 @@
# Paper: Nested Browser-Use Learning for Agentic Information Seeking
---
type: paper
title: Nested Browser-Use Learning for Agentic Information Seeking
authors: Baixuan Li, Jialong Wu, Wenbiao Yin, Kuan Li, Zhongwang Zhang, Huifeng Yin, Zhengwei Tao, Liwen Zhang, et al.
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2512.23647
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-12-29
updated_at: 2025-12-29
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
- cs.AI
- cs.IR
- cs.MA
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, computer-use, rag, reasoning, tool-use
- arXiv categories: cs.CL, cs.AI, cs.IR, cs.MA
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2512.23647
@@ -0,0 +1,64 @@
# Paper: State-of-the-art Small Language Coder Model: Mify-Coder
---
type: paper
title: "State-of-the-art Small Language Coder Model: Mify-Coder"
authors: Abhinav Parmar, Abhisek Panigrahi, Abhishek Kumar Dwivedi, Abhishek Bhattacharya, Adarsh Ramachandra, Aditya Choudhary, Aditya Garg, Aditya Raj, et al.
year: 2025
venue: arXiv
url: https://arxiv.org/abs/2512.23747
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2025-12-26
updated_at: 2025-12-26
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- coding-agent
- computer-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
- cs.AI
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, agent-safety, coding-agent, computer-use, workflow-agent
- arXiv categories: cs.SE, cs.AI, cs.CL
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2512.23747
@@ -0,0 +1,60 @@
# Paper: Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
---
type: paper
title: "Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity"
authors: Doyoung Kim, Zhiwei Ren, Jie Hao, Zhongkai Sun, Lichao Wang, Xiyao Ma, Zack Ye, Xu Han, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2601.00268
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-01-01
updated_at: 2026-01-01
status: queued
relevance: high
topics:
- agent-evaluation
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, tool-use
- arXiv categories: cs.CL, cs.AI
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2601.00268
@@ -0,0 +1,64 @@
# Paper: STELP: Secure Transpilation and Execution of LLM-Generated Programs
---
type: paper
title: "STELP: Secure Transpilation and Execution of LLM-Generated Programs"
authors: Swapnil Shinde, Sahil Wadhwa, Andy Luo, Akshay Gupta, Mohammad Shahed Sorower
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2601.05467
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-01-09
updated_at: 2026-01-15
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- multi-agent
- planning
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, agent-safety, multi-agent, planning, reasoning, tool-use
- arXiv categories: cs.SE, cs.AI
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2601.05467
@@ -0,0 +1,61 @@
# Paper: Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks
---
type: paper
title: "Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks"
authors: Elias Lumer, Faheem Nizar, Akshaya Jangiti, Kevin Frank, Anmol Gulati, Mandar Phadate, Vamse Kumar Subbiah
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2601.06007
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-01-09
updated_at: 2026-01-31
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- planning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, computer-use, planning, tool-use
- arXiv categories: cs.CL
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2601.06007
@@ -0,0 +1,60 @@
# Paper: CEDAR: Context Engineering for Agentic Data Science
---
type: paper
title: "CEDAR: Context Engineering for Agentic Data Science"
authors: Rishiraj Saha Roy, Chris Hinze, Luzian Hahn, Fabian Kuech
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2601.06606
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-01-10
updated_at: 2026-04-22
status: queued
relevance: high
topics:
- planning
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: planning, workflow-agent
- arXiv categories: cs.LG, cs.AI
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2601.06606
@@ -0,0 +1,62 @@
# Paper: PaperGuide: Making Small Language-Model Paper-Reading Agents More Efficient
---
type: paper
title: "PaperGuide: Making Small Language-Model Paper-Reading Agents More Efficient"
authors: Zijian Wang, Tiancheng Huang, Hanqi Li, Da Ma, Lu Chen, Kai Yu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2601.12988
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-01-19
updated_at: 2026-01-19
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- planning
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, computer-use, planning, reasoning, tool-use
- arXiv categories: cs.LG
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2601.12988
@@ -0,0 +1,63 @@
# Paper: MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
---
type: paper
title: "MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks"
authors: Zixuan Ke, Yifei Ming, Austin Xu, Ryan Chin, Xuan-Phi Nguyen, Prathyusha Jwalapuram, Jiayu Wang, Semih Yavuz, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2601.14652
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-01-21
updated_at: 2026-05-21
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- multi-agent
- reasoning
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.CL
- cs.MA
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, computer-use, multi-agent, reasoning
- arXiv categories: cs.AI, cs.CL, cs.MA
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2601.14652
@@ -0,0 +1,61 @@
# Paper: AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
---
type: paper
title: "AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?"
authors: Hao Li, Ruoyao Wen, Shanghao Shi, Ning Zhang, Yevgeniy Vorobeychik, Chaowei Xiao
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.03117
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-03
updated_at: 2026-05-07
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- planning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, planning, tool-use
- arXiv categories: cs.CR
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.03117
@@ -0,0 +1,63 @@
# Paper: TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking
---
type: paper
title: "TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking"
authors: Yu Cheng, Yongkang Hu, Jiuan Zhou, Yushuo Zhang, Yihang Chen, Huichi Zhou, Mingang Chen, Zhizhong Zhang, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.03224
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-03
updated_at: 2026-06-06
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- computer-use
- memory
- reasoning
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, computer-use, memory, reasoning
- arXiv categories: cs.AI, cs.LG
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.03224
@@ -0,0 +1,63 @@
# Paper: AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration
---
type: paper
title: "AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration"
authors: Jianhao Ruan, Zhihao Xu, Yiran Peng, Fashen Ren, Zhaoyang Yu, Xinbing Liang, Jinyu Xiang, Yongru Chen, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.03786
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-03
updated_at: 2026-02-07
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- planning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, coding-agent, planning, tool-use, workflow-agent
- arXiv categories: cs.AI, cs.CL
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.03786
@@ -0,0 +1,61 @@
# Paper: SocialVeil: Probing Social Intelligence of Language Agents under Communication Barriers
---
type: paper
title: "SocialVeil: Probing Social Intelligence of Language Agents under Communication Barriers"
authors: Keyang Xuan, Pengda Wang, Chongrui Ye, Haofei Yu, Tal August, Jiaxuan You
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.05115
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-04
updated_at: 2026-02-04
status: queued
relevance: high
topics:
- agent-evaluation
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, rag, tool-use
- arXiv categories: cs.AI, cs.CL
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.05115
@@ -0,0 +1,61 @@
# Paper: PieArena: Ranking and Profiling Language Agents in Realistic Negotiation Scenarios
---
type: paper
title: "PieArena: Ranking and Profiling Language Agents in Realistic Negotiation Scenarios"
authors: Chris Zhu, Sasha Cui, Will Sanok Dufallo, Runzhi Jin, Zhen Xu, Linjun Zhang, Daylian Cain
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.05302
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-05
updated_at: 2026-06-01
status: queued
relevance: high
topics:
- agent-evaluation
- multi-agent
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, multi-agent, reasoning, tool-use
- arXiv categories: cs.AI
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.05302
@@ -0,0 +1,62 @@
# Paper: Spider-Sense: Intrinsic Risk Sensing for Efficient Agent Defense with Hierarchical Adaptive Screening
---
type: paper
title: "Spider-Sense: Intrinsic Risk Sensing for Efficient Agent Defense with Hierarchical Adaptive Screening"
authors: Zhenxiong Yu, Zhi Yang, Zhiheng Jin, Shuhe Wang, Heng Zhang, Yanlin Fei, Lingfeng Zeng, Fangqi Lou, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.05386
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-05
updated_at: 2026-02-06
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, reasoning, tool-use
- arXiv categories: cs.CR, cs.AI
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.05386
@@ -0,0 +1,60 @@
# Paper: NAAMSE: Framework for Evolutionary Security Evaluation of Agents
---
type: paper
title: "NAAMSE: Framework for Evolutionary Security Evaluation of Agents"
authors: Kunal Pai, Parth Shah, Harshil Patel
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.07391
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-07
updated_at: 2026-03-08
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.MA
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety
- arXiv categories: cs.AI, cs.MA
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.07391
@@ -0,0 +1,64 @@
# Paper: Agent-Fence: Mapping Security Vulnerabilities Across Deep Research Agents
---
type: paper
title: "Agent-Fence: Mapping Security Vulnerabilities Across Deep Research Agents"
authors: Sai Puppala, Ismail Hossain, Md Jahangir Alam, Yoonpyo Lee, Jay Yoo, Tanzim Ahad, Syed Bahauddin Alam, Sajedul Talukder
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.07652
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-07
updated_at: 2026-02-07
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- memory
- planning
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, memory, planning, rag, tool-use
- arXiv categories: cs.CR, cs.AI
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.07652
@@ -0,0 +1,61 @@
# Paper: LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth
---
type: paper
title: "LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth"
authors: Weihao Zeng, Yuzhen Huang, Junxian He
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.07962
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-08
updated_at: 2026-02-08
status: queued
relevance: high
topics:
- agent-evaluation
- planning
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, planning, rag, tool-use
- arXiv categories: cs.AI
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.07962
@@ -0,0 +1,62 @@
# Paper: Spectral Guardrails for Agents in the Wild: Detecting Tool Use Hallucinations via Attention Topology
---
type: paper
title: "Spectral Guardrails for Agents in the Wild: Detecting Tool Use Hallucinations via Attention Topology"
authors: Valentin Noël
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.08082
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-08
updated_at: 2026-02-08
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
- cs.AI
- eess.SP
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, tool-use
- arXiv categories: cs.LG, cs.AI, eess.SP
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.08082
@@ -0,0 +1,63 @@
# Paper: From Assistant to Double Agent: Formalizing and Benchmarking Attacks on OpenClaw for Personalized Local AI Agent
---
type: paper
title: "From Assistant to Double Agent: Formalizing and Benchmarking Attacks on OpenClaw for Personalized Local AI Agent"
authors: Yuhang Wang, Feiming Xu, Zheng Lin, Guangyu He, Yuzhe Huang, Haichang Gao, Zhenxing Niu, Shiguo Lian, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.08412
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-09
updated_at: 2026-02-11
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- memory
- planning
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 21
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, memory, planning, rag, tool-use
- arXiv categories: cs.AI
- collection score: 21
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.08412
@@ -0,0 +1,62 @@
# Paper: AgentTrace: A Structured Logging Framework for Agent System Observability
---
type: paper
title: "AgentTrace: A Structured Logging Framework for Agent System Observability"
authors: Adam AlSayyad, Kelvin Yuxiang Huang, Richik Pal
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.10133
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-07
updated_at: 2026-02-07
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, reasoning, tool-use
- arXiv categories: cs.SE, cs.AI
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.10133
@@ -0,0 +1,61 @@
# Paper: AIR: Improving Agent Safety through Incident Response
---
type: paper
title: "AIR: Improving Agent Safety through Incident Response"
authors: Zibo Xiao, Jun Sun, Junjie Chen
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.11749
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-12
updated_at: 2026-06-20
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- computer-use
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, computer-use, tool-use
- arXiv categories: cs.AI
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.11749
@@ -0,0 +1,65 @@
# Paper: Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
---
type: paper
title: "Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents"
authors: Xu Li, Simon Yu, Minzhou Pan, Yiyou Sun, Bo Li, Dawn Song, Xue Lin, Weiyan Shi
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.13379
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-13
updated_at: 2026-06-10
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
- cs.AI
- cs.CL
- cs.LG
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, rag, tool-use
- arXiv categories: cs.CR, cs.AI, cs.CL, cs.LG, cs.SE
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.13379
@@ -0,0 +1,62 @@
# Paper: REMem: Reasoning with Episodic Memory in Language Agent
---
type: paper
title: "REMem: Reasoning with Episodic Memory in Language Agent"
authors: Yiheng Shu, Saisri Padmaja Jonnalagedda, Xiang Gao, Bernal Jiménez Gutiérrez, Weijian Qi, Kamalika Das, Huan Sun, Yu Su
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.13530
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-13
updated_at: 2026-02-28
status: queued
relevance: high
topics:
- agent-evaluation
- memory
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 19
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, memory, rag, reasoning, tool-use
- arXiv categories: cs.AI
- collection score: 19
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.13530
@@ -0,0 +1,59 @@
# Paper: HyFunc: Accelerating LLM-based Function Calls for Agentic AI through Hybrid-Model Cascade and Dynamic Templating
---
type: paper
title: "HyFunc: Accelerating LLM-based Function Calls for Agentic AI through Hybrid-Model Cascade and Dynamic Templating"
authors: Weibin Liao, Jian-guang Lou, Haoyi Xiong
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.13665
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-14
updated_at: 2026-02-14
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, computer-use
- arXiv categories: cs.AI
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.13665
@@ -0,0 +1,62 @@
# Paper: REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agents
---
type: paper
title: "REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agents"
authors: Zheng Chu, Xiao Wang, Jack Hong, Huiming Fan, Yuqi Huang, Yue Yang, Guohai Xu, Chenxiao Zhao, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.14234
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-15
updated_at: 2026-02-15
status: queued
relevance: high
topics:
- agent-evaluation
- planning
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, planning, rag, tool-use
- arXiv categories: cs.AI, cs.CL
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.14234
@@ -0,0 +1,62 @@
# Paper: MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in Model Context Protocol Agents
---
type: paper
title: "MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in Model Context Protocol Agents"
authors: Zhenhong Zhou, Yuanhe Zhang, Hongwei Cai, Moayad Aloqaily, Ouns Bouachir, Linsey Pang, Prakhar Mehrotra, Kun Wang, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.14281
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-15
updated_at: 2026-02-24
status: queued
relevance: high
topics:
- agent-safety
- computer-use
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-safety, computer-use, reasoning, tool-use
- arXiv categories: cs.CR, cs.CL
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.14281
@@ -0,0 +1,59 @@
# Paper: Narrow Fine-Tuning Erodes Safety Alignment in Vision-Language Agents
---
type: paper
title: Narrow Fine-Tuning Erodes Safety Alignment in Vision-Language Agents
authors: Idhant Gulati, Shivam Raval
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.16931
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-18
updated_at: 2026-03-15
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, agent-safety
- arXiv categories: cs.AI
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.16931
@@ -0,0 +1,61 @@
# Paper: Beyond single-channel agentic benchmarking
---
type: paper
title: Beyond single-channel agentic benchmarking
authors: Nelu D. Radpour
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.18456
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-05
updated_at: 2026-02-05
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CY
- cs.AI
- cs.HC
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety
- arXiv categories: cs.CY, cs.AI, cs.HC
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.18456
@@ -0,0 +1,62 @@
# Paper: Capable but Unreliable: Canonical Path Deviation as a Causal Mechanism of Agent Failure in Long-Horizon Tasks
---
type: paper
title: "Capable but Unreliable: Canonical Path Deviation as a Causal Mechanism of Agent Failure in Long-Horizon Tasks"
authors: Wilson Y. Lee
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.19008
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-22
updated_at: 2026-02-22
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- planning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, computer-use, planning, tool-use
- arXiv categories: cs.CL, cs.LG
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.19008
@@ -0,0 +1,63 @@
# Paper: "Are You Sure?": An Empirical Study of Human Perception Vulnerability in LLM-Driven Agentic Systems
---
type: paper
title: "\"Are You Sure?\": An Empirical Study of Human Perception Vulnerability in LLM-Driven Agentic Systems"
authors: Xinfeng Li, Shenyu Dai, Kelong Zheng, Yue Xiao, Gelei Deng, Wei Dong, Xiaofeng Wang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.21127
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-24
updated_at: 2026-02-24
status: queued
relevance: high
topics:
- agent-safety
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.HC
- cs.AI
- cs.CR
- cs.SI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-safety, tool-use, workflow-agent
- arXiv categories: cs.HC, cs.AI, cs.CR, cs.SI
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.21127
@@ -0,0 +1,61 @@
# Paper: ParamMem: Augmenting Language Agents with Parametric Reflective Memory
---
type: paper
title: "ParamMem: Augmenting Language Agents with Parametric Reflective Memory"
authors: Tianjun Yao, Yongqiang Chen, Yujia Zheng, Pan Li, Zhiqiang Shen, Kun Zhang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2602.23320
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-26
updated_at: 2026-02-27
status: queued
relevance: high
topics:
- agent-evaluation
- memory
- reasoning
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
- cs.MA
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, memory, reasoning
- arXiv categories: cs.LG, cs.MA
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2602.23320
@@ -0,0 +1,61 @@
# Paper: Thought Virus: Viral Misalignment via Subliminal Prompting in Multi-Agent Systems
---
type: paper
title: "Thought Virus: Viral Misalignment via Subliminal Prompting in Multi-Agent Systems"
authors: Moritz Weckbecker, Jonas Müller, Ben Hagag, Michael Mulet
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.00131
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-23
updated_at: 2026-02-23
status: queued
relevance: high
topics:
- agent-safety
- multi-agent
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.MA
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-safety, multi-agent, tool-use
- arXiv categories: cs.MA, cs.AI
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.00131
@@ -0,0 +1,63 @@
# Paper: TraceSIR: A Multi-Agent Framework for Structured Analysis and Reporting of Agentic Execution Traces
---
type: paper
title: "TraceSIR: A Multi-Agent Framework for Structured Analysis and Reporting of Agentic Execution Traces"
authors: Shu-Xun Yang, Cunxiang Wang, Haoke Zhang, Wenbo Yu, Lindong Wu, Jiayi Gui, Dayong Yang, Yukuo Cen, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.00623
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-28
updated_at: 2026-02-28
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- multi-agent
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 19
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, coding-agent, multi-agent, reasoning, tool-use
- arXiv categories: cs.AI, cs.CL
- collection score: 19
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.00623
@@ -0,0 +1,62 @@
# Paper: The Synthetic Web: Adversarially-Curated Mini-Internets for Diagnosing Epistemic Weaknesses of Language Agents
---
type: paper
title: "The Synthetic Web: Adversarially-Curated Mini-Internets for Diagnosing Epistemic Weaknesses of Language Agents"
authors: Shrey Shah, Levent Ozgur
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.00801
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-28
updated_at: 2026-02-28
status: queued
relevance: high
topics:
- agent-evaluation
- embodied-agent
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.IR
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, embodied-agent, rag, tool-use
- arXiv categories: cs.AI, cs.IR
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.00801
@@ -0,0 +1,64 @@
# Paper: Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents
---
type: paper
title: Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents
authors: Yuxin Liu, Mingye Zhu, Siyuan Liu, Bo Hu, Lei Zhang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.01438
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-02
updated_at: 2026-03-02
status: queued
relevance: high
topics:
- agent-safety
- coding-agent
- computer-use
- planning
- rag
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-safety, coding-agent, computer-use, planning, rag, world-model
- arXiv categories: cs.CL, cs.AI
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.01438
@@ -0,0 +1,61 @@
# Paper: FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents
---
type: paper
title: "FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents"
authors: Qizheng Li, Yifei Zhang, Xiao Yang, Xu Yang, Zhuo Wang, Weiqing Liu, Jiang Bian
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.01712
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-02
updated_at: 2026-05-20
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- planning
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, coding-agent, planning
- arXiv categories: cs.AI, cs.LG
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.01712
@@ -0,0 +1,60 @@
# Paper: A Natural Language Agentic Approach to Study Affective Polarization
---
type: paper
title: A Natural Language Agentic Approach to Study Affective Polarization
authors: Stephanie Anneris Malvicini, Ewelina Gajewska, Arda Derbent, Katarzyna Budzynska, Jarosław A. Chudziak, Maria Vanina Martinez
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.02711
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-03
updated_at: 2026-03-03
status: queued
relevance: high
topics:
- multi-agent
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: multi-agent, rag, tool-use
- arXiv categories: cs.AI
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.02711
@@ -0,0 +1,63 @@
# Paper: The Controllability Trap: A Governance Framework for Military AI Agents
---
type: paper
title: "The Controllability Trap: A Governance Framework for Military AI Agents"
authors: Subramanyam Sahoo
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.03515
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-03
updated_at: 2026-03-03
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- planning
- tool-use
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CY
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, planning, tool-use, world-model
- arXiv categories: cs.CY, cs.AI
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.03515
@@ -0,0 +1,61 @@
# Paper: MAGE: Meta-Reinforcement Learning for Language Agents toward Strategic Exploration and Exploitation
---
type: paper
title: "MAGE: Meta-Reinforcement Learning for Language Agents toward Strategic Exploration and Exploitation"
authors: Lu Yang, Zelai Xu, Minyang Xie, Jiaxuan Gao, Zhao Shok, Yu Wang, Yi Wu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.03680
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-04
updated_at: 2026-03-04
status: queued
relevance: high
topics:
- memory
- multi-agent
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: memory, multi-agent, reasoning, tool-use
- arXiv categories: cs.AI
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.03680
@@ -0,0 +1,63 @@
# Paper: EigenData: A Self-Evolving Multi-Agent Platform for Function-Calling Data Synthesis, Auditing, and Repair
---
type: paper
title: "EigenData: A Self-Evolving Multi-Agent Platform for Function-Calling Data Synthesis, Auditing, and Repair"
authors: Jiaao Chen, Jingyuan Qi, Mingye Gao, Wei-Chen Wang, Hanrui Wang, Di Jin
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.05553
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-05
updated_at: 2026-03-05
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- multi-agent
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
- cs.AI
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, coding-agent, multi-agent, tool-use
- arXiv categories: cs.SE, cs.AI, cs.CL
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.05553
@@ -0,0 +1,61 @@
# Paper: Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
---
type: paper
title: "Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent"
authors: Bowei Xia, Mengkang Hu, Shijian Wang, Jiarui Jin, Wenxiang Jiao, Yuan Lu, Kexin Li, Ping Luo
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.05578
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-05
updated_at: 2026-03-05
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, computer-use, tool-use
- arXiv categories: cs.SE, cs.AI
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.05578
@@ -0,0 +1,64 @@
# Paper: From Thinker to Society: Security in Hierarchical Autonomy Evolution of AI Agents
---
type: paper
title: "From Thinker to Society: Security in Hierarchical Autonomy Evolution of AI Agents"
authors: Xiaolei Zhang, Lu Zhou, Xiaogang Xu, Jiafei Wu, Tianyu Du, Heqing Huang, Hao Peng, Zhe Liu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.07496
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-08
updated_at: 2026-03-21
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- computer-use
- multi-agent
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, computer-use, multi-agent, reasoning, tool-use
- arXiv categories: cs.CR, cs.AI
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.07496
@@ -0,0 +1,62 @@
# Paper: AgentRaft: Automated Detection of Data Over-Exposure in LLM Agents
---
type: paper
title: "AgentRaft: Automated Detection of Data Over-Exposure in LLM Agents"
authors: Yixi Lin, Jiangrong Wu, Yuhong Nan, Xueqiang Wang, Xinyuan Zhang, Zibin Zheng
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.07557
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-08
updated_at: 2026-03-08
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, agent-safety, rag, reasoning, tool-use
- arXiv categories: cs.SE
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.07557
@@ -0,0 +1,63 @@
# Paper: \$OneMillion-Bench: How Far are Language Agents from Human Experts?
---
type: paper
title: "\\$OneMillion-Bench: How Far are Language Agents from Human Experts?"
authors: Qianyu Yang, Yang Liu, Jiaqi Li, Jun Bai, Hao Chen, Kaiyuan Chen, Tiliang Duan, Jiayun Dong, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.07980
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-09
updated_at: 2026-03-09
status: queued
relevance: high
topics:
- agent-evaluation
- planning
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
- cs.AI
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, planning, reasoning, tool-use
- arXiv categories: cs.LG, cs.AI, cs.CL
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.07980
@@ -0,0 +1,62 @@
# Paper: KernelCraft: Benchmarking for Agentic Close-to-Metal Kernel Generation on Emerging Hardware
---
type: paper
title: "KernelCraft: Benchmarking for Agentic Close-to-Metal Kernel Generation on Emerging Hardware"
authors: Jiayi Nie, Haoran Wu, Yao Lai, Zeyu Cao, Cheng Zhang, Binglei Lou, Erwei Wang, Jianyi Cheng, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.08721
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-10
updated_at: 2026-05-29
status: queued
relevance: high
topics:
- agent-evaluation
- reasoning
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AR
- cs.LG
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, reasoning, workflow-agent
- arXiv categories: cs.AR, cs.LG, cs.SE
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.08721
@@ -0,0 +1,66 @@
# Paper: Security Considerations for Multi-agent Systems
---
type: paper
title: Security Considerations for Multi-agent Systems
authors: Tam Nguyen, Moses Ndebugre, Dheeraj Arremsetty
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.09002
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-09
updated_at: 2026-04-26
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- computer-use
- memory
- multi-agent
- planning
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, computer-use, memory, multi-agent, planning, rag, tool-use
- arXiv categories: cs.CR, cs.AI
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.09002
@@ -0,0 +1,63 @@
# Paper: Human-AI Co-reasoning for Clinical Diagnosis with Evidence-Integrated Language Agent
---
type: paper
title: Human-AI Co-reasoning for Clinical Diagnosis with Evidence-Integrated Language Agent
authors: Zhongzhen Huang, Yan Ling, Hong Chen, Ye Feng, Li Wu, Linjie Mu, Shaoting Zhang, Xiaofan Zhang, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.10492
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-11
updated_at: 2026-03-18
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- multi-agent
- rag
- reasoning
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, agent-safety, multi-agent, rag, reasoning, workflow-agent
- arXiv categories: cs.CL
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.10492
@@ -0,0 +1,61 @@
# Paper: The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
---
type: paper
title: "The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey"
authors: Juhee Kim, Xiaoyuan Liu, Zhun Wang, Shi Qiu, Bo Li, Wenbo Guo, Dawn Song
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.11088
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-11
updated_at: 2026-03-11
status: queued
relevance: high
topics:
- agent-safety
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-safety, tool-use, workflow-agent
- arXiv categories: cs.CR, cs.AI
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.11088
@@ -0,0 +1,61 @@
# Paper: QUARE: Quality-Aware Requirements Analysis through Multi-Agent Dialectical Negotiation
---
type: paper
title: "QUARE: Quality-Aware Requirements Analysis through Multi-Agent Dialectical Negotiation"
authors: Haowei Cheng, Milhan Kim, Foutse Khomh, Teeradaj Racharak, Nobukazu Yoshioka, Naoyasu Ubayashi, Hironori Washizaki
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.11890
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-12
updated_at: 2026-06-05
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- multi-agent
- rag
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, multi-agent, rag
- arXiv categories: cs.SE
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.11890
@@ -0,0 +1,61 @@
# Paper: CCTU: A Benchmark for Tool Use under Complex Constraints
---
type: paper
title: "CCTU: A Benchmark for Tool Use under Complex Constraints"
authors: Junjie Ye, Guoqiang Zhang, Wenjie Fu, Tao Gui, Qi Zhang, Xuanjing Huang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.15309
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-16
updated_at: 2026-03-16
status: queued
relevance: high
topics:
- agent-evaluation
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, rag, tool-use
- arXiv categories: cs.CL, cs.AI
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.15309
@@ -0,0 +1,59 @@
# Paper: Compiled Memory: Not More Information, but More Precise Instructions for Language Agents
---
type: paper
title: "Compiled Memory: Not More Information, but More Precise Instructions for Language Agents"
authors: James Rhodes, George Kang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.15666
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-12
updated_at: 2026-03-12
status: queued
relevance: high
topics:
- memory
- rag
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: memory, rag
- arXiv categories: cs.AI
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.15666
@@ -0,0 +1,62 @@
# Paper: Differential Harm Propensity in Personalized LLM Agents: The Curious Case of Mental Health Disclosure
---
type: paper
title: "Differential Harm Propensity in Personalized LLM Agents: The Curious Case of Mental Health Disclosure"
authors: Caglar Yildirim
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.16734
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-17
updated_at: 2026-03-17
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- memory
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 20
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, memory, rag, tool-use
- arXiv categories: cs.AI
- collection score: 20
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.16734
@@ -0,0 +1,61 @@
# Paper: Agentic Cognitive Profiling: Realigning Automated Alzheimer's Disease Detection with Clinical Construct Validity
---
type: paper
title: "Agentic Cognitive Profiling: Realigning Automated Alzheimer's Disease Detection with Clinical Construct Validity"
authors: Jiawen Kang, Kun Li, Dongrui Han, Jinchao Li, Junan Li, Lingwei Meng, Xixin Wu, Helen Meng
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.17392
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-18
updated_at: 2026-03-18
status: queued
relevance: high
topics:
- agent-evaluation
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.MA
- cs.IR
- q-bio.NC
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: agent-evaluation, tool-use
- arXiv categories: cs.MA, cs.IR, q-bio.NC
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.17392
@@ -0,0 +1,63 @@
# Paper: Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
---
type: paper
title: Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
authors: Xuan Chen, Lu Yan, Ruqi Zhang, Xiangyu Zhang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.18245
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-18
updated_at: 2026-03-18
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- rag
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.SE
- cs.CR
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 18
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, rag, tool-use, workflow-agent
- arXiv categories: cs.SE, cs.CR
- collection score: 18
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.18245
@@ -0,0 +1,61 @@
# Paper: A Framework for Formalizing LLM Agent Security
---
type: paper
title: A Framework for Formalizing LLM Agent Security
authors: Vincent Siu, Jingxuan He, Kyle Montgomery, Zhun Wang, Neil Gong, Chenguang Wang, Dawn Song
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.19469
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-19
updated_at: 2026-03-19
status: queued
relevance: high
topics:
- agent-safety
- memory
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-safety, memory, tool-use
- arXiv categories: cs.CR, cs.AI
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.19469
@@ -0,0 +1,62 @@
# Paper: TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents
---
type: paper
title: "TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents"
authors: Shaojie Zhuang, Lu Yin, Guangshun Wei, Yunpeng Li, Xilu Wang, Yuanfeng Zhou
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.19684
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-20
updated_at: 2026-06-23
status: queued
relevance: high
topics:
- agent-evaluation
- coding-agent
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CV
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, coding-agent, rag, reasoning, tool-use
- arXiv categories: cs.CV
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.19684
@@ -0,0 +1,62 @@
# Paper: AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling
---
type: paper
title: "AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling"
authors: Liang Ding
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.21357
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-22
updated_at: 2026-05-10
status: queued
relevance: high
topics:
- computer-use
- memory
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: computer-use, memory, tool-use, workflow-agent
- arXiv categories: cs.AI, cs.CL
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.21357
@@ -0,0 +1,64 @@
# Paper: Toward a Theory of Hierarchical Memory for Language Agents
---
type: paper
title: Toward a Theory of Hierarchical Memory for Language Agents
authors: Yashar Talebirad, Ali Parsaee, Csongor Y. Szepesvari, Amirhossein Nadiri, Osmar Zaiane
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.21564
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-23
updated_at: 2026-03-23
status: queued
relevance: high
topics:
- agent-evaluation
- memory
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.IR
- cs.AI
- cs.IT
- cs.SI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, memory, rag, tool-use
- arXiv categories: cs.IR, cs.AI, cs.IT, cs.SI
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.21564
@@ -0,0 +1,60 @@
# Paper: Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning
---
type: paper
title: Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning
authors: Tommaso Galliena, Stefano Rosa, Tommaso Apicella, Pietro Morerio, Alessio Del Bue, Lorenzo Natale
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.24257
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-25
updated_at: 2026-03-30
status: queued
relevance: high
topics:
- agent-evaluation
- embodied-agent
- memory
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CV
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, embodied-agent, memory
- arXiv categories: cs.CV
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.24257
@@ -0,0 +1,62 @@
# Paper: SafeGuard ASF: SR Agentic Humanoid Robot System for Autonomous Industrial Safety
---
type: paper
title: "SafeGuard ASF: SR Agentic Humanoid Robot System for Autonomous Industrial Safety"
authors: Thanh Nguyen Canh, Thang Tran Viet, Thanh Tuan Tran, Ben Wei Lim
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.25353
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-26
updated_at: 2026-03-26
status: queued
relevance: high
topics:
- agent-safety
- embodied-agent
- reasoning
- tool-use
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.RO
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-safety, embodied-agent, reasoning, tool-use, world-model
- arXiv categories: cs.RO
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.25353
@@ -0,0 +1,60 @@
# Paper: SafetyDrift: Predicting When AI Agents Cross the Line Before They Actually Do
---
type: paper
title: "SafetyDrift: Predicting When AI Agents Cross the Line Before They Actually Do"
authors: Aditya Dhodapkar, Farhaan Pishori
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.27148
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-28
updated_at: 2026-03-28
status: queued
relevance: high
topics:
- agent-safety
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-safety, tool-use
- arXiv categories: cs.CR, cs.AI
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.27148
@@ -0,0 +1,63 @@
# Paper: Evaluating Privilege Usage of Agents with Real-World Tools
---
type: paper
title: Evaluating Privilege Usage of Agents with Real-World Tools
authors: Quan Zhang, Lianhang Fu, Lvsi Lian, Gwihwan Go, Yujue Wang, Chijin Zhou, Yu Jiang, Geguang Pu
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.28166
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-30
updated_at: 2026-04-20
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- rag
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, rag, tool-use, workflow-agent
- arXiv categories: cs.CR, cs.AI
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.28166
@@ -0,0 +1,64 @@
# Paper: Synergy: A Next-Generation General-Purpose Agent for Open Agentic Web
---
type: paper
title: "Synergy: A Next-Generation General-Purpose Agent for Open Agentic Web"
authors: Xiaohang Nie, Zihan Guo, Kezhuo Yang, Zhichong Zheng, Bochen Ge, Shuai Pan, Zeyi Chen, Youling Xiang, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.28428
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-30
updated_at: 2026-03-30
status: queued
relevance: high
topics:
- coding-agent
- embodied-agent
- memory
- multi-agent
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CY
- cs.MA
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: function-calling
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling
- inferred topics: coding-agent, embodied-agent, memory, multi-agent, rag, tool-use
- arXiv categories: cs.CY, cs.MA
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.28428
@@ -0,0 +1,64 @@
# Paper: Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing
---
type: paper
title: Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing
authors: Alex Zongo, Filippos Fotiadis, Ufuk Topcu, Peng Wei
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2603.28900
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-03-30
updated_at: 2026-03-30
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- multi-agent
- world-model
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.RO
- cs.AI
- cs.LG
- eess.SY
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 13
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, multi-agent, world-model
- arXiv categories: cs.RO, cs.AI, cs.LG, eess.SY
- collection score: 13
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2603.28900
@@ -0,0 +1,62 @@
# Paper: ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
---
type: paper
title: "ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis"
authors: Yu Li, Haoyu Luo, Yuejin Xie, Yuqian Fu, Zhonghao Yang, Shuai Shao, Qihan Ren, Wanying Qu, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2604.02022
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-04-02
updated_at: 2026-05-13
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- planning
- rag
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 20
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, planning, rag, tool-use
- arXiv categories: cs.AI
- collection score: 20
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2604.02022
@@ -0,0 +1,61 @@
# Paper: Brief Is Better: Non-Monotonic Chain-of-Thought Budget Effects in Function-Calling Language Agents
---
type: paper
title: "Brief Is Better: Non-Monotonic Chain-of-Thought Budget Effects in Function-Calling Language Agents"
authors: Xuan Qi
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2604.02155
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-04-02
updated_at: 2026-04-02
status: queued
relevance: high
topics:
- agent-evaluation
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: function-calling, language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: function-calling, language-agent
- inferred topics: agent-evaluation, rag, reasoning, tool-use
- arXiv categories: cs.CL
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2604.02155
@@ -0,0 +1,63 @@
# Paper: Co-Evolution of Policy and Internal Reward for Language Agents
---
type: paper
title: Co-Evolution of Policy and Internal Reward for Language Agents
authors: Xinyu Wang, Hanwei Wu, Jingwei Song, Shuyuan Zhang, Jiayi Zhang, Fanqi Kong, Tung Sum Thomas Kwok, Xiao-Wen Chang, et al.
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2604.03098
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-04-03
updated_at: 2026-04-03
status: queued
relevance: high
topics:
- agent-evaluation
- computer-use
- planning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
- cs.AI
- cs.CL
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, computer-use, planning, tool-use
- arXiv categories: cs.LG, cs.AI, cs.CL
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2604.03098
@@ -0,0 +1,62 @@
# Paper: DRAFT: Task Decoupled Latent Reasoning for Agent Safety
---
type: paper
title: "DRAFT: Task Decoupled Latent Reasoning for Agent Safety"
authors: Lin Wang, Junfeng Fang, Dan Zhang, Fei Shen, Xiang Wang, Tat-Seng Chua
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2604.03242
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-02-11
updated_at: 2026-02-11
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- rag
- reasoning
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.LG
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 17
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, rag, reasoning, tool-use
- arXiv categories: cs.LG
- collection score: 17
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2604.03242
@@ -0,0 +1,62 @@
# Paper: Profile-Then-Reason: Bounded Semantic Complexity for Tool-Augmented Language Agents
---
type: paper
title: "Profile-Then-Reason: Bounded Semantic Complexity for Tool-Augmented Language Agents"
authors: Paulo Akira F. Enabe
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2604.04131
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-04-05
updated_at: 2026-04-05
status: queued
relevance: high
topics:
- agent-evaluation
- rag
- reasoning
- tool-use
- workflow-agent
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 16
collection_queries: language-agent
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: language-agent
- inferred topics: agent-evaluation, rag, reasoning, tool-use, workflow-agent
- arXiv categories: cs.AI
- collection score: 16
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2604.04131
@@ -0,0 +1,60 @@
# Paper: ShieldNet: Network-Level Guardrails against Emerging Supply-Chain Injections in Agentic Systems
---
type: paper
title: "ShieldNet: Network-Level Guardrails against Emerging Supply-Chain Injections in Agentic Systems"
authors: Zhuowen Yuan, Zhaorun Chen, Zhen Xiang, Nathaniel D. Bastian, Seyyed Hadi Hashemi, Chaowei Xiao, Wenbo Guo, Bo Li
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2604.04426
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-04-06
updated_at: 2026-04-06
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- tool-use
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.AI
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 15
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, tool-use
- arXiv categories: cs.AI
- collection score: 15
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2604.04426
@@ -0,0 +1,60 @@
# Paper: ARuleCon: Agentic Security Rule Conversion
---
type: paper
title: "ARuleCon: Agentic Security Rule Conversion"
authors: Ming Xu, Hongtai Wang, Yanpei Guo, Zhengmin Yu, Weili Han, Hoon Wei Lim, Jin Song Dong, Jiaheng Zhang
year: 2026
venue: arXiv
url: https://arxiv.org/abs/2604.06762
code_url:
source: arxiv
collected_at: 2026-07-08
published_at: 2026-04-08
updated_at: 2026-04-08
status: queued
relevance: high
topics:
- agent-evaluation
- agent-safety
- rag
methods:
-
benchmarks:
-
models:
-
datasets:
- cs.CR
related_concepts:
-
related_jobs:
-
related_experiments:
-
related_projects:
-
collection_score: 14
collection_queries: agent-safety
---
## One-line Takeaway
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
## Why Collected
- matched queries: agent-safety
- inferred topics: agent-evaluation, agent-safety, rag
- arXiv categories: cs.CR
- collection score: 14
## Review Checklist
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
- Should it be promoted from `queued` to `skimmed` or `summarized`?
## Links
- arXiv: https://arxiv.org/abs/2604.06762

Some files were not shown because too many files have changed in this diff Show More