Expand arXiv paper corpus
This commit is contained in:
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models"
|
||||
authors: Hafsteinn Einarsson
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2507.20395
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-07-27
|
||||
updated_at: 2025-07-27
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- embodied-agent
|
||||
- reasoning
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, computer-use, embodied-agent, reasoning
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2507.20395
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: MIMII-Agent: Leveraging LLMs with Function Calling for Relative Evaluation of Anomalous Sound Detection
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "MIMII-Agent: Leveraging LLMs with Function Calling for Relative Evaluation of Anomalous Sound Detection"
|
||||
authors: Harsh Purohit, Tomoya Nishida, Kota Dohi, Takashi Endo, Yohei Kawaguchi
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2507.20666
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-07-28
|
||||
updated_at: 2025-07-28
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- eess.AS
|
||||
- cs.AI
|
||||
- cs.LG
|
||||
- cs.SD
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, rag, tool-use
|
||||
- arXiv categories: eess.AS, cs.AI, cs.LG, cs.SD
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2507.20666
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: MCPToolBench++: A Large Scale AI Agent Model Context Protocol MCP Tool Use Benchmark
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "MCPToolBench++: A Large Scale AI Agent Model Context Protocol MCP Tool Use Benchmark"
|
||||
authors: Shiqing Fan, Xichen Ding, Liang Zhang, Linjian Mo
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2508.07575
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-08-11
|
||||
updated_at: 2025-08-11
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 21
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, computer-use, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 21
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2508.07575
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Hell or High Water: Evaluating Agentic Recovery from External Failures
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Hell or High Water: Evaluating Agentic Recovery from External Failures"
|
||||
authors: Andrew Wang, Sophia Hager, Adi Asija, Daniel Khashabi, Nicholas Andrews
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2508.11027
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-08-14
|
||||
updated_at: 2025-08-14
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- planning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, planning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2508.11027
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction"
|
||||
authors: Xingshan Zeng, Weiwen Liu, Lingzhi Wang, Liangyou Li, Fei Mi, Yasheng Wang, Lifeng Shang, Xin Jiang, et al.
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2508.12685
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-08-18
|
||||
updated_at: 2026-02-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- tool-use
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
- cs.AI
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: tool-use, world-model
|
||||
- arXiv categories: cs.CL, cs.AI, cs.LG
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2508.12685
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: PowerChain: A Verifiable Agentic AI System for Automating Distribution Grid Analyses
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "PowerChain: A Verifiable Agentic AI System for Automating Distribution Grid Analyses"
|
||||
authors: Emmanuel O. Badmus, Peng Sang, Dimitrios Stamoulis, Amritanshu Pandey
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2508.17094
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-08-23
|
||||
updated_at: 2025-10-21
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- planning
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- eess.SY
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: planning, rag, reasoning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI, eess.SY
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2508.17094
|
||||
+66
@@ -0,0 +1,66 @@
|
||||
# Paper: AppCopilot: Toward General, Accurate, Long-Horizon, and Efficient Mobile Agent
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "AppCopilot: Toward General, Accurate, Long-Horizon, and Efficient Mobile Agent"
|
||||
authors: Jingru Fan, Yufan Dang, Jingyao Wu, Huatao Li, Runde Yang, Xiyuan Yang, Yuheng Wang, Chen Qian
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2509.02444
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-09-02
|
||||
updated_at: 2025-10-17
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- computer-use
|
||||
- memory
|
||||
- multi-agent
|
||||
- planning
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.CL
|
||||
- cs.CV
|
||||
- cs.HC
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: computer-use, memory, multi-agent, planning, reasoning, tool-use
|
||||
- arXiv categories: cs.AI, cs.CL, cs.CV, cs.HC
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2509.02444
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: GridMind: LLMs-Powered Agents for Power System Analysis and Operations
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "GridMind: LLMs-Powered Agents for Power System Analysis and Operations"
|
||||
authors: Hongwei Jin, Kibaek Kim, Jonghwan Kwon
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2509.02494
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-09-02
|
||||
updated_at: 2025-09-02
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- multi-agent
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, multi-agent, workflow-agent
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2509.02494
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: GeoJSON Agents:A Multi-Agent LLM Architecture for Geospatial Analysis-Function Calling vs Code Generation
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "GeoJSON Agents:A Multi-Agent LLM Architecture for Geospatial Analysis-Function Calling vs Code Generation"
|
||||
authors: Qianqian Luo, Qingming Lin, Liuchang Xu, Sensen Wu, Ruichen Mao, Chao Wang, Hailin Feng, Bo Huang, et al.
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2509.08863
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-09-10
|
||||
updated_at: 2025-12-03
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- multi-agent
|
||||
- planning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, multi-agent, planning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.SE
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2509.08863
|
||||
+64
@@ -0,0 +1,64 @@
|
||||
# Paper: AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise"
|
||||
authors: Tara Bogavelli, Roshnee Sharma, Hari Subramani
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2509.10769
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-09-13
|
||||
updated_at: 2026-01-06
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- memory
|
||||
- multi-agent
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.CL
|
||||
- cs.MA
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, memory, multi-agent, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI, cs.CL, cs.MA
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2509.10769
|
||||
+59
@@ -0,0 +1,59 @@
|
||||
# Paper: Towards General Agentic Intelligence via Environment Scaling
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Towards General Agentic Intelligence via Environment Scaling
|
||||
authors: Runnan Fang, Shihao Cai, Baixuan Li, Jialong Wu, Guangyu Li, Wenbiao Yin, Xinyu Wang, Xiaobin Wang, et al.
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2509.13311
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-09-16
|
||||
updated_at: 2025-09-16
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, tool-use
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2509.13311
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation"
|
||||
authors: Thales Sales Almeida, João Guilherme Alves Santos, Thiago Laitz, Giovana Kerche Bonás
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2509.14477
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-09-17
|
||||
updated_at: 2025-09-17
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- reasoning
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, computer-use, reasoning
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2509.14477
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: CORE: Full-Path Evaluation of LLM Agents Beyond Final State
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "CORE: Full-Path Evaluation of LLM Agents Beyond Final State"
|
||||
authors: Panagiotis Michelakis, Yiannis Hadjiyiannis, Dimitrios Stamoulis
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2509.20998
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-09-25
|
||||
updated_at: 2025-09-25
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- tool-use
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 20
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, agent-safety, tool-use, world-model
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 20
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2509.20998
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling"
|
||||
authors: Seiji Maekawa, Jackson Hassell, Pouya Pezeshkpour, Tom Mitchell, Estevam Hruschka
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2509.26553
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-09-30
|
||||
updated_at: 2026-02-06
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
- cs.PL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, reasoning, tool-use
|
||||
- arXiv categories: cs.CL, cs.PL
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2509.26553
|
||||
+65
@@ -0,0 +1,65 @@
|
||||
# Paper: Small Language Models for Agentic Systems: A Survey of Architectures, Capabilities, and Deployment Trade offs
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Small Language Models for Agentic Systems: A Survey of Architectures, Capabilities, and Deployment Trade offs"
|
||||
authors: Raghav Sharma, Manan Mehta
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2510.03847
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-10-04
|
||||
updated_at: 2025-10-04
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- computer-use
|
||||
- planning
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 19
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, coding-agent, computer-use, planning, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.AI, cs.LG
|
||||
- collection score: 19
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2510.03847
|
||||
+59
@@ -0,0 +1,59 @@
|
||||
# Paper: AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework"
|
||||
authors: Hanchen Zhang, Xiao Liu, Bowen Lv, Xueqiao Sun, Bohao Jing, Iat Long Iong, Zhenyu Hou, Zehan Qi, et al.
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2510.04206
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-10-05
|
||||
updated_at: 2025-10-05
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: rag, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2510.04206
|
||||
@@ -0,0 +1,61 @@
|
||||
# Paper: LLM Agents Beyond Utility: An Open-Ended Perspective
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "LLM Agents Beyond Utility: An Open-Ended Perspective"
|
||||
authors: Asen Nachkov, Xi Wang, Luc Van Gool
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2510.14548
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-10-16
|
||||
updated_at: 2025-10-16
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- memory
|
||||
- planning
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: memory, planning, reasoning, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2510.14548
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications"
|
||||
authors: Zhuohang Bian, Feiyang Wu, Zhuoran Li, Teng Ma, Youwei Zhuo
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2510.18586
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-10-21
|
||||
updated_at: 2026-05-20
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- memory
|
||||
- multi-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.DC
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, computer-use, memory, multi-agent
|
||||
- arXiv categories: cs.DC
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2510.18586
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: EU-Agent-Bench: Measuring Illegal Behavior of LLM Agents Under EU Law
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "EU-Agent-Bench: Measuring Illegal Behavior of LLM Agents Under EU Law"
|
||||
authors: Ilija Lichkovski, Alexander Müller, Mariam Ibrahim, Tiwai Mhundwa
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2510.21524
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-10-24
|
||||
updated_at: 2025-10-24
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, agent-safety, rag, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2510.21524
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion
|
||||
authors: Haoyi Qiu, Yilun Zhou, Pranav Narayanan Venkit, Kung-Hsiang Huang, Jiaxin Zhang, Nanyun Peng, Chien-Sheng Wu
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2510.22768
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-10-26
|
||||
updated_at: 2026-06-02
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- multi-agent
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, agent-safety, multi-agent, rag, tool-use
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2510.22768
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use"
|
||||
authors: Zengzhuang Xu, Bingguang Hao, Zechuan Wang, Yuntao Wen, Xinyi Xu, Yang Liu, Long Chen, Dong Wang, et al.
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2510.24645
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-10-28
|
||||
updated_at: 2025-11-16
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- multi-agent
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, computer-use, multi-agent, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2510.24645
|
||||
@@ -0,0 +1,60 @@
|
||||
# Paper: ToolRM: Towards Agentic Tool-Use Reward Modeling
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "ToolRM: Towards Agentic Tool-Use Reward Modeling"
|
||||
authors: Renhao Li, Jianhong Tu, Yang Su, Yantao Liu, Fei Huang, Hamid Alinejad-Rokny, Derek F. Wong, Junyang Lin, et al.
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2510.26167
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-10-30
|
||||
updated_at: 2026-01-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, tool-use
|
||||
- arXiv categories: cs.AI, cs.CL
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2510.26167
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: Test-Time Adaptation for LLM Agents via Environment Interaction
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Test-Time Adaptation for LLM Agents via Environment Interaction
|
||||
authors: Arthur Chen, Zuxin Liu, Jianguo Zhang, Akshara Prabhakar, Zhiwei Liu, Shelby Heinecke, Silvio Savarese, Victor Zhong, et al.
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2511.04847
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-11-06
|
||||
updated_at: 2026-02-22
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- embodied-agent
|
||||
- rag
|
||||
- tool-use
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, agent-safety, embodied-agent, rag, tool-use, world-model
|
||||
- arXiv categories: cs.LG
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2511.04847
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: Refine and Align: Confidence Calibration through Multi-Agent Interaction in VQA
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Refine and Align: Confidence Calibration through Multi-Agent Interaction in VQA"
|
||||
authors: Ayush Pandey, Jai Bardhan, Ishita Jain, Ramya S Hebbalaguppe, Rohan Raju Dhanakshirur, Lovekesh Vig
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2511.11169
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-11-14
|
||||
updated_at: 2025-11-14
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- embodied-agent
|
||||
- multi-agent
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CV
|
||||
- cs.AI
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, embodied-agent, multi-agent, tool-use
|
||||
- arXiv categories: cs.CV, cs.AI, cs.LG
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2511.11169
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
|
||||
authors: Zimo Ji, Xunguang Wang, Zongjie Li, Pingchuan Ma, Yudong Gao, Daoyuan Wu, Xincheng Yan, Tian Tian, et al.
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2511.15203
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-11-19
|
||||
updated_at: 2025-11-19
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, agent-safety, rag, tool-use
|
||||
- arXiv categories: cs.CR, cs.AI
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2511.15203
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge Devices
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge Devices"
|
||||
authors: Mohd Ariful Haque, Fahad Rahman, Kishor Datta Gupta, Khalil Shujaee, Roy George
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2511.22138
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-11-27
|
||||
updated_at: 2025-11-27
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, agent-safety, tool-use
|
||||
- arXiv categories: cs.LG
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2511.22138
|
||||
+65
@@ -0,0 +1,65 @@
|
||||
# Paper: IACT: A Self-Organizing Recursive Model for General AI Agents: A Technical White Paper on the Architecture Behind kragent.ai
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "IACT: A Self-Organizing Recursive Model for General AI Agents: A Technical White Paper on the Architecture Behind kragent.ai"
|
||||
authors: Pengju Lu
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2512.02605
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-12-02
|
||||
updated_at: 2025-12-02
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- memory
|
||||
- rag
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.MA
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, computer-use, memory, rag, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI, cs.MA, cs.SE
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2512.02605
|
||||
+65
@@ -0,0 +1,65 @@
|
||||
# Paper: MedAI: Evaluating TxAgent's Therapeutic Agentic Reasoning in the NeurIPS CURE-Bench Competition
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "MedAI: Evaluating TxAgent's Therapeutic Agentic Reasoning in the NeurIPS CURE-Bench Competition"
|
||||
authors: Tim Cofala, Christian Kalfar, Jingge Xiao, Johanna Schrader, Michelle Tang, Wolfgang Nejdl
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2512.11682
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-12-12
|
||||
updated_at: 2026-06-15
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- computer-use
|
||||
- planning
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, agent-safety, computer-use, planning, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.AI, cs.LG
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2512.11682
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Close the Loop: Synthesizing Infinite Tool-Use Data via Multi-Agent Role-Playing
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Close the Loop: Synthesizing Infinite Tool-Use Data via Multi-Agent Role-Playing"
|
||||
authors: Yuwen Li, Wei Zhang, Zelong Huang, Mason Yang, Jiajun Wu, Shawn Guo, Huahao Hu, Lingyi Sun, et al.
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2512.23611
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-12-29
|
||||
updated_at: 2025-12-29
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- multi-agent
|
||||
- rag
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, multi-agent, rag, tool-use, workflow-agent
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2512.23611
|
||||
+65
@@ -0,0 +1,65 @@
|
||||
# Paper: Nested Browser-Use Learning for Agentic Information Seeking
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Nested Browser-Use Learning for Agentic Information Seeking
|
||||
authors: Baixuan Li, Jialong Wu, Wenbiao Yin, Kuan Li, Zhongwang Zhang, Huifeng Yin, Zhengwei Tao, Liwen Zhang, et al.
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2512.23647
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-12-29
|
||||
updated_at: 2025-12-29
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
- cs.AI
|
||||
- cs.IR
|
||||
- cs.MA
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, computer-use, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.CL, cs.AI, cs.IR, cs.MA
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2512.23647
|
||||
+64
@@ -0,0 +1,64 @@
|
||||
# Paper: State-of-the-art Small Language Coder Model: Mify-Coder
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "State-of-the-art Small Language Coder Model: Mify-Coder"
|
||||
authors: Abhinav Parmar, Abhisek Panigrahi, Abhishek Kumar Dwivedi, Abhishek Bhattacharya, Adarsh Ramachandra, Aditya Choudhary, Aditya Garg, Aditya Raj, et al.
|
||||
year: 2025
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2512.23747
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2025-12-26
|
||||
updated_at: 2025-12-26
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- coding-agent
|
||||
- computer-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
- cs.AI
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, agent-safety, coding-agent, computer-use, workflow-agent
|
||||
- arXiv categories: cs.SE, cs.AI, cs.CL
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2512.23747
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity"
|
||||
authors: Doyoung Kim, Zhiwei Ren, Jie Hao, Zhongkai Sun, Lichao Wang, Xiyao Ma, Zack Ye, Xu Han, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2601.00268
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-01-01
|
||||
updated_at: 2026-01-01
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, tool-use
|
||||
- arXiv categories: cs.CL, cs.AI
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2601.00268
|
||||
+64
@@ -0,0 +1,64 @@
|
||||
# Paper: STELP: Secure Transpilation and Execution of LLM-Generated Programs
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "STELP: Secure Transpilation and Execution of LLM-Generated Programs"
|
||||
authors: Swapnil Shinde, Sahil Wadhwa, Andy Luo, Akshay Gupta, Mohammad Shahed Sorower
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2601.05467
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-01-09
|
||||
updated_at: 2026-01-15
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- multi-agent
|
||||
- planning
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, agent-safety, multi-agent, planning, reasoning, tool-use
|
||||
- arXiv categories: cs.SE, cs.AI
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2601.05467
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks"
|
||||
authors: Elias Lumer, Faheem Nizar, Akshaya Jangiti, Kevin Frank, Anmol Gulati, Mandar Phadate, Vamse Kumar Subbiah
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2601.06007
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-01-09
|
||||
updated_at: 2026-01-31
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- planning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, computer-use, planning, tool-use
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2601.06007
|
||||
@@ -0,0 +1,60 @@
|
||||
# Paper: CEDAR: Context Engineering for Agentic Data Science
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "CEDAR: Context Engineering for Agentic Data Science"
|
||||
authors: Rishiraj Saha Roy, Chris Hinze, Luzian Hahn, Fabian Kuech
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2601.06606
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-01-10
|
||||
updated_at: 2026-04-22
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- planning
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: planning, workflow-agent
|
||||
- arXiv categories: cs.LG, cs.AI
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2601.06606
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: PaperGuide: Making Small Language-Model Paper-Reading Agents More Efficient
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "PaperGuide: Making Small Language-Model Paper-Reading Agents More Efficient"
|
||||
authors: Zijian Wang, Tiancheng Huang, Hanqi Li, Da Ma, Lu Chen, Kai Yu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2601.12988
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-01-19
|
||||
updated_at: 2026-01-19
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- planning
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, computer-use, planning, reasoning, tool-use
|
||||
- arXiv categories: cs.LG
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2601.12988
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks"
|
||||
authors: Zixuan Ke, Yifei Ming, Austin Xu, Ryan Chin, Xuan-Phi Nguyen, Prathyusha Jwalapuram, Jiayu Wang, Semih Yavuz, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2601.14652
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-01-21
|
||||
updated_at: 2026-05-21
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- multi-agent
|
||||
- reasoning
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.CL
|
||||
- cs.MA
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, computer-use, multi-agent, reasoning
|
||||
- arXiv categories: cs.AI, cs.CL, cs.MA
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2601.14652
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?"
|
||||
authors: Hao Li, Ruoyao Wen, Shanghao Shi, Ning Zhang, Yevgeniy Vorobeychik, Chaowei Xiao
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.03117
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-03
|
||||
updated_at: 2026-05-07
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- planning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, planning, tool-use
|
||||
- arXiv categories: cs.CR
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.03117
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking"
|
||||
authors: Yu Cheng, Yongkang Hu, Jiuan Zhou, Yushuo Zhang, Yihang Chen, Huichi Zhou, Mingang Chen, Zhizhong Zhang, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.03224
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-03
|
||||
updated_at: 2026-06-06
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- computer-use
|
||||
- memory
|
||||
- reasoning
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, computer-use, memory, reasoning
|
||||
- arXiv categories: cs.AI, cs.LG
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.03224
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration"
|
||||
authors: Jianhao Ruan, Zhihao Xu, Yiran Peng, Fashen Ren, Zhaoyang Yu, Xinbing Liang, Jinyu Xiang, Yongru Chen, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.03786
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-03
|
||||
updated_at: 2026-02-07
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- planning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, coding-agent, planning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI, cs.CL
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.03786
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: SocialVeil: Probing Social Intelligence of Language Agents under Communication Barriers
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "SocialVeil: Probing Social Intelligence of Language Agents under Communication Barriers"
|
||||
authors: Keyang Xuan, Pengda Wang, Chongrui Ye, Haofei Yu, Tal August, Jiaxuan You
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.05115
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-04
|
||||
updated_at: 2026-02-04
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, rag, tool-use
|
||||
- arXiv categories: cs.AI, cs.CL
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.05115
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: PieArena: Ranking and Profiling Language Agents in Realistic Negotiation Scenarios
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "PieArena: Ranking and Profiling Language Agents in Realistic Negotiation Scenarios"
|
||||
authors: Chris Zhu, Sasha Cui, Will Sanok Dufallo, Runzhi Jin, Zhen Xu, Linjun Zhang, Daylian Cain
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.05302
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-05
|
||||
updated_at: 2026-06-01
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- multi-agent
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, multi-agent, reasoning, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.05302
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Spider-Sense: Intrinsic Risk Sensing for Efficient Agent Defense with Hierarchical Adaptive Screening
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Spider-Sense: Intrinsic Risk Sensing for Efficient Agent Defense with Hierarchical Adaptive Screening"
|
||||
authors: Zhenxiong Yu, Zhi Yang, Zhiheng Jin, Shuhe Wang, Heng Zhang, Yanlin Fei, Lingfeng Zeng, Fangqi Lou, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.05386
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-05
|
||||
updated_at: 2026-02-06
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, reasoning, tool-use
|
||||
- arXiv categories: cs.CR, cs.AI
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.05386
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: NAAMSE: Framework for Evolutionary Security Evaluation of Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "NAAMSE: Framework for Evolutionary Security Evaluation of Agents"
|
||||
authors: Kunal Pai, Parth Shah, Harshil Patel
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.07391
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-07
|
||||
updated_at: 2026-03-08
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.MA
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety
|
||||
- arXiv categories: cs.AI, cs.MA
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.07391
|
||||
+64
@@ -0,0 +1,64 @@
|
||||
# Paper: Agent-Fence: Mapping Security Vulnerabilities Across Deep Research Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Agent-Fence: Mapping Security Vulnerabilities Across Deep Research Agents"
|
||||
authors: Sai Puppala, Ismail Hossain, Md Jahangir Alam, Yoonpyo Lee, Jay Yoo, Tanzim Ahad, Syed Bahauddin Alam, Sajedul Talukder
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.07652
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-07
|
||||
updated_at: 2026-02-07
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- memory
|
||||
- planning
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, memory, planning, rag, tool-use
|
||||
- arXiv categories: cs.CR, cs.AI
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.07652
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth"
|
||||
authors: Weihao Zeng, Yuzhen Huang, Junxian He
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.07962
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-08
|
||||
updated_at: 2026-02-08
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- planning
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, planning, rag, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.07962
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Spectral Guardrails for Agents in the Wild: Detecting Tool Use Hallucinations via Attention Topology
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Spectral Guardrails for Agents in the Wild: Detecting Tool Use Hallucinations via Attention Topology"
|
||||
authors: Valentin Noël
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.08082
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-08
|
||||
updated_at: 2026-02-08
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
- cs.AI
|
||||
- eess.SP
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, tool-use
|
||||
- arXiv categories: cs.LG, cs.AI, eess.SP
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.08082
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: From Assistant to Double Agent: Formalizing and Benchmarking Attacks on OpenClaw for Personalized Local AI Agent
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "From Assistant to Double Agent: Formalizing and Benchmarking Attacks on OpenClaw for Personalized Local AI Agent"
|
||||
authors: Yuhang Wang, Feiming Xu, Zheng Lin, Guangyu He, Yuzhe Huang, Haichang Gao, Zhenxing Niu, Shiguo Lian, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.08412
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-09
|
||||
updated_at: 2026-02-11
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- memory
|
||||
- planning
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 21
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, memory, planning, rag, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 21
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.08412
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: AgentTrace: A Structured Logging Framework for Agent System Observability
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "AgentTrace: A Structured Logging Framework for Agent System Observability"
|
||||
authors: Adam AlSayyad, Kelvin Yuxiang Huang, Richik Pal
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.10133
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-07
|
||||
updated_at: 2026-02-07
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, reasoning, tool-use
|
||||
- arXiv categories: cs.SE, cs.AI
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.10133
|
||||
@@ -0,0 +1,61 @@
|
||||
# Paper: AIR: Improving Agent Safety through Incident Response
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "AIR: Improving Agent Safety through Incident Response"
|
||||
authors: Zibo Xiao, Jun Sun, Junjie Chen
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.11749
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-12
|
||||
updated_at: 2026-06-20
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- computer-use
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, computer-use, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.11749
|
||||
+65
@@ -0,0 +1,65 @@
|
||||
# Paper: Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents"
|
||||
authors: Xu Li, Simon Yu, Minzhou Pan, Yiyou Sun, Bo Li, Dawn Song, Xue Lin, Weiyan Shi
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.13379
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-13
|
||||
updated_at: 2026-06-10
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
- cs.AI
|
||||
- cs.CL
|
||||
- cs.LG
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, rag, tool-use
|
||||
- arXiv categories: cs.CR, cs.AI, cs.CL, cs.LG, cs.SE
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.13379
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: REMem: Reasoning with Episodic Memory in Language Agent
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "REMem: Reasoning with Episodic Memory in Language Agent"
|
||||
authors: Yiheng Shu, Saisri Padmaja Jonnalagedda, Xiang Gao, Bernal Jiménez Gutiérrez, Weijian Qi, Kamalika Das, Huan Sun, Yu Su
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.13530
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-13
|
||||
updated_at: 2026-02-28
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- memory
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 19
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, memory, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 19
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.13530
|
||||
+59
@@ -0,0 +1,59 @@
|
||||
# Paper: HyFunc: Accelerating LLM-based Function Calls for Agentic AI through Hybrid-Model Cascade and Dynamic Templating
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "HyFunc: Accelerating LLM-based Function Calls for Agentic AI through Hybrid-Model Cascade and Dynamic Templating"
|
||||
authors: Weibin Liao, Jian-guang Lou, Haoyi Xiong
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.13665
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-14
|
||||
updated_at: 2026-02-14
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, computer-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.13665
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agents"
|
||||
authors: Zheng Chu, Xiao Wang, Jack Hong, Huiming Fan, Yuqi Huang, Yue Yang, Guohai Xu, Chenxiao Zhao, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.14234
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-15
|
||||
updated_at: 2026-02-15
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- planning
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, planning, rag, tool-use
|
||||
- arXiv categories: cs.AI, cs.CL
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.14234
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in Model Context Protocol Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in Model Context Protocol Agents"
|
||||
authors: Zhenhong Zhou, Yuanhe Zhang, Hongwei Cai, Moayad Aloqaily, Ouns Bouachir, Linsey Pang, Prakhar Mehrotra, Kun Wang, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.14281
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-15
|
||||
updated_at: 2026-02-24
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-safety
|
||||
- computer-use
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-safety, computer-use, reasoning, tool-use
|
||||
- arXiv categories: cs.CR, cs.CL
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.14281
|
||||
+59
@@ -0,0 +1,59 @@
|
||||
# Paper: Narrow Fine-Tuning Erodes Safety Alignment in Vision-Language Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Narrow Fine-Tuning Erodes Safety Alignment in Vision-Language Agents
|
||||
authors: Idhant Gulati, Shivam Raval
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.16931
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-18
|
||||
updated_at: 2026-03-15
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, agent-safety
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.16931
|
||||
@@ -0,0 +1,61 @@
|
||||
# Paper: Beyond single-channel agentic benchmarking
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Beyond single-channel agentic benchmarking
|
||||
authors: Nelu D. Radpour
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.18456
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-05
|
||||
updated_at: 2026-02-05
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CY
|
||||
- cs.AI
|
||||
- cs.HC
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety
|
||||
- arXiv categories: cs.CY, cs.AI, cs.HC
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.18456
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Capable but Unreliable: Canonical Path Deviation as a Causal Mechanism of Agent Failure in Long-Horizon Tasks
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Capable but Unreliable: Canonical Path Deviation as a Causal Mechanism of Agent Failure in Long-Horizon Tasks"
|
||||
authors: Wilson Y. Lee
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.19008
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-22
|
||||
updated_at: 2026-02-22
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- planning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, computer-use, planning, tool-use
|
||||
- arXiv categories: cs.CL, cs.LG
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.19008
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: "Are You Sure?": An Empirical Study of Human Perception Vulnerability in LLM-Driven Agentic Systems
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "\"Are You Sure?\": An Empirical Study of Human Perception Vulnerability in LLM-Driven Agentic Systems"
|
||||
authors: Xinfeng Li, Shenyu Dai, Kelong Zheng, Yue Xiao, Gelei Deng, Wei Dong, Xiaofeng Wang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.21127
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-24
|
||||
updated_at: 2026-02-24
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-safety
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.HC
|
||||
- cs.AI
|
||||
- cs.CR
|
||||
- cs.SI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-safety, tool-use, workflow-agent
|
||||
- arXiv categories: cs.HC, cs.AI, cs.CR, cs.SI
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.21127
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: ParamMem: Augmenting Language Agents with Parametric Reflective Memory
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "ParamMem: Augmenting Language Agents with Parametric Reflective Memory"
|
||||
authors: Tianjun Yao, Yongqiang Chen, Yujia Zheng, Pan Li, Zhiqiang Shen, Kun Zhang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2602.23320
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-26
|
||||
updated_at: 2026-02-27
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- memory
|
||||
- reasoning
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
- cs.MA
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, memory, reasoning
|
||||
- arXiv categories: cs.LG, cs.MA
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2602.23320
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Thought Virus: Viral Misalignment via Subliminal Prompting in Multi-Agent Systems
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Thought Virus: Viral Misalignment via Subliminal Prompting in Multi-Agent Systems"
|
||||
authors: Moritz Weckbecker, Jonas Müller, Ben Hagag, Michael Mulet
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.00131
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-23
|
||||
updated_at: 2026-02-23
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-safety
|
||||
- multi-agent
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.MA
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-safety, multi-agent, tool-use
|
||||
- arXiv categories: cs.MA, cs.AI
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.00131
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: TraceSIR: A Multi-Agent Framework for Structured Analysis and Reporting of Agentic Execution Traces
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "TraceSIR: A Multi-Agent Framework for Structured Analysis and Reporting of Agentic Execution Traces"
|
||||
authors: Shu-Xun Yang, Cunxiang Wang, Haoke Zhang, Wenbo Yu, Lindong Wu, Jiayi Gui, Dayong Yang, Yukuo Cen, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.00623
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-28
|
||||
updated_at: 2026-02-28
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- multi-agent
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 19
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, coding-agent, multi-agent, reasoning, tool-use
|
||||
- arXiv categories: cs.AI, cs.CL
|
||||
- collection score: 19
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.00623
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: The Synthetic Web: Adversarially-Curated Mini-Internets for Diagnosing Epistemic Weaknesses of Language Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "The Synthetic Web: Adversarially-Curated Mini-Internets for Diagnosing Epistemic Weaknesses of Language Agents"
|
||||
authors: Shrey Shah, Levent Ozgur
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.00801
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-28
|
||||
updated_at: 2026-02-28
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- embodied-agent
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.IR
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, embodied-agent, rag, tool-use
|
||||
- arXiv categories: cs.AI, cs.IR
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.00801
|
||||
+64
@@ -0,0 +1,64 @@
|
||||
# Paper: Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents
|
||||
authors: Yuxin Liu, Mingye Zhu, Siyuan Liu, Bo Hu, Lei Zhang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.01438
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-02
|
||||
updated_at: 2026-03-02
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-safety
|
||||
- coding-agent
|
||||
- computer-use
|
||||
- planning
|
||||
- rag
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-safety, coding-agent, computer-use, planning, rag, world-model
|
||||
- arXiv categories: cs.CL, cs.AI
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.01438
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents"
|
||||
authors: Qizheng Li, Yifei Zhang, Xiao Yang, Xu Yang, Zhuo Wang, Weiqing Liu, Jiang Bian
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.01712
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-02
|
||||
updated_at: 2026-05-20
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- planning
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, coding-agent, planning
|
||||
- arXiv categories: cs.AI, cs.LG
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.01712
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: A Natural Language Agentic Approach to Study Affective Polarization
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: A Natural Language Agentic Approach to Study Affective Polarization
|
||||
authors: Stephanie Anneris Malvicini, Ewelina Gajewska, Arda Derbent, Katarzyna Budzynska, Jarosław A. Chudziak, Maria Vanina Martinez
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.02711
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-03
|
||||
updated_at: 2026-03-03
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- multi-agent
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: multi-agent, rag, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.02711
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: The Controllability Trap: A Governance Framework for Military AI Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "The Controllability Trap: A Governance Framework for Military AI Agents"
|
||||
authors: Subramanyam Sahoo
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.03515
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-03
|
||||
updated_at: 2026-03-03
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- planning
|
||||
- tool-use
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CY
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, planning, tool-use, world-model
|
||||
- arXiv categories: cs.CY, cs.AI
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.03515
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: MAGE: Meta-Reinforcement Learning for Language Agents toward Strategic Exploration and Exploitation
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "MAGE: Meta-Reinforcement Learning for Language Agents toward Strategic Exploration and Exploitation"
|
||||
authors: Lu Yang, Zelai Xu, Minyang Xie, Jiaxuan Gao, Zhao Shok, Yu Wang, Yi Wu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.03680
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-04
|
||||
updated_at: 2026-03-04
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- memory
|
||||
- multi-agent
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: memory, multi-agent, reasoning, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.03680
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: EigenData: A Self-Evolving Multi-Agent Platform for Function-Calling Data Synthesis, Auditing, and Repair
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "EigenData: A Self-Evolving Multi-Agent Platform for Function-Calling Data Synthesis, Auditing, and Repair"
|
||||
authors: Jiaao Chen, Jingyuan Qi, Mingye Gao, Wei-Chen Wang, Hanrui Wang, Di Jin
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.05553
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-05
|
||||
updated_at: 2026-03-05
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- multi-agent
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
- cs.AI
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, coding-agent, multi-agent, tool-use
|
||||
- arXiv categories: cs.SE, cs.AI, cs.CL
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.05553
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent"
|
||||
authors: Bowei Xia, Mengkang Hu, Shijian Wang, Jiarui Jin, Wenxiang Jiao, Yuan Lu, Kexin Li, Ping Luo
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.05578
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-05
|
||||
updated_at: 2026-03-05
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, computer-use, tool-use
|
||||
- arXiv categories: cs.SE, cs.AI
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.05578
|
||||
+64
@@ -0,0 +1,64 @@
|
||||
# Paper: From Thinker to Society: Security in Hierarchical Autonomy Evolution of AI Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "From Thinker to Society: Security in Hierarchical Autonomy Evolution of AI Agents"
|
||||
authors: Xiaolei Zhang, Lu Zhou, Xiaogang Xu, Jiafei Wu, Tianyu Du, Heqing Huang, Hao Peng, Zhe Liu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.07496
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-08
|
||||
updated_at: 2026-03-21
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- computer-use
|
||||
- multi-agent
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, computer-use, multi-agent, reasoning, tool-use
|
||||
- arXiv categories: cs.CR, cs.AI
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.07496
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: AgentRaft: Automated Detection of Data Over-Exposure in LLM Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "AgentRaft: Automated Detection of Data Over-Exposure in LLM Agents"
|
||||
authors: Yixi Lin, Jiangrong Wu, Yuhong Nan, Xueqiang Wang, Xinyuan Zhang, Zibin Zheng
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.07557
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-08
|
||||
updated_at: 2026-03-08
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, agent-safety, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.SE
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.07557
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: \$OneMillion-Bench: How Far are Language Agents from Human Experts?
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "\\$OneMillion-Bench: How Far are Language Agents from Human Experts?"
|
||||
authors: Qianyu Yang, Yang Liu, Jiaqi Li, Jun Bai, Hao Chen, Kaiyuan Chen, Tiliang Duan, Jiayun Dong, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.07980
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-09
|
||||
updated_at: 2026-03-09
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- planning
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
- cs.AI
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, planning, reasoning, tool-use
|
||||
- arXiv categories: cs.LG, cs.AI, cs.CL
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.07980
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: KernelCraft: Benchmarking for Agentic Close-to-Metal Kernel Generation on Emerging Hardware
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "KernelCraft: Benchmarking for Agentic Close-to-Metal Kernel Generation on Emerging Hardware"
|
||||
authors: Jiayi Nie, Haoran Wu, Yao Lai, Zeyu Cao, Cheng Zhang, Binglei Lou, Erwei Wang, Jianyi Cheng, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.08721
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-10
|
||||
updated_at: 2026-05-29
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- reasoning
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AR
|
||||
- cs.LG
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, reasoning, workflow-agent
|
||||
- arXiv categories: cs.AR, cs.LG, cs.SE
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.08721
|
||||
@@ -0,0 +1,66 @@
|
||||
# Paper: Security Considerations for Multi-agent Systems
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Security Considerations for Multi-agent Systems
|
||||
authors: Tam Nguyen, Moses Ndebugre, Dheeraj Arremsetty
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.09002
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-09
|
||||
updated_at: 2026-04-26
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- computer-use
|
||||
- memory
|
||||
- multi-agent
|
||||
- planning
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, computer-use, memory, multi-agent, planning, rag, tool-use
|
||||
- arXiv categories: cs.CR, cs.AI
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.09002
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: Human-AI Co-reasoning for Clinical Diagnosis with Evidence-Integrated Language Agent
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Human-AI Co-reasoning for Clinical Diagnosis with Evidence-Integrated Language Agent
|
||||
authors: Zhongzhen Huang, Yan Ling, Hong Chen, Ye Feng, Li Wu, Linjie Mu, Shaoting Zhang, Xiaofan Zhang, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.10492
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-11
|
||||
updated_at: 2026-03-18
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- multi-agent
|
||||
- rag
|
||||
- reasoning
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, agent-safety, multi-agent, rag, reasoning, workflow-agent
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.10492
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey"
|
||||
authors: Juhee Kim, Xiaoyuan Liu, Zhun Wang, Shi Qiu, Bo Li, Wenbo Guo, Dawn Song
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.11088
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-11
|
||||
updated_at: 2026-03-11
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-safety
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-safety, tool-use, workflow-agent
|
||||
- arXiv categories: cs.CR, cs.AI
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.11088
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: QUARE: Quality-Aware Requirements Analysis through Multi-Agent Dialectical Negotiation
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "QUARE: Quality-Aware Requirements Analysis through Multi-Agent Dialectical Negotiation"
|
||||
authors: Haowei Cheng, Milhan Kim, Foutse Khomh, Teeradaj Racharak, Nobukazu Yoshioka, Naoyasu Ubayashi, Hironori Washizaki
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.11890
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-12
|
||||
updated_at: 2026-06-05
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- multi-agent
|
||||
- rag
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, multi-agent, rag
|
||||
- arXiv categories: cs.SE
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.11890
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: CCTU: A Benchmark for Tool Use under Complex Constraints
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "CCTU: A Benchmark for Tool Use under Complex Constraints"
|
||||
authors: Junjie Ye, Guoqiang Zhang, Wenjie Fu, Tao Gui, Qi Zhang, Xuanjing Huang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.15309
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-16
|
||||
updated_at: 2026-03-16
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, rag, tool-use
|
||||
- arXiv categories: cs.CL, cs.AI
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.15309
|
||||
+59
@@ -0,0 +1,59 @@
|
||||
# Paper: Compiled Memory: Not More Information, but More Precise Instructions for Language Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Compiled Memory: Not More Information, but More Precise Instructions for Language Agents"
|
||||
authors: James Rhodes, George Kang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.15666
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-12
|
||||
updated_at: 2026-03-12
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- memory
|
||||
- rag
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: memory, rag
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.15666
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Differential Harm Propensity in Personalized LLM Agents: The Curious Case of Mental Health Disclosure
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Differential Harm Propensity in Personalized LLM Agents: The Curious Case of Mental Health Disclosure"
|
||||
authors: Caglar Yildirim
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.16734
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-17
|
||||
updated_at: 2026-03-17
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- memory
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 20
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, memory, rag, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 20
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.16734
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Agentic Cognitive Profiling: Realigning Automated Alzheimer's Disease Detection with Clinical Construct Validity
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Agentic Cognitive Profiling: Realigning Automated Alzheimer's Disease Detection with Clinical Construct Validity"
|
||||
authors: Jiawen Kang, Kun Li, Dongrui Han, Jinchao Li, Junan Li, Lingwei Meng, Xixin Wu, Helen Meng
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.17392
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-18
|
||||
updated_at: 2026-03-18
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.MA
|
||||
- cs.IR
|
||||
- q-bio.NC
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: agent-evaluation, tool-use
|
||||
- arXiv categories: cs.MA, cs.IR, q-bio.NC
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.17392
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
|
||||
authors: Xuan Chen, Lu Yan, Ruqi Zhang, Xiangyu Zhang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.18245
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-18
|
||||
updated_at: 2026-03-18
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- rag
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
- cs.CR
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, rag, tool-use, workflow-agent
|
||||
- arXiv categories: cs.SE, cs.CR
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.18245
|
||||
@@ -0,0 +1,61 @@
|
||||
# Paper: A Framework for Formalizing LLM Agent Security
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: A Framework for Formalizing LLM Agent Security
|
||||
authors: Vincent Siu, Jingxuan He, Kyle Montgomery, Zhun Wang, Neil Gong, Chenguang Wang, Dawn Song
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.19469
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-19
|
||||
updated_at: 2026-03-19
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-safety
|
||||
- memory
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-safety, memory, tool-use
|
||||
- arXiv categories: cs.CR, cs.AI
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.19469
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents"
|
||||
authors: Shaojie Zhuang, Lu Yin, Guangshun Wei, Yunpeng Li, Xilu Wang, Yuanfeng Zhou
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.19684
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-20
|
||||
updated_at: 2026-06-23
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CV
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, coding-agent, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.CV
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.19684
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling"
|
||||
authors: Liang Ding
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.21357
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-22
|
||||
updated_at: 2026-05-10
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- computer-use
|
||||
- memory
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: computer-use, memory, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI, cs.CL
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.21357
|
||||
+64
@@ -0,0 +1,64 @@
|
||||
# Paper: Toward a Theory of Hierarchical Memory for Language Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Toward a Theory of Hierarchical Memory for Language Agents
|
||||
authors: Yashar Talebirad, Ali Parsaee, Csongor Y. Szepesvari, Amirhossein Nadiri, Osmar Zaiane
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.21564
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-23
|
||||
updated_at: 2026-03-23
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- memory
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.IR
|
||||
- cs.AI
|
||||
- cs.IT
|
||||
- cs.SI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, memory, rag, tool-use
|
||||
- arXiv categories: cs.IR, cs.AI, cs.IT, cs.SI
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.21564
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning
|
||||
authors: Tommaso Galliena, Stefano Rosa, Tommaso Apicella, Pietro Morerio, Alessio Del Bue, Lorenzo Natale
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.24257
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-25
|
||||
updated_at: 2026-03-30
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- embodied-agent
|
||||
- memory
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CV
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, embodied-agent, memory
|
||||
- arXiv categories: cs.CV
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.24257
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: SafeGuard ASF: SR Agentic Humanoid Robot System for Autonomous Industrial Safety
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "SafeGuard ASF: SR Agentic Humanoid Robot System for Autonomous Industrial Safety"
|
||||
authors: Thanh Nguyen Canh, Thang Tran Viet, Thanh Tuan Tran, Ben Wei Lim
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.25353
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-26
|
||||
updated_at: 2026-03-26
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-safety
|
||||
- embodied-agent
|
||||
- reasoning
|
||||
- tool-use
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.RO
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-safety, embodied-agent, reasoning, tool-use, world-model
|
||||
- arXiv categories: cs.RO
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.25353
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: SafetyDrift: Predicting When AI Agents Cross the Line Before They Actually Do
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "SafetyDrift: Predicting When AI Agents Cross the Line Before They Actually Do"
|
||||
authors: Aditya Dhodapkar, Farhaan Pishori
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.27148
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-28
|
||||
updated_at: 2026-03-28
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-safety
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-safety, tool-use
|
||||
- arXiv categories: cs.CR, cs.AI
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.27148
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: Evaluating Privilege Usage of Agents with Real-World Tools
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Evaluating Privilege Usage of Agents with Real-World Tools
|
||||
authors: Quan Zhang, Lianhang Fu, Lvsi Lian, Gwihwan Go, Yujue Wang, Chijin Zhou, Yu Jiang, Geguang Pu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.28166
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-30
|
||||
updated_at: 2026-04-20
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- rag
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, rag, tool-use, workflow-agent
|
||||
- arXiv categories: cs.CR, cs.AI
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.28166
|
||||
+64
@@ -0,0 +1,64 @@
|
||||
# Paper: Synergy: A Next-Generation General-Purpose Agent for Open Agentic Web
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Synergy: A Next-Generation General-Purpose Agent for Open Agentic Web"
|
||||
authors: Xiaohang Nie, Zihan Guo, Kezhuo Yang, Zhichong Zheng, Bochen Ge, Shuai Pan, Zeyi Chen, Youling Xiang, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.28428
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-30
|
||||
updated_at: 2026-03-30
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- coding-agent
|
||||
- embodied-agent
|
||||
- memory
|
||||
- multi-agent
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CY
|
||||
- cs.MA
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: function-calling
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling
|
||||
- inferred topics: coding-agent, embodied-agent, memory, multi-agent, rag, tool-use
|
||||
- arXiv categories: cs.CY, cs.MA
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.28428
|
||||
+64
@@ -0,0 +1,64 @@
|
||||
# Paper: Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing
|
||||
authors: Alex Zongo, Filippos Fotiadis, Ufuk Topcu, Peng Wei
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2603.28900
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-03-30
|
||||
updated_at: 2026-03-30
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- multi-agent
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.RO
|
||||
- cs.AI
|
||||
- cs.LG
|
||||
- eess.SY
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, multi-agent, world-model
|
||||
- arXiv categories: cs.RO, cs.AI, cs.LG, eess.SY
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2603.28900
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis"
|
||||
authors: Yu Li, Haoyu Luo, Yuejin Xie, Yuqian Fu, Zhonghao Yang, Shuai Shao, Qihan Ren, Wanying Qu, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2604.02022
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-04-02
|
||||
updated_at: 2026-05-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- planning
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 20
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, planning, rag, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 20
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2604.02022
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Brief Is Better: Non-Monotonic Chain-of-Thought Budget Effects in Function-Calling Language Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Brief Is Better: Non-Monotonic Chain-of-Thought Budget Effects in Function-Calling Language Agents"
|
||||
authors: Xuan Qi
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2604.02155
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-04-02
|
||||
updated_at: 2026-04-02
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: function-calling, language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: function-calling, language-agent
|
||||
- inferred topics: agent-evaluation, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2604.02155
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: Co-Evolution of Policy and Internal Reward for Language Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Co-Evolution of Policy and Internal Reward for Language Agents
|
||||
authors: Xinyu Wang, Hanwei Wu, Jingwei Song, Shuyuan Zhang, Jiayi Zhang, Fanqi Kong, Tung Sum Thomas Kwok, Xiao-Wen Chang, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2604.03098
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-04-03
|
||||
updated_at: 2026-04-03
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- planning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
- cs.AI
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, computer-use, planning, tool-use
|
||||
- arXiv categories: cs.LG, cs.AI, cs.CL
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2604.03098
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: DRAFT: Task Decoupled Latent Reasoning for Agent Safety
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "DRAFT: Task Decoupled Latent Reasoning for Agent Safety"
|
||||
authors: Lin Wang, Junfeng Fang, Dan Zhang, Fei Shen, Xiang Wang, Tat-Seng Chua
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2604.03242
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-02-11
|
||||
updated_at: 2026-02-11
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.LG
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2604.03242
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Profile-Then-Reason: Bounded Semantic Complexity for Tool-Augmented Language Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Profile-Then-Reason: Bounded Semantic Complexity for Tool-Augmented Language Agents"
|
||||
authors: Paulo Akira F. Enabe
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2604.04131
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-04-05
|
||||
updated_at: 2026-04-05
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent
|
||||
- inferred topics: agent-evaluation, rag, reasoning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2604.04131
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: ShieldNet: Network-Level Guardrails against Emerging Supply-Chain Injections in Agentic Systems
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "ShieldNet: Network-Level Guardrails against Emerging Supply-Chain Injections in Agentic Systems"
|
||||
authors: Zhuowen Yuan, Zhaorun Chen, Zhen Xiang, Nathaniel D. Bastian, Seyyed Hadi Hashemi, Chaowei Xiao, Wenbo Guo, Bo Li
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2604.04426
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-04-06
|
||||
updated_at: 2026-04-06
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2604.04426
|
||||
@@ -0,0 +1,60 @@
|
||||
# Paper: ARuleCon: Agentic Security Rule Conversion
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "ARuleCon: Agentic Security Rule Conversion"
|
||||
authors: Ming Xu, Hongtai Wang, Yanpei Guo, Zhengmin Yu, Weili Han, Hoon Wei Lim, Jin Song Dong, Jiaheng Zhang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2604.06762
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-08
|
||||
published_at: 2026-04-08
|
||||
updated_at: 2026-04-08
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- rag
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: agent-safety
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety
|
||||
- inferred topics: agent-evaluation, agent-safety, rag
|
||||
- arXiv categories: cs.CR
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2604.06762
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user