Refresh papers and define Agent evaluation
This commit is contained in:
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation"
|
||||
authors: Andrey Podivilov, Vadim Lomshakov, Sergey Savin, Matvei Startsev, Roman Pozharskiy, Maksim Parshin, Sergey Nikolenko
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.06624
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-14
|
||||
updated_at: 2026-07-14
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- planning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 19
|
||||
collection_queries: agent-evaluation, coding-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-evaluation, coding-agent
|
||||
- inferred topics: agent-evaluation, coding-agent, planning, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 19
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.06624
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents"
|
||||
authors: Vikas Reddy, Sumanth Reddy Challaram, Abhishek Basu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.07405
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-11
|
||||
updated_at: 2026-07-11
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: tool-use
|
||||
- inferred topics: agent-evaluation, reasoning, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.07405
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Out of Sight: Compression-Aware Content Protection against Agentic Crawlers
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Out of Sight: Compression-Aware Content Protection against Agentic Crawlers"
|
||||
authors: Xuefei Wang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.08180
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-09
|
||||
updated_at: 2026-07-09
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- computer-use
|
||||
- memory
|
||||
- reasoning
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: agentic-ai
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agentic-ai
|
||||
- inferred topics: computer-use, memory, reasoning, workflow-agent
|
||||
- arXiv categories: cs.CR
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.08180
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: Multi-Agent Firewall Architecture for Privacy Protection of Sensitive Data in Interactions with Language Models
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Multi-Agent Firewall Architecture for Privacy Protection of Sensitive Data in Interactions with Language Models
|
||||
authors: Hugo García Cuesta, Pablo Mateo Torrejón, Alfonso Sánchez-Macián
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.08282
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-09
|
||||
updated_at: 2026-07-09
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- computer-use
|
||||
- multi-agent
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: multi-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: multi-agent-llm
|
||||
- inferred topics: agent-evaluation, agent-safety, computer-use, multi-agent, tool-use, workflow-agent
|
||||
- arXiv categories: cs.CR
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.08282
|
||||
+65
@@ -0,0 +1,65 @@
|
||||
# Paper: Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents"
|
||||
authors: Yixian Zhang, Huanming Zhang, Feng Gao, Xiao Li, Zhihao Liu, Chunyang Zhu, Jiaxing Qiu, Yuchen Yan, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.08448
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-15
|
||||
updated_at: 2026-07-15
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- coding-agent
|
||||
- computer-use
|
||||
- embodied-agent
|
||||
- memory
|
||||
- planning
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.RO
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 20
|
||||
collection_queries: coding-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: coding-agent
|
||||
- inferred topics: coding-agent, computer-use, embodied-agent, memory, planning, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.RO
|
||||
- collection score: 20
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.08448
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets"
|
||||
authors: Shilin Ou, Yifan Xu, Luyao Zhang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.08681
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-09
|
||||
updated_at: 2026-07-09
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- planning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: agent-evaluation, agentic-ai, autonomous-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-evaluation, agentic-ai, autonomous-agent-llm
|
||||
- inferred topics: agent-evaluation, agent-safety, planning, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.08681
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents"
|
||||
authors: Yifan Wu, Lizhu Zhang, Yuhang Zhou, Mingyi Wang, Bo Peng, Serena Li, Xiangjun Fan, Zhuokai Zhao
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.08716
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-09
|
||||
updated_at: 2026-07-09
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- memory
|
||||
- planning
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: agent-memory
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-memory
|
||||
- inferred topics: agent-evaluation, computer-use, memory, planning, rag, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.08716
|
||||
+64
@@ -0,0 +1,64 @@
|
||||
# Paper: GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning"
|
||||
authors: Maureese Williams, Dymitr Nowicki
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.08894
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-09
|
||||
updated_at: 2026-07-09
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- coding-agent
|
||||
- computer-use
|
||||
- embodied-agent
|
||||
- planning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: language-agent, planning-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent, planning-agent
|
||||
- inferred topics: coding-agent, computer-use, embodied-agent, planning, tool-use, workflow-agent, world-model
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.08894
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution"
|
||||
authors: Ning Liu, Kalle Kujanpää, Zhaoxuan Zhu, P Aditya Sreekar, Kaiwen Liu, Chuanneng Sun, Jorge Marchena Menendez, Matthew Bales, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.08960
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-09
|
||||
updated_at: 2026-07-09
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- memory
|
||||
- multi-agent
|
||||
- reasoning
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 19
|
||||
collection_queries: multi-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: multi-agent-llm
|
||||
- inferred topics: agent-evaluation, computer-use, memory, multi-agent, reasoning
|
||||
- arXiv categories: cs.LG
|
||||
- collection score: 19
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.08960
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation"
|
||||
authors: Sijia Gu, Noor Nashid, Ali Mesbah
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.08983
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-09
|
||||
updated_at: 2026-07-09
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: coding-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: coding-agent
|
||||
- inferred topics: agent-evaluation, coding-agent, rag, tool-use
|
||||
- arXiv categories: cs.SE
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.08983
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls"
|
||||
authors: Saroj Gopali, Bipin Chhetri, Deepika Giri, Sima Siami-Namini, Akbar Siami Namin
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.09076
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-09
|
||||
updated_at: 2026-07-09
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- planning
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: agentic-ai
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agentic-ai
|
||||
- inferred topics: agent-evaluation, agent-safety, planning, reasoning, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.09076
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: AgentKGV: Agentic LLM-RAG Framework with Two-Stage Training for the Fact Verification of Knowledge Graphs
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "AgentKGV: Agentic LLM-RAG Framework with Two-Stage Training for the Fact Verification of Knowledge Graphs"
|
||||
authors: Yumin Heo, Hyeon-gu Lee, Sumin Seo, Youngjoong Ko
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.09092
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-10
|
||||
updated_at: 2026-07-10
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: rag-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: rag-agent
|
||||
- inferred topics: agent-evaluation, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.09092
|
||||
+65
@@ -0,0 +1,65 @@
|
||||
# Paper: Multi-Agent LLM Collaboration for Unit Test Generation via Human-Testing-Inspired Workflows
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Multi-Agent LLM Collaboration for Unit Test Generation via Human-Testing-Inspired Workflows
|
||||
authors: Quanjun Zhang, Ye Shang, Siqi Gu, Jianyi Zhou, Chunrong Fang, Zhenyu Chen, Liang Xiao
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.09101
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-10
|
||||
updated_at: 2026-07-10
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- multi-agent
|
||||
- planning
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: multi-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: multi-agent-llm
|
||||
- inferred topics: agent-evaluation, coding-agent, multi-agent, planning, rag, reasoning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.SE
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.09101
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling"
|
||||
authors: Peng Kuang, Haibo Jin, Xiaoyu Han, Yanli Wang, Xiaopeng Yuan, Ye Yu, Kaidi Xu, Haohan Wang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.09153
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-10
|
||||
updated_at: 2026-07-10
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- computer-use
|
||||
- memory
|
||||
- multi-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: multi-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: multi-agent-llm
|
||||
- inferred topics: agent-evaluation, coding-agent, computer-use, memory, multi-agent
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.09153
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: Malaika: Understanding Malware through Tri-Grounded Agentic Reasoning
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Malaika: Understanding Malware through Tri-Grounded Agentic Reasoning"
|
||||
authors: Xingzhi Qian, Xinran Zheng, Yiling He, Lorenzo Cavallaro
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.09179
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-10
|
||||
updated_at: 2026-07-10
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- multi-agent
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: multi-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: multi-agent-llm
|
||||
- inferred topics: agent-evaluation, agent-safety, multi-agent, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.CR
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.09179
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents"
|
||||
authors: Izumi Takahara, Teruyasu Mizoguchi
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.09195
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-10
|
||||
updated_at: 2026-07-10
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- planning
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: planning-agent, tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: planning-agent, tool-use
|
||||
- inferred topics: agent-evaluation, planning, reasoning, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.09195
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: Writing Bug Reports for Software Repair Agents: What Information Matters Most?
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Writing Bug Reports for Software Repair Agents: What Information Matters Most?"
|
||||
authors: Vincenzo Luigi Bruno, Alessandro Giagnorio, Daniele Bifolco, Leon Wienges, Massimiliano Di Penta, Gabriele Bavota
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.09553
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-10
|
||||
updated_at: 2026-07-10
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: coding-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: coding-agent
|
||||
- inferred topics: agent-evaluation, coding-agent, workflow-agent
|
||||
- arXiv categories: cs.SE
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.09553
|
||||
+65
@@ -0,0 +1,65 @@
|
||||
# Paper: VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents"
|
||||
authors: Katherine Swinea, Kshitiz Aryal, Lopamudra Praharaj, Maanak Gupta
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.09653
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-10
|
||||
updated_at: 2026-07-10
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- multi-agent
|
||||
- planning
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 20
|
||||
collection_queries: multi-agent-llm, planning-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: multi-agent-llm, planning-agent
|
||||
- inferred topics: agent-evaluation, agent-safety, multi-agent, planning, rag, reasoning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.CR
|
||||
- collection score: 20
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.09653
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Do These Violent Delights Have Violent Ends? Measuring the Post-Merge Fate of Agentic Code
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Do These Violent Delights Have Violent Ends? Measuring the Post-Merge Fate of Agentic Code
|
||||
authors: Chunqiu Steven Xia, Courtney Miller
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.09902
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-10
|
||||
updated_at: 2026-07-10
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- coding-agent
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: coding-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: coding-agent
|
||||
- inferred topics: agent-evaluation, agent-safety, coding-agent, rag, tool-use
|
||||
- arXiv categories: cs.SE
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.09902
|
||||
+59
@@ -0,0 +1,59 @@
|
||||
# Paper: Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?"
|
||||
authors: Jiale Liu, Huajun Xi, Shaokun Zhang, Yifan Zeng, Tianwei Yue, Chi Wang, Jian Kang, Qingyun Wu, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.09996
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-10
|
||||
updated_at: 2026-07-10
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: ai-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: ai-agent
|
||||
- inferred topics: agent-evaluation, computer-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.09996
|
||||
+59
@@ -0,0 +1,59 @@
|
||||
# Paper: Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation"
|
||||
authors: Dongping Liu, Aoyu Zhang, Luyao Zhang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.10057
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-10
|
||||
updated_at: 2026-07-10
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- reasoning
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- quant-ph
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: ai-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: ai-agent
|
||||
- inferred topics: agent-evaluation, reasoning
|
||||
- arXiv categories: quant-ph
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.10057
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: TGMS: An Agent-Native Bi-Temporal Graph Management System
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "TGMS: An Agent-Native Bi-Temporal Graph Management System"
|
||||
authors: Xiaofei Zhang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.10265
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-11
|
||||
updated_at: 2026-07-11
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- planning
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.DB
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: planning-agent, rag-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: planning-agent, rag-agent
|
||||
- inferred topics: agent-evaluation, planning, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.DB
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.10265
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: Can Agentic Trading Systems Pay for Their Own Intelligence?
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Can Agentic Trading Systems Pay for Their Own Intelligence?
|
||||
authors: Qiqi Duan, Changlun Li, Chen Wang, Fan Zhang, Mengxiang Wang, Dayi Miao, Peixian Ma, Jiangpeng Yan, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.10286
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-11
|
||||
updated_at: 2026-07-11
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: tool-use
|
||||
- inferred topics: agent-evaluation, reasoning, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.10286
|
||||
@@ -0,0 +1,61 @@
|
||||
# Paper: GRASP: GRanularity-Aware Search Policy for Agentic RAG
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "GRASP: GRanularity-Aware Search Policy for Agentic RAG"
|
||||
authors: Varun Gandhi, Jaewook Lee, Shantanu Todmal, Franck Dernoncourt, Ryan Rossi, Zichao Wang, Andrew Lan
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.10463
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-11
|
||||
updated_at: 2026-07-11
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: rag-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: rag-agent
|
||||
- inferred topics: agent-evaluation, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.10463
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations"
|
||||
authors: Ruksat Khan Shayoni, Muhammad Faraz Shoaib, S M Asif Hossain, M. F. Mridha
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.10490
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-11
|
||||
updated_at: 2026-07-11
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: tool-use
|
||||
- inferred topics: agent-evaluation, agent-safety, tool-use
|
||||
- arXiv categories: cs.CR
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.10490
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference"
|
||||
authors: Venkatesha Matam, Keon Kim
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.10582
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-12
|
||||
updated_at: 2026-07-12
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- memory
|
||||
- planning
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: planning-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: planning-agent
|
||||
- inferred topics: agent-evaluation, coding-agent, memory, planning, reasoning, tool-use
|
||||
- arXiv categories: cs.LG
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.10582
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory"
|
||||
authors: Yixiong Chen, Xinyi Bai, Alan Yuille
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.10608
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-12
|
||||
updated_at: 2026-07-12
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- memory
|
||||
- planning
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: ai-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: ai-agent
|
||||
- inferred topics: agent-evaluation, computer-use, memory, planning, rag, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.10608
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging"
|
||||
authors: Siyi Chen, Jiahe Ying, Yixuan Jia, Yuxuan Gu, Enze Ye, Weimin Bai, Zhijun Zeng, Shaochi Ren, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.10789
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-12
|
||||
updated_at: 2026-07-12
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- planning
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: coding-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: coding-agent
|
||||
- inferred topics: agent-evaluation, coding-agent, planning
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.10789
|
||||
+66
@@ -0,0 +1,66 @@
|
||||
# Paper: A Multi-Agent Framework for Zero-Dimensional Reduced-Order Model Planning
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: A Multi-Agent Framework for Zero-Dimensional Reduced-Order Model Planning
|
||||
authors: Bingteng Sun, Hao Yin, Yiling Chen, Renjie Xiao, Lei Xie, Shanyou Wang, Ruonan Wang, Shubao Chen, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.10994
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-12
|
||||
updated_at: 2026-07-12
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- embodied-agent
|
||||
- multi-agent
|
||||
- planning
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 22
|
||||
collection_queries: multi-agent-llm, planning-agent, rag-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: multi-agent-llm, planning-agent, rag-agent
|
||||
- inferred topics: agent-evaluation, computer-use, embodied-agent, multi-agent, planning, rag, reasoning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.LG
|
||||
- collection score: 22
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.10994
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services"
|
||||
authors: Yuzhe Guo, Mengzhou Wu, Yuan Cao, Jialei Wei, Dezhi Ran, Wei Yang, Tao Xie
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.11042
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-12
|
||||
updated_at: 2026-07-12
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: coding-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: coding-agent
|
||||
- inferred topics: agent-evaluation, coding-agent, tool-use
|
||||
- arXiv categories: cs.SE
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.11042
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Retrieval-Oriented Code Representations in Agentic Bug Localization
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Retrieval-Oriented Code Representations in Agentic Bug Localization
|
||||
authors: Genevieve Caumartin, Tse-Hsun, Chen, Diego Elias Costa
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.11046
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-12
|
||||
updated_at: 2026-07-12
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- computer-use
|
||||
- rag
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: coding-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: coding-agent
|
||||
- inferred topics: agent-evaluation, coding-agent, computer-use, rag
|
||||
- arXiv categories: cs.SE
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.11046
|
||||
@@ -0,0 +1,64 @@
|
||||
# Paper: VIA: Visual Interface Agent for Robot Control
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "VIA: Visual Interface Agent for Robot Control"
|
||||
authors: Hengyuan Hu, Priya Sundaresan, Jensen Gao, Dorsa Sadigh
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.11119
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- coding-agent
|
||||
- computer-use
|
||||
- embodied-agent
|
||||
- planning
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.RO
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: web-gui-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: web-gui-agent
|
||||
- inferred topics: coding-agent, computer-use, embodied-agent, planning, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.RO
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.11119
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: ToolAtlas: Learning Once, Reusing Everywhere with Tool-Side Memory
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "ToolAtlas: Learning Once, Reusing Everywhere with Tool-Side Memory"
|
||||
authors: Yue Fang, Zhibang Yang, Fangkai Yang, Xiaoting Qin, Liqun Li, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.11126
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- memory
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: tool-use
|
||||
- inferred topics: agent-evaluation, computer-use, memory, tool-use
|
||||
- arXiv categories: cs.LG
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.11126
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management"
|
||||
authors: Changlun Li, Peixian Ma, Qiqi Duan, Zhenyu Lin, Peineng Wu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.11141
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- multi-agent
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: multi-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: multi-agent-llm
|
||||
- inferred topics: agent-evaluation, agent-safety, multi-agent, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.11141
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL"
|
||||
authors: Bowen Lv, Xiao Liu, Yanyu Ren, Hanyu Lai, Bohao Jing, Hanchen Zhang, Yanxiao Zhao, Shuntian Yao, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.11185
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- computer-use
|
||||
- multi-agent
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: web-gui-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: web-gui-agent
|
||||
- inferred topics: computer-use, multi-agent, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.11185
|
||||
@@ -0,0 +1,61 @@
|
||||
# Paper: Multi-Agent LLMs Fail to Explore Each Other
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Multi-Agent LLMs Fail to Explore Each Other
|
||||
authors: Hyeong Kyu Choi, Jiatong Li, Wendi Li, Xin Eric Wang, Sharon Li
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.11250
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- multi-agent
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.MA
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: multi-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: multi-agent-llm
|
||||
- inferred topics: agent-evaluation, computer-use, multi-agent, tool-use
|
||||
- arXiv categories: cs.MA
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.11250
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure"
|
||||
authors: Wenyi Wu, Sibo Zhu, Kun Zhou, Aayush Salvi, Zixuan Song, Biwei Huang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.11388
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- computer-use
|
||||
- planning
|
||||
- reasoning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: web-gui-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: web-gui-agent
|
||||
- inferred topics: computer-use, planning, reasoning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.11388
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: ToFu: A White-Box, Token-Efficient Agent Harness for Researchers
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "ToFu: A White-Box, Token-Efficient Agent Harness for Researchers"
|
||||
authors: Junhao Ruan, Yuan Ge, Bei Li, Yongjing Yin, Yuchun Fan, Xin Chen, Jingang Wang, Chenglong Wang, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.11423
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: tool-use
|
||||
- inferred topics: agent-evaluation, coding-agent, tool-use, workflow-agent
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.11423
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: UMoE:Unlocking Every Expert in Domain-Specific Training
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "UMoE:Unlocking Every Expert in Domain-Specific Training"
|
||||
authors: Xuefeng Li, Pengfei Liu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.11444
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: coding-agent, tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: coding-agent, tool-use
|
||||
- inferred topics: agent-evaluation, coding-agent, rag, tool-use
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.11444
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Heuristic Learning for Active Flow Control Using Coding Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Heuristic Learning for Active Flow Control Using Coding Agents
|
||||
authors: Paul Garnier, Jonathan Viquerat, Elie Hachem
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.11565
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- tool-use
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: coding-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: coding-agent
|
||||
- inferred topics: agent-evaluation, coding-agent, tool-use, world-model
|
||||
- arXiv categories: cs.LG
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.11565
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems"
|
||||
authors: Yibo Hu, Ren Wang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.11751
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- coding-agent
|
||||
- multi-agent
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: multi-agent-llm, tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: multi-agent-llm, tool-use
|
||||
- inferred topics: agent-evaluation, agent-safety, coding-agent, multi-agent, rag, tool-use
|
||||
- arXiv categories: cs.CR
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.11751
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability"
|
||||
authors: Said Elnaffar, Farzad Rashidi
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.12056
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- rag
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: ai-agent, web-gui-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: ai-agent, web-gui-agent
|
||||
- inferred topics: agent-evaluation, computer-use, rag, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.12056
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Beyond Test Presence: Assessing the Quality and Robustness of Agent-Generated Tests in Open-Source Projects
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Beyond Test Presence: Assessing the Quality and Robustness of Agent-Generated Tests in Open-Source Projects"
|
||||
authors: Preet Jhanglani, Zeel Kaushal Desai, Vidhi Kansara, Eman Abdullah AlOmar
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.12068
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- coding-agent
|
||||
- rag
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: ai-agent, coding-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: ai-agent, coding-agent
|
||||
- inferred topics: agent-evaluation, agent-safety, coding-agent, rag
|
||||
- arXiv categories: cs.SE
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.12068
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: An Agentic AI Scientific Community for Automated Neural Operator Discovery
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: An Agentic AI Scientific Community for Automated Neural Operator Discovery
|
||||
authors: Luis Loo, Ulisses Braga-Neto
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.12122
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- planning
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: agentic-ai
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agentic-ai
|
||||
- inferred topics: agent-evaluation, planning, world-model
|
||||
- arXiv categories: cs.LG
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.12122
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: Fine-Tuned Multi-Agent Framework for Detecting OCEAN in Life Narratives
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Fine-Tuned Multi-Agent Framework for Detecting OCEAN in Life Narratives
|
||||
authors: Rasiq Hussain, Darshil Italiya, Joshua Oltmanns, Mehak Gupta
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.12215
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- multi-agent
|
||||
- reasoning
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: multi-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: multi-agent-llm
|
||||
- inferred topics: agent-evaluation, multi-agent, reasoning
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.12215
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents"
|
||||
authors: Ning Liu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.12267
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- memory
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 19
|
||||
collection_queries: language-agent, tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent, tool-use
|
||||
- inferred topics: agent-evaluation, memory, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.LG
|
||||
- collection score: 19
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.12267
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents"
|
||||
authors: Yaopei Zeng, Congchao Wang, JianHang Chen, Nan Wang, Yurui Chang, Lu Lin
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.12397
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-14
|
||||
updated_at: 2026-07-14
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- memory
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: agent-evaluation
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-evaluation
|
||||
- inferred topics: agent-evaluation, memory, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.12397
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions"
|
||||
authors: Huihao Jing, Wenbin Hu, Shaojin Chen, Haochen Shi, Sirui Zhang, Hanyu Yang, Changxuan Fan, Zhongwei Xie, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.12406
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-14
|
||||
updated_at: 2026-07-14
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- memory
|
||||
- rag
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: agentic-ai
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agentic-ai
|
||||
- inferred topics: agent-evaluation, agent-safety, memory, rag, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.12406
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
|
||||
authors: Yubo Wang, Jiarong Liang, Yuxuan Zhang, Xuye Liu, Cong Wei, Yuyu Zhang, Ping Nie, Wenhu Chen
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.12463
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-19
|
||||
updated_at: 2026-07-19
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: coding-agent, function-calling, tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: coding-agent, function-calling, tool-use
|
||||
- inferred topics: agent-evaluation, coding-agent, reasoning, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.12463
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: PalmClaw: A Native On-Device Agent Framework for Mobile Phones
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "PalmClaw: A Native On-Device Agent Framework for Mobile Phones"
|
||||
authors: Hongru Cai, Yongqi Li, Ran Wei, Wenjie Li
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.13027
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-14
|
||||
updated_at: 2026-07-14
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- computer-use
|
||||
- memory
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: tool-use
|
||||
- inferred topics: computer-use, memory, tool-use, workflow-agent
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.13027
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification"
|
||||
authors: SingGuard Team
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.13081
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- computer-use
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: agent-safety, agentic-ai
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-safety, agentic-ai
|
||||
- inferred topics: agent-evaluation, agent-safety, computer-use, reasoning, tool-use
|
||||
- arXiv categories: cs.CR
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.13081
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testing
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testing"
|
||||
authors: Ananda Dhakal, Krish Neupane, Aarjan Chaudhary
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.13085
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-13
|
||||
updated_at: 2026-07-13
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- coding-agent
|
||||
- rag
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: coding-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: coding-agent
|
||||
- inferred topics: agent-evaluation, agent-safety, coding-agent, rag
|
||||
- arXiv categories: cs.CR
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.13085
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality"
|
||||
authors: Suzhen Zhong, Shayan Noei, Bram Adams, Ying Zou
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.13196
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-14
|
||||
updated_at: 2026-07-14
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- computer-use
|
||||
- multi-agent
|
||||
- planning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: ai-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: ai-agent
|
||||
- inferred topics: computer-use, multi-agent, planning, tool-use
|
||||
- arXiv categories: cs.SE
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.13196
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents"
|
||||
authors: Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu, Levina Li, Dong Liu, Xiao Liang, Rui Sun, Yubei Li, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.13591
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-15
|
||||
updated_at: 2026-07-15
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- memory
|
||||
- planning
|
||||
- rag
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 19
|
||||
collection_queries: planning-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: planning-agent
|
||||
- inferred topics: agent-evaluation, memory, planning, rag
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 19
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.13591
|
||||
+59
@@ -0,0 +1,59 @@
|
||||
# Paper: STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle"
|
||||
authors: Sagar Deb, Ashwanth Krishnan
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.13618
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-15
|
||||
updated_at: 2026-07-15
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: llm-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: llm-agent
|
||||
- inferred topics: agent-evaluation, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.13618
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities"
|
||||
authors: Kai Chen, Zichen Ding, Jiaye Ge, Shufan Jiang, Mo Li, Qingqiu Li, Zehao Li, Zonglin Li, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.13705
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-20
|
||||
updated_at: 2026-07-20
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: autonomous-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: autonomous-agent-llm
|
||||
- inferred topics: agent-evaluation, rag, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.13705
|
||||
@@ -0,0 +1,64 @@
|
||||
# Paper: SPyCE: Skill-Policy Co-evolution for Multimodal Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "SPyCE: Skill-Policy Co-evolution for Multimodal Agents"
|
||||
authors: Ru Zhang, Weijie Qiu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.13854
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-15
|
||||
updated_at: 2026-07-15
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- memory
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: tool-use
|
||||
- inferred topics: agent-evaluation, computer-use, memory, rag, reasoning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.13854
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Agent Skill Security: Threat Models, Attacks, Defenses, and Evaluation
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Agent Skill Security: Threat Models, Attacks, Defenses, and Evaluation"
|
||||
authors: Sanket Badhe, Priyanka Tiwari
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.13987
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-15
|
||||
updated_at: 2026-07-15
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- coding-agent
|
||||
- planning
|
||||
- rag
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: llm-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: llm-agent
|
||||
- inferred topics: agent-evaluation, agent-safety, coding-agent, planning, rag
|
||||
- arXiv categories: cs.CR
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.13987
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability"
|
||||
authors: Weiting Liu, Jieyi Bi, Wanqi Zhou, Jianfeng Feng, Yining Ma, Ai Han, Wenlian Lu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.14145
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-14
|
||||
updated_at: 2026-07-14
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- planning
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: tool-use
|
||||
- inferred topics: agent-evaluation, planning, reasoning, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.14145
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Towards Reliable AI-Assisted Analog Design: Template-Constrained LLM Agents for SAR ADC Generation
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Towards Reliable AI-Assisted Analog Design: Template-Constrained LLM Agents for SAR ADC Generation"
|
||||
authors: Dimple Vijay Kochar, Hae-Seung Lee, Anantha P. Chandrakasan
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.14165
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-15
|
||||
updated_at: 2026-07-15
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- planning
|
||||
- rag
|
||||
- workflow-agent
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: llm-agent, planning-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: llm-agent, planning-agent
|
||||
- inferred topics: planning, rag, workflow-agent, world-model
|
||||
- arXiv categories: cs.SE
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.14165
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System"
|
||||
authors: Yutong He, Daibo Li, Guohong Li, Jiahe Geng, Zhengyang Huang, Can Ren, Zekun Zhang, Yifan Liu, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.14178
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-15
|
||||
updated_at: 2026-07-15
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- multi-agent
|
||||
- rag
|
||||
- reasoning
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: ai-agent, autonomous-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: ai-agent, autonomous-agent-llm
|
||||
- inferred topics: agent-evaluation, multi-agent, rag, reasoning
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.14178
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: MonteRET: AI Agent Enhancing Multimodal LLMs with Multi-granularity Knowledge Retrieval for Chest CT Report Generation
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "MonteRET: AI Agent Enhancing Multimodal LLMs with Multi-granularity Knowledge Retrieval for Chest CT Report Generation"
|
||||
authors: Yi Lin, Yihao Ding, Elana Benishay, Elefterios Trikantzopoulos, David Nauheim, Hanley Ong, Jiang Bian, Hua Xu, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.14264
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-15
|
||||
updated_at: 2026-07-15
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- computer-use
|
||||
- rag
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CV
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: ai-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: ai-agent
|
||||
- inferred topics: agent-evaluation, agent-safety, computer-use, rag
|
||||
- arXiv categories: cs.CV
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.14264
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making"
|
||||
authors: Amirhosein Ghasemabadi, Ruichen Chen, Bahador Rashidi, Di Niu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.14277
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-15
|
||||
updated_at: 2026-07-15
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: llm-agent, tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: llm-agent, tool-use
|
||||
- inferred topics: agent-evaluation, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.14277
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: CIPHER: A Decoupled Exploration-Selection Framework for Test-Time Scaling of Data Science Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "CIPHER: A Decoupled Exploration-Selection Framework for Test-Time Scaling of Data Science Agents"
|
||||
authors: Maxime Heuillet, Sharadind Peddiraju
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.14386
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-15
|
||||
updated_at: 2026-07-15
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- rag
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: ai-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: ai-agent
|
||||
- inferred topics: agent-evaluation, rag, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.14386
|
||||
+59
@@ -0,0 +1,59 @@
|
||||
# Paper: Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution"
|
||||
authors: Harris Borman, Herman Wandabwa, Fusun Yu, Sandeepa Kannangara, Justin Liu, Anna Leontjeva, Ritchie Ng
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.14456
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-15
|
||||
updated_at: 2026-07-15
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: agentic-ai, tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agentic-ai, tool-use
|
||||
- inferred topics: tool-use, workflow-agent
|
||||
- arXiv categories: cs.SE
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.14456
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents"
|
||||
authors: Hy Vision Team, Huawen Shen, Zhengyang Tang, Shangpin Peng, Liang Wu, Anran Zhang, Weinong Wang, Yiduo Guo, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.14548
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-16
|
||||
updated_at: 2026-07-16
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- computer-use
|
||||
- embodied-agent
|
||||
- planning
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CV
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: web-gui-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: web-gui-agent
|
||||
- inferred topics: computer-use, embodied-agent, planning, reasoning, tool-use
|
||||
- arXiv categories: cs.CV
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.14548
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents"
|
||||
authors: Shiyu Ying, Xuejie Cao, Yingfan Ma, Yuanhao Dong, Wenyu Chen, Bowen Song, Lin Zhu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.14573
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-21
|
||||
updated_at: 2026-07-21
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- coding-agent
|
||||
- computer-use
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 19
|
||||
collection_queries: coding-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: coding-agent
|
||||
- inferred topics: agent-evaluation, agent-safety, coding-agent, computer-use, rag, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 19
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.14573
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers"
|
||||
authors: Huanxi Liu, Kun Hu, Jiaqi Liao, Qiang Wang, Pengfei Qian, YuanZhao Zhai, Dawei Feng, Bo Ding, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.14642
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-16
|
||||
updated_at: 2026-07-16
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- planning
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 21
|
||||
collection_queries: llm-agent, planning-agent, tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: llm-agent, planning-agent, tool-use
|
||||
- inferred topics: agent-evaluation, planning, rag, reasoning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 21
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.14642
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents"
|
||||
authors: Jifeng Gao, Kang Xia, Yi Zhang, Xiaobin Hong, Mingkai Lin, Xingshen Wei, Wenzhong Li, Sanglu Lu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.14651
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-16
|
||||
updated_at: 2026-07-16
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- memory
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 20
|
||||
collection_queries: llm-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: llm-agent
|
||||
- inferred topics: agent-evaluation, agent-safety, memory, rag, tool-use
|
||||
- arXiv categories: cs.CR
|
||||
- collection score: 20
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.14651
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning"
|
||||
authors: Jinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.14777
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-16
|
||||
updated_at: 2026-07-16
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- computer-use
|
||||
- planning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: tool-use
|
||||
- inferred topics: computer-use, planning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.14777
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios"
|
||||
authors: Chengyu Shen, Yujie Fu, Gangtao Xin, Yanheng Hou, Wenlong Fei, Guojie Zhu, Jiawei Li, Hongcheng Gao, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.14989
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-16
|
||||
updated_at: 2026-07-16
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- planning
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 19
|
||||
collection_queries: agent-evaluation, ai-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-evaluation, ai-agent
|
||||
- inferred topics: agent-evaluation, planning, rag, tool-use
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 19
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.14989
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: BrainPilot: Automating Brain Discovery with Agentic Research
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "BrainPilot: Automating Brain Discovery with Agentic Research"
|
||||
authors: Haoxuan Li, Tianci Gao, Jianhe Li, Yang Fan, Runze Shi, Weiran Wang, Tianxiang Zhao, Zezhao Wu, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.15079
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-17
|
||||
updated_at: 2026-07-17
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- multi-agent
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 20
|
||||
collection_queries: ai-agent, tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: ai-agent, tool-use
|
||||
- inferred topics: agent-evaluation, multi-agent, rag, reasoning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 20
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.15079
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Plover: Steering GUI Agents through Plan-Centric Interaction
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Plover: Steering GUI Agents through Plan-Centric Interaction"
|
||||
authors: Madhumitha Venkatesan, Shicheng Wen, Jiajing Guo, Jorge Piazentin Ono, Liu Ren, Dongyu Liu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.15193
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-16
|
||||
updated_at: 2026-07-16
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- planning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: web-gui-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: web-gui-agent
|
||||
- inferred topics: agent-evaluation, computer-use, planning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.15193
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors"
|
||||
authors: Christos Korgialas, Gabriel Lee Jun Rong, Dion Jia Xu Ho, Pai Chet Ng, Xiaoxiao Miao, Konstantinos N. Plataniotis
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.15246
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-16
|
||||
updated_at: 2026-07-16
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- multi-agent
|
||||
- rag
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CV
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 17
|
||||
collection_queries: agent-evaluation, multi-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-evaluation, multi-agent-llm
|
||||
- inferred topics: agent-evaluation, multi-agent, rag
|
||||
- arXiv categories: cs.CV
|
||||
- collection score: 17
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.15246
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents"
|
||||
authors: Paul Kassianik, Blaine Nelson, Yaron Singer
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.15263
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-17
|
||||
updated_at: 2026-07-17
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- embodied-agent
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CR
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: agent-evaluation, tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-evaluation, tool-use
|
||||
- inferred topics: agent-evaluation, agent-safety, embodied-agent, reasoning, tool-use
|
||||
- arXiv categories: cs.CR
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.15263
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery"
|
||||
authors: Raunak B Sinha
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.15367
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-16
|
||||
updated_at: 2026-07-16
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-safety
|
||||
- computer-use
|
||||
- multi-agent
|
||||
- planning
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: multi-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: multi-agent-llm
|
||||
- inferred topics: agent-safety, computer-use, multi-agent, planning, rag, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.15367
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation"
|
||||
authors: Jasmine Brazilek, Maheep Chaudhary, Zoe Lu, Miles Tidmarsh
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.15434
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-22
|
||||
updated_at: 2026-07-22
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- multi-agent
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.MA
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 21
|
||||
collection_queries: agent-evaluation, ai-agent, multi-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-evaluation, ai-agent, multi-agent-llm
|
||||
- inferred topics: agent-evaluation, multi-agent, tool-use
|
||||
- arXiv categories: cs.MA
|
||||
- collection score: 21
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.15434
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: A Tool-Invariant Framework for Teaching and Assessing Computational Methods in the Age of Agentic AI
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: A Tool-Invariant Framework for Teaching and Assessing Computational Methods in the Age of Agentic AI
|
||||
authors: Larry Engelhardt
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.15518
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-16
|
||||
updated_at: 2026-07-16
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- tool-use
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- physics.ed-ph
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: agentic-ai, language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agentic-ai, language-agent
|
||||
- inferred topics: agent-evaluation, coding-agent, tool-use, world-model
|
||||
- arXiv categories: physics.ed-ph
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.15518
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Symbolic Predicate-Guided Language Agents for Inverse Design of Perovskite Oxides
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Symbolic Predicate-Guided Language Agents for Inverse Design of Perovskite Oxides
|
||||
authors: Dong Hyeon Mok, Seoin Back, Victor Fung, Guoxiang Hu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.15535
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-16
|
||||
updated_at: 2026-07-16
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- multi-agent
|
||||
- rag
|
||||
- reasoning
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cond-mat.mtrl-sci
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 19
|
||||
collection_queries: language-agent, llm-agent, multi-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: language-agent, llm-agent, multi-agent-llm
|
||||
- inferred topics: agent-evaluation, computer-use, multi-agent, rag, reasoning
|
||||
- arXiv categories: cond-mat.mtrl-sci
|
||||
- collection score: 19
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.15535
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents"
|
||||
authors: Yanze Wang, Pengfei Yao, Tianyi Sun, Chuanrui Hu, Yan Xiao, Yunyun Han, Yifan Chen, Jun Sun, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.15557
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-23
|
||||
updated_at: 2026-07-23
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- rag
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: llm-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: llm-agent
|
||||
- inferred topics: agent-evaluation, agent-safety, rag
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.15557
|
||||
@@ -0,0 +1,62 @@
|
||||
# Paper: Scalable LLM Agent Tool Access in the Cloud
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Scalable LLM Agent Tool Access in the Cloud
|
||||
authors: Mingxin Li, Enge Song, Yueshang Zuo, Xiaodong Liu, Rong Wen, Qiang Fu, Gianni Antichi, Jian He, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.15593
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-16
|
||||
updated_at: 2026-07-16
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- memory
|
||||
- planning
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.DC
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 15
|
||||
collection_queries: llm-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: llm-agent
|
||||
- inferred topics: agent-evaluation, memory, planning, rag, tool-use
|
||||
- arXiv categories: cs.DC
|
||||
- collection score: 15
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.15593
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning"
|
||||
authors: Shuaiyu Zhou, Fengpeng Yue, Zengjie Hu, Yuanzhe Shen, Chenyang Zhang, feng hong, Cao Liu, Ke Zeng
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.15660
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-17
|
||||
updated_at: 2026-07-17
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- planning
|
||||
- reasoning
|
||||
- tool-use
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 20
|
||||
collection_queries: agent-evaluation, llm-agent, tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-evaluation, llm-agent, tool-use
|
||||
- inferred topics: agent-evaluation, planning, reasoning, tool-use, world-model
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 20
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.15660
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents"
|
||||
authors: Lujia Zhang, Xingzhou Chen, Hongwei Feng
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.15715
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-17
|
||||
updated_at: 2026-07-17
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- memory
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: llm-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: llm-agent
|
||||
- inferred topics: agent-evaluation, memory, rag, reasoning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.15715
|
||||
+64
@@ -0,0 +1,64 @@
|
||||
# Paper: AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets"
|
||||
authors: Ming Chen, Pranav Pai
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.15781
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-24
|
||||
updated_at: 2026-07-24
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- coding-agent
|
||||
- multi-agent
|
||||
- planning
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 21
|
||||
collection_queries: multi-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: multi-agent-llm
|
||||
- inferred topics: agent-evaluation, agent-safety, coding-agent, multi-agent, planning, rag, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 21
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.15781
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning"
|
||||
authors: Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, Karim Lakhani
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.16057
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-20
|
||||
updated_at: 2026-07-20
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CL
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: tool-use
|
||||
- inferred topics: agent-evaluation, coding-agent, reasoning, tool-use
|
||||
- arXiv categories: cs.CL
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.16057
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: When Do Multi-Agent Systems Help? An Information Bottleneck Perspective
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: When Do Multi-Agent Systems Help? An Information Bottleneck Perspective
|
||||
authors: Wendi Yu, Lianhao Zhou, Xiangjue Dong, Sai Sudarshan Barath, Declan Staunton, Byung-Jun Yoon, Xiaoning Qian, James Caverlee, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.16133
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-17
|
||||
updated_at: 2026-07-17
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- multi-agent
|
||||
- reasoning
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.LG
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: multi-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: multi-agent-llm
|
||||
- inferred topics: agent-evaluation, multi-agent, reasoning
|
||||
- arXiv categories: cs.LG
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.16133
|
||||
@@ -0,0 +1,63 @@
|
||||
# Paper: Just A Rather Very Intelligent Spoken Agent
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Just A Rather Very Intelligent Spoken Agent
|
||||
authors: Chen Chen, Zhehuai Chen
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.16610
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-17
|
||||
updated_at: 2026-07-17
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- multi-agent
|
||||
- planning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: agentic-ai, ai-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agentic-ai, ai-agent
|
||||
- inferred topics: agent-evaluation, computer-use, multi-agent, planning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.16610
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
# Paper: Model-Driven Discipline for Multi-Agent LLMs: Requirement-to-Verification Generation of Traceable System Models
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Model-Driven Discipline for Multi-Agent LLMs: Requirement-to-Verification Generation of Traceable System Models"
|
||||
authors: Ran Wei, Le Zhu, Haochi Wang, Ruizhe Yang, Jiapeng Guan, Siyuan Ji, Yuchen Hu, Zhe Jiang, et al.
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.16708
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-18
|
||||
updated_at: 2026-07-18
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- multi-agent
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.SE
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: multi-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: multi-agent-llm
|
||||
- inferred topics: agent-evaluation, agent-safety, multi-agent, tool-use
|
||||
- arXiv categories: cs.SE
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.16708
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: AgentBrew: Lifelong Knowledge Brewing from Strong Teachers to Weak LLM Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "AgentBrew: Lifelong Knowledge Brewing from Strong Teachers to Weak LLM Agents"
|
||||
authors: Yangqin Jiang, Chao Huang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.16851
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-18
|
||||
updated_at: 2026-07-18
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- computer-use
|
||||
- memory
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: llm-agent, tool-use
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: llm-agent, tool-use
|
||||
- inferred topics: agent-evaluation, coding-agent, computer-use, memory, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.16851
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: Environment-free Synthetic Data Generation for API-Calling Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Environment-free Synthetic Data Generation for API-Calling Agents
|
||||
authors: Seanie Lee, Sanjoy Chowdhury, Chao Jiang, Cheng-Yu Hsieh, Ting-Yao Hu, Alexander T Toshev, Oncel Tuzel, Raviteja Vemulapalli
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.16900
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-21
|
||||
updated_at: 2026-07-21
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- memory
|
||||
- rag
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: llm-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: llm-agent
|
||||
- inferred topics: agent-evaluation, memory, rag, tool-use, workflow-agent, world-model
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.16900
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding"
|
||||
authors: Yaohan Yang, Minglei Shi, Borui Zhang, Jie Zhou, Jiwen Lu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.17050
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-18
|
||||
updated_at: 2026-07-18
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- planning
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.CV
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 16
|
||||
collection_queries: agent-evaluation, web-gui-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-evaluation, web-gui-agent
|
||||
- inferred topics: agent-evaluation, computer-use, planning, rag, tool-use
|
||||
- arXiv categories: cs.CV
|
||||
- collection score: 16
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.17050
|
||||
@@ -0,0 +1,60 @@
|
||||
# Paper: A Diagnostic Framework for AI Agent Behavior
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: A Diagnostic Framework for AI Agent Behavior
|
||||
authors: Xichen Zhang, Yingjie Zhang, Tianshu Sun
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.17149
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-19
|
||||
updated_at: 2026-07-19
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- memory
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: ai-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: ai-agent
|
||||
- inferred topics: agent-evaluation, memory, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.17149
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Paper: SAGA: Synthetic Agentic Graph Architecture for Temporal Benchmark Generation
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "SAGA: Synthetic Agentic Graph Architecture for Temporal Benchmark Generation"
|
||||
authors: Jiacheng Ding, Xiaofei Zhang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.17288
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-19
|
||||
updated_at: 2026-07-19
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- rag
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.DB
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 18
|
||||
collection_queries: llm-agent, rag-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: llm-agent, rag-agent
|
||||
- inferred topics: agent-evaluation, agent-safety, rag
|
||||
- arXiv categories: cs.DB
|
||||
- collection score: 18
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.17288
|
||||
+64
@@ -0,0 +1,64 @@
|
||||
# Paper: Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning"
|
||||
authors: Zhihao Liu, Tianyu Wang, Xi Vincent Wang, Lihui Wang
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.17331
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-19
|
||||
updated_at: 2026-07-19
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- multi-agent
|
||||
- planning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 24
|
||||
collection_queries: llm-agent, multi-agent-llm, planning-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: llm-agent, multi-agent-llm, planning-agent
|
||||
- inferred topics: agent-evaluation, agent-safety, multi-agent, planning, tool-use, workflow-agent, world-model
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 24
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.17331
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Paper: Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions
|
||||
authors: Chen Xia, Zexi Kuang, Yuqing Hu
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.17437
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-19
|
||||
updated_at: 2026-07-19
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- memory
|
||||
- planning
|
||||
- rag
|
||||
- reasoning
|
||||
- world-model
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 19
|
||||
collection_queries: llm-agent, planning-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: llm-agent, planning-agent
|
||||
- inferred topics: agent-evaluation, memory, planning, rag, reasoning, world-model
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 19
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.17437
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Sidekick: Designing Communication for Effective Multitasking with Computer Use Agents
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Sidekick: Designing Communication for Effective Multitasking with Computer Use Agents"
|
||||
authors: Ruei-Che Chang, Wenqian Xu, Dingzeyu Li, Bryan Wang, Anhong Guo
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.17527
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-20
|
||||
updated_at: 2026-07-20
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- computer-use
|
||||
- multi-agent
|
||||
- planning
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.HC
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 13
|
||||
collection_queries: web-gui-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: web-gui-agent
|
||||
- inferred topics: computer-use, multi-agent, planning, reasoning, tool-use
|
||||
- arXiv categories: cs.HC
|
||||
- collection score: 13
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.17527
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows
|
||||
authors: Jinyuan Deng, Zhengrui Chen, Xufeng Wei, Tianyu Xing, Chenyi Wen, Qi Sun, Cheng Zhuo
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.17528
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-22
|
||||
updated_at: 2026-07-22
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- coding-agent
|
||||
- planning
|
||||
- tool-use
|
||||
- workflow-agent
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 23
|
||||
collection_queries: ai-agent, coding-agent, llm-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: ai-agent, coding-agent, llm-agent
|
||||
- inferred topics: agent-evaluation, coding-agent, planning, tool-use, workflow-agent
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 23
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.17528
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory
|
||||
authors: Qingcan Kang, Mingyang Liu, Shixiong Kai, Kaichao Liang, Zhentao Tang, Yuqi Cui, Tao Zhong, Mingxuan Yuan
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.17545
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-20
|
||||
updated_at: 2026-07-20
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- agent-safety
|
||||
- memory
|
||||
- rag
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 20
|
||||
collection_queries: agent-memory, language-agent
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-memory, language-agent
|
||||
- inferred topics: agent-evaluation, agent-safety, memory, rag, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 20
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.17545
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Mechanistic Attention Guidance for Agent Memory Refinement
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: Mechanistic Attention Guidance for Agent Memory Refinement
|
||||
authors: Yechao Hong, Haiquan Qiu, Yaqing Wang, Quanming Yao
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.17621
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-20
|
||||
updated_at: 2026-07-20
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- computer-use
|
||||
- memory
|
||||
- rag
|
||||
- reasoning
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: agent-memory
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: agent-memory
|
||||
- inferred topics: agent-evaluation, computer-use, memory, rag, reasoning
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.17621
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
# Paper: Exploratory and Assimilating Reflection: Reflective Recall Cycle for Long-term Memory
|
||||
|
||||
---
|
||||
type: paper
|
||||
title: "Exploratory and Assimilating Reflection: Reflective Recall Cycle for Long-term Memory"
|
||||
authors: Ganesh Senrayan, Moyuru Yamada, Ishan Jindal, Kiran Purohit
|
||||
year: 2026
|
||||
venue: arXiv
|
||||
url: https://arxiv.org/abs/2607.17879
|
||||
code_url:
|
||||
source: arxiv
|
||||
collected_at: 2026-07-27
|
||||
published_at: 2026-07-20
|
||||
updated_at: 2026-07-20
|
||||
status: queued
|
||||
relevance: high
|
||||
topics:
|
||||
- agent-evaluation
|
||||
- memory
|
||||
- rag
|
||||
- reasoning
|
||||
- tool-use
|
||||
methods:
|
||||
-
|
||||
benchmarks:
|
||||
-
|
||||
models:
|
||||
-
|
||||
datasets:
|
||||
- cs.AI
|
||||
related_concepts:
|
||||
-
|
||||
related_jobs:
|
||||
-
|
||||
related_experiments:
|
||||
-
|
||||
related_projects:
|
||||
-
|
||||
collection_score: 14
|
||||
collection_queries: autonomous-agent-llm
|
||||
---
|
||||
|
||||
## One-line Takeaway
|
||||
|
||||
Auto-collected from arXiv because it matched the Agent collection queries. Needs human skim.
|
||||
|
||||
## Why Collected
|
||||
|
||||
- matched queries: autonomous-agent-llm
|
||||
- inferred topics: agent-evaluation, memory, rag, reasoning, tool-use
|
||||
- arXiv categories: cs.AI
|
||||
- collection score: 14
|
||||
|
||||
## Review Checklist
|
||||
|
||||
- Does this paper directly inform Agent architecture, evaluation, memory, tools, safety, coding agents, GUI/browser agents, or multi-agent workflows?
|
||||
- Does it include a benchmark, dataset, code, or reproducible experimental setup?
|
||||
- Should it be promoted from `queued` to `skimmed` or `summarized`?
|
||||
|
||||
## Links
|
||||
|
||||
- arXiv: https://arxiv.org/abs/2607.17879
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user