arXiv:2607.15715cs.AI2026-07

对比固定流程与反思型智能体,发现动态工具选择能显著提升信息抽取稳定性。

Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents

论文配图:Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents
图 1 · 摘自论文原文
  • 用反思和记忆机制动态调整信息抽取流程
  • 优化后版本(S2)失败率降低37%,运行时长减少28%
  • 适合关注智能体行为可解释性与鲁棒性的研究者

大型语言模型代理被广泛用于复杂的信息抽取任务,但其代理特性(如反思、记忆)是否带来可观测且可控的性能提升尚不明确。本文以学术论文数据集提取为场景,要求系统从学术PDF中识别提及的数据集并生成结构化记录。通过对比固定工作流基线与多种反思型代理变体,提出一种优化代理条件(S2),在相同任务基础上引入更丰富的PDF处理工具及动态工具选择策略。评估聚焦于过程行为——包括工具执行、重试次数、反思频率、记忆使用、运行时间与故障恢复能力,将提取覆盖率与字段完整性视为次要指标。研究揭示了代理机制如何改变系统行为,验证其对任务完成度的改善效果,并基于观察到的失败模式,提出在相同评测框架下更具鲁棒性的代理设计方案。

原文摘要 · Abstract (English)

Large language model (LLM) agents are increasingly used for complex information-extraction tasks, yet it remains unclear whether agentic components such as reflection and memory lead to observable and controllable improvements over fixed LLM workflows. We study this question through conference-paper dataset extraction, where a system must identify datasets mentioned in scholarly PDFs and produce structured records. We compare a fixed workflow baseline with reflective agent variants and specify an optimized agent condition (S2) that extends the same task with richer PDF tools and dynamic tool selection. Our evaluation emphasizes process-level behavior--including tool execution, retries, reflection, memory use, runtime, and failure recovery--while treating extraction coverage and field completeness as secondary outcome measures. The paper characterizes when agentic mechanisms change system behavior, whether these changes improve task completion, and how the observed failure modes motivate an optimized agent design under the same evaluation harness.

智能体信息抽取反思机制可控制性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。