搜索代理常误判任务完成,新框架可识别并纠正其错误信念。
When Is Enough Not Enough? Illusory Completion in Search Agents
- 引入认知账本追踪每个约束的证据与代理信念
- 发现四类常见失败模式,如忽略反例和过早退出
- 实时约束追踪使准确率提升11.6%,漏检减少26.5%
近期搜索代理通过多轮推理和搜索工具在多跳与长时序基准上表现优异,但其是否能可靠地跨所有要求进行推理仍不明确。本文研究多约束问题下代理对多个条件的跟踪、验证与维持能力。结果发现,幻觉完成现象频繁出现:代理虽未解决或违反某些约束,却误认为任务已完成,导致答案验证不足。为此,我们提出认知账本(Epistemic Ledger)评估框架,全程追踪每条约束的证据支持与代理信念。分析揭示四种典型失败模式:空洞断言、忽略反证、推理停滞和提前退出。基于此,我们设计LiveLedger——执行时的显式约束状态追踪机制。该简单干预显著提升性能,在多约束问题上使未验证答案减少最多26.5%,整体准确率最高提升11.6%。
原文摘要 · Abstract (English)
Recent search agents leverage multi-turn reasoning and search tools to achieve strong performance on multi-hop and long-horizon benchmarks. Yet it remains unclear whether they reliably reason across all requirements by tracking, verifying, and maintaining multiple conditions in these questions. We study this capability under multi-constraint problems, where valid answers must satisfy several constraints simultaneously. We find that illusory completion frequently occurs, wherein agents believe tasks are complete despite unresolved or violated constraints, leading to underverified answers. To diagnose this behavior, we introduce the Epistemic Ledger, an evaluation framework that tracks evidential support and agents' beliefs for each constraint throughout multi-turn reasoning. Our analysis reveals four recurring failure patterns: bare assertions, overlooked refutations, stagnation, and premature exit. Motivated by these findings, we examine whether explicit constraint-state tracking during execution mitigates these failures via LiveLedger, an inference-time tracker. This simple intervention consistently improves performance, substantially reducing underverified answers (by up to 26.5%) and improving overall accuracy (by up to 11.6%) on multi-constraint problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。