定位AI研究系统中引用错误的源头,提升报告可信度。
Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research
- 通过局部测试每个代理的输入输出,定位错误产生环节。
- 84.7%的最终错误源于协调器,其中31%为幻觉,其余为引用错误。
- 发现不同代理错误类型有规律可循,适合调试和优化多代理系统。
深度研究(DR)系统通过协调多个代理从网络搜索并合成信息生成长篇带引文的报告。引文是评估报告可信度的主要手段,但现有系统引文召回率低下。由于DR系统是复杂的多代理架构,信息在代理间传递如同传话游戏,内容与引文均可能被扭曲。本文提出一种评估方法,通过局部测试代理调用对自身输入的忠实性与可验证性,精确定位错误来源。此外,我们构建了四类错误分类体系:幻觉、未引用输入依赖、未引用输出、引文不足。在三个顶尖开源DR系统上应用该方法,获得可操作的诊断结果。几乎每个代理都犯了大量错误,唯独单文档摘要代理表现较好。错误类型在不同代理间系统性差异明显,协调器错误以引用问题为主。在AI-Q中,84.7%的最终报告错误源自协调器,约31%为幻觉,其余为引用错误。基于这些洞见,我们证明两种简单干预可提升引文召回率5%,且不降低输出质量。
原文摘要 · Abstract (English)
Deep research (DR) systems produce long-form cited reports by orchestrating multiple agents that search and synthesize information from the web. Citations are the primary mechanism for evaluating the faithfulness of these reports, yet current DR systems exhibit poor citation recall. Moreover, improving citation recall is challenging because DR systems are complex multi-agent architectures where information passes through agents like a telephone game, and both content and citations can get corrupted along the way. We propose an evaluation method that pinpoints which agent introduced each error by locally testing agent invocations for faithfulness and verifiability relative to their own inputs. Furthermore, we propose a four-type taxonomy to categorize the discovered errors: hallucination, uncited input reliance, uncited output, or insufficient citations. Applying our method to three top-ranked open-source DR systems, we obtain actionable diagnostics. Almost every agent makes a lot of mistakes with the exception being those that summarize a single document. We find that the dominant error type varies systematically across agents, where the orchestrator mistakes are mostly citation-related. We find that 84.7% of final-report errors in AI-Q originate at the orchestrator, roughly 31% of them hallucinations and the rest citation mistakes. Guided by these insights, we demonstrate that two simple interventions raise citation recall by 5% without degrading output quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。