提出DeepStress框架,系统测试搜索代理在劣质证据下的鲁棒性。
DeepStress: Stress-Testing Deep Search Agents

- 用可控合成环境替代检索模块,精准调控证据可信度、相关性和事实性。
- 在HotpotQA和BrowseCompPlus上发现不同代理对不可靠信息的处理能力差异显著。
- 提出新评估指标,揭示参数知识与检索知识间的冲突交互机制。
尽管搜索代理在多步问答中表现优异,但其对低质量证据的鲁棒性仍缺乏深入研究。这一现象在真实基准中罕见,却可能在实际应用中导致严重失败。为此,本文提出DeepStress,一种通过用受控合成环境替换搜索代理的检索模块,以调节挑战性证据出现频率的压力测试框架。该框架可控制影响文档可靠性的三个维度:可信度、相关性和事实性。在HotpotQA和BrowseCompPlus上测试多个搜索代理,结果表明代理在处理不可靠信息方面存在显著差异,并提出新的评估指标,更准确反映系统表现及参数知识与检索知识之间的冲突交互。
原文摘要 · Abstract (English)
While search agents demonstrate impressive capabilities in multi-step question answering, their robustness to poor-quality evidence remains under-explored. This phenomenon occurs rarely in realistic benchmarks but can lead to dramatic failure in real life applications. Therefore in this study we propose DeepStress, a stress testing framework that controls the frequency of challenging evidence by replacing the retrieval module of search agents with a controlled synthetic environment. We use this framework to control three dimensions that can affect document reliability: trustworthiness, relevance, and factuality. Testing several search agents on HotpotQA and BrowseCompPlus, we demonstrate that agents exhibit substantial differences in their ability to handle unreliable information and propose new metrics that better document systems outcomes as well as the interactions between conflicting parametric and retrieved knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。