测试大模型在虚假信息干扰下的推理能力,发现其易被误导而放弃证据链验证。
DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments

- 设计100个任务,模拟真实网页中混入看似合理的错误信息。
- 模型准确率下降66-88个百分点,因停止证据链验证而采信错误答案。
- 适合关注AI研究可信性、开放网络应用的研究者和开发者。
深度研究代理在开放网络上运行时,常面临冗余摘要、过时报告和误导性文档的干扰。现有评估难以揭示代理在看似正常但故意植入的虚假文档环境下,是否仍能保持严谨的证据标准。我们提出DRNOISE,一个包含100个任务的基准测试,用于评估在误导性证据下的答案恢复能力。每个任务都有唯一正确答案,并由两条间接记录链支持;在噪声条件下,新增一条看似合理且直接给出矛盾答案的文档。该基准涵盖十类证据操作。尽管多个代理在干净任务上表现良好,单次干预导致其准确率下降66-88个百分点。追踪分析表明,验证惰性是主要失败模式:代理虽检索到真实记录,却未完成并整合证据链,转而依赖看似答案的文档。通用验证提示可缓解但无法消除差距。该场景尤其贴近开放网络部署实际,因误导性内容常以普通页面形式出现,而非显式攻击。因此,可靠深度研究不仅需要检索与引用,更需主动比对直接声明与底层证据。
原文摘要 · Abstract (English)
Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents. Existing evaluations offer limited insight into whether agents preserve sound evidential standards when an ordinary-looking false document is deliberately seeded into a searchable environment and offers a direct shortcut to a conflicting answer. We introduce DRNOISE, a 100-task benchmark for answer recovery under misleading evidence. Each task has a unique gold answer supported by two corroborating indirect record chains; the paired noisy condition adds one plausible document that states a conflicting answer directly. The benchmark spans ten families of evidence operations. Across agents with strong clean-task performance, this single intervention causes 66-88 percentage-point accuracy drops. Trace analyses identify verification inertia as the dominant failure mode: agents often retrieve truthful records but stop before completing and reconciling the evidence chain, instead deferring to the answer-like document. Generic verification prompts reduce but do not close this gap. The setting is especially relevant to open-web deployment, where plausible falsehoods arrive through ordinary-looking pages rather than explicit attacks. Reliable deep research therefore requires more than retrieval and citation; it requires active reconciliation of direct claims with record-level evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。