arXiv:2607.20891cs.AI2026-07被引 1

测试大模型研究中假信息如何误导结论,发现仅一个假文档就让错误采纳率升至54.7%。

Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions

论文配图:Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions
图 1 · 摘自论文原文
  • 构建含虚假结论的误导性文档集,模拟真实研究中的可信假信息。
  • 引入一个假文档使错误结论采纳率从0%升至54.7%,且受来源权威性影响显著。
  • 强调研究中需持续验证中间证据,适合关注大模型可靠性与安全的研究者。

深度研究智能体通过迭代规划、检索证据和生成报告进行长周期调查。然而,它们能否抵御被植入工作流中的看似可信但事实错误的信息仍不明确。为此,我们提出MisKnow-Agent框架,构建支持人工审核后虚假结论的任务特定文档,控制权威线索和来源风格。应用于DeepResearch Bench任务,经筛选生成5,933份误导性文档。评估DeerFlow、WebThinker及Gemini Deep Research三种模型,使用报告级错误结论采纳率(FCAR)衡量仅报告支持虚假结论的比例。在所有配置中,引入一份误导文档使平均FCAR从无注入对照组的0%升至54.7%。FCAR随研究阶段、框架设计、来源权威性和呈现风格变化显著,而搜索结果排名及额外文档数量影响有限。尽管跨模型验证能识别多数误导内容,但智能体仍会采纳相应虚假结论。预研与后研防御可降低但无法消除采纳,提示需在证据进入中间状态和最终合成时持续验证。代码与数据集已公开于https://github.com/whfeLingYu/MisKnow-Agent 和 https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge。

原文摘要 · Abstract (English)

Deep Research agents conduct long-horizon investigations by iteratively planning, retrieving evidence, and generating reports. However, it remains unclear whether they can resist apparently credible but factually false information introduced into these workflows. To study this failure mode, we introduce MisKnow-Agent, a controlled evaluation framework that constructs task-specific documents supporting manually audited false conclusions with controlled authority cues and source styles. Applied to the tasks from DeepResearch Bench, it generates 5,933 misleading documents after filtering. We evaluate DeerFlow and WebThinker with three backbone LLMs, together with Gemini Deep Research, using a report-level false-conclusion adoption rate (FCAR) that counts only reports endorsing the false conclusion. Across the configurations, introducing one misleading document increases the mean FCAR from 0\% in the no-injection control to 54.7\%. FCAR varies substantially with lifecycle stage and framework design, and also with source authority and presentation style, whereas search-result rank and additional documents beyond the first have limited influence. Although cross-model verification consistently classifies retained instances as misleading, Deep Research agents can still adopt the corresponding false conclusions during long-horizon research. Pre- and post-research defenses reduce FCAR but do not eliminate adoption, motivating continuous verification when evidence enters intermediate research states and final synthesis. To facilitate reproducibility, our code and dataset are publicly available at https://github.com/whfeLingYu/MisKnow-Agent and https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge, respectively.

大模型可靠性虚假信息研究智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。