arXiv:2603.00801cs.AIcs.IR2026-03被引 3

构建虚拟网络测试语言模型在虚假信息攻击下的可靠性。

The Synthetic Web: Adversarially-Curated Mini-Internets for Diagnosing Epistemic Weaknesses of Language Agents

  • 用程序生成含真实标签的微型互联网,模拟虚假信息干扰。
  • 六种前沿模型受骗后准确率骤降,且无法通过多查来纠正。
  • 适合研究模型抗操纵能力或高风险场景部署的学者使用。

语言代理日益作为具备网页搜索、浏览和信息整合能力的系统运行,但其面对不可靠或对抗性内容时的鲁棒性仍不明确。现有基准评估功能导航或静态事实性,无法因果分离对抗性排名带来的脆弱性,当前检索增强生成的缓解策略也缺乏在此条件下的验证。本文提出Synthetic Web Benchmark,一个程序化生成的环境,包含数千个带可信度与真实性标签的超链接文章,包含过程级交互轨迹,并通过污染过滤消除训练数据泄露。通过在可控搜索排名中注入单一高可信度虚假文章,我们测量了六种前沿模型在对抗暴露下的因果影响。结果揭示灾难性失败:即便拥有无限真实来源,准确率仍急剧下降,搜索次数极少提升且严重误校准。这些发现暴露了当前前沿模型处理冲突信息的根本缺陷,对高风险领域部署具有直接意义。该基准支持系统分析此类失效模式,并提供对抗性排名下的可控测试平台,填补了研究空白。本工作为开发搜索鲁棒且认知谦逊的语言代理建立了可复现基线。

原文摘要 · Abstract (English)

Language agents increasingly act as web-enabled systems that search, browse, and synthesize information from diverse sources. However, these sources can include unreliable or adversarial content, and the robustness of agents to adversarial ranking - where misleading information appears prominently in search results - remains poorly understood. Existing benchmarks evaluate functional navigation or static factuality but cannot causally isolate this vulnerability, and current mitigation strategies for retrieval-augmented generation remain largely untested under such conditions. We introduce Synthetic Web Benchmark, a procedurally generated environment comprising thousands of hyperlinked articles with ground-truth labels for credibility and factuality, process-level interaction traces, and contamination filtering to eliminate training-data leakage. By injecting a single high-plausibility misinformation article into a controllable search rank, we measure the causal effect of adversarial exposure in six frontier models. The results reveal catastrophic failures: accuracy collapses despite unlimited access to truthful sources, with minimal search escalation and severe miscalibration. These findings expose fundamental limitations in how current frontier models handle conflicting information, with immediate implications for deployment in high-stakes domains. Our benchmark enables systematic analysis of these failure modes and provides a controlled testbed for evaluating mitigation strategies under adversarial ranking - a gap in current research. This work establishes a reproducible baseline for developing search-robust and epistemically humble agents capable of resisting manipulation in high-stakes domains.

语言模型对抗攻击虚假信息评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。