arXiv:2607.20730cs.CRcs.AI2026-07

提出GPE框架,评估大模型在可控伪造证据下的事实核查鲁棒性。

GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-Style Poisoning

论文配图:GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-Style Poisoning
图 1 · 摘自论文原文
  • 构建多领域事实核查数据集,可控制证据来源与污染比例。
  • 实验发现纯净化评估无法暴露对抗环境中的性能下降和效率权衡。
  • 适合研究大模型安全、对抗样本防御与可信信息验证的学者。

大型语言模型越来越多地依赖搜索工具获取最新信息,这带来了新的攻击面:被检索的文档可能被操控。这一风险因生成式引擎优化(GEO)的发展而加剧,该技术可使特定内容更易被检索、引用并被模型采纳。现有事实核查基准和评估框架缺乏可控证据环境,难以评估模型在GEO污染下的鲁棒性。为此,我们提出GPE,包含一个多领域事实核查基准和一个可控制证据源与污染比例的评估框架。在多种验证方法和污染攻击下的实验表明,GPE揭示了仅通过纯净评估无法观察到的鲁棒性退化与效率权衡,证实了在对抗性证据环境下评估事实核查的必要性。

原文摘要 · Abstract (English)

Large language models increasingly use search tools to retrieve up-to-date information, introducing a new attack surface in which retrieved documents can be manipulated. This risk is amplified by the development of generative engine optimization, which can make selected content more likely to be retrieved, cited, and adopted by models. Existing fact-verification benchmarks and evaluation frameworks do not provide the controlled evidence environments needed to assess robustness against GEO poisoning. We therefore propose GPE, which consists of a multi-domain fact-verification benchmark and an evaluation framework for controlling evidence sources and poisoning ratios. Experiments across multiple verification methods and poisoning attacks demonstrate that GPE exposes robustness degradation and efficiency trade-offs that cannot be observed through clean evaluation alone, confirming the need to evaluate fact verification under adversarial evidence environments.

事实核查对抗攻击大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。