提出新型对抗攻击框架,让假话骗过搜索型大模型事实核查系统。
DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems
- 设计代理式攻击框架,仅通过输入误导搜索与推理流程。
- 在真实系统上使准确率从78.7%降至53.7%,性能大幅下降。
- 攻击具备跨系统迁移能力,适合研究模型鲁棒性者参考。
基于搜索的大型语言模型(LLM)事实核查系统在动态获取外部证据方面展现出强大潜力。然而,这类系统在对抗攻击下的鲁棒性尚未充分理解。本文在仅输入威胁模型下,研究针对搜索型LLM事实核查系统的对抗性陈述攻击。我们提出DECEIVE-AFC,一种基于代理的对抗攻击框架,融合新颖的陈述级攻击策略与对抗性陈述有效性评估原则。该框架系统探索干扰搜索行为、证据检索及LLM推理的对抗攻击路径,无需访问证据源或模型内部信息。在基准数据集与真实系统上的广泛评估表明,我们的攻击显著降低验证性能,使准确率从78.7%降至53.7%,且显著优于现有基于陈述的攻击基线,具备强跨系统迁移能力。
原文摘要 · Abstract (English)
Fact-checking systems with search-enabled large language models (LLMs) have shown strong potential for verifying claims by dynamically retrieving external evidence. However, the robustness of such systems against adversarial attack remains insufficiently understood. In this work, we study adversarial claim attacks against search-enabled LLM-based fact-checking systems under a realistic input-only threat model. We propose DECEIVE-AFC, an agent-based adversarial attack framework that integrates novel claim-level attack strategies and adversarial claim validity evaluation principles. DECEIVE-AFC systematically explores adversarial attack trajectories that disrupt search behavior, evidence retrieval, and LLM-based reasoning without relying on access to evidence sources or model internals. Extensive evaluations on benchmark datasets and real-world systems demonstrate that our attacks substantially degrade verification performance, reducing accuracy from 78.7% to 53.7%, and significantly outperform existing claim-based attack baselines with strong cross-system transferability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。