攻击智能事实核查系统,用伪造证据误导其判断。
Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System
- 利用大模型模仿系统分解逻辑,生成针对性恶意证据。
- 在不同攻击预算下,成功率比现有方法高8.9%至21.2%。
- 揭示当前事实核查系统的安全漏洞,适合安全研究者关注。
最先进的事实核查系统通过基于大语言模型(LLM)的自主代理,将复杂声明拆解为子声明,逐个验证后整合结果并给出理由。这类系统的安全性至关重要,一旦被攻破可能加剧虚假信息传播,但相关研究仍不足。本文提出一种新型威胁模型,并设计首个针对此类系统的投毒攻击框架——Fact2Fiction。该框架利用大模型模拟系统的分解策略,结合系统生成的解释文本,定制恶意证据以破坏子声明验证。大量实验表明,Fact2Fiction在不同投毒预算下均实现8.9%至21.2%更高的攻击成功率,暴露了现有事实核查系统中的安全缺陷,凸显了防御机制的必要性。
原文摘要 · Abstract (English)
State-of-the-art (SOTA) fact-checking systems combat misinformation by employing autonomous LLM-based agents to decompose complex claims into smaller sub-claims, verify each sub-claim individually, and aggregate the partial results to produce verdicts with justifications (explanations for the verdicts). The security of these systems is crucial, as compromised fact-checkers can amplify misinformation, but remains largely underexplored. To bridge this gap, this work introduces a novel threat model against such fact-checking systems and presents \textsc{Fact2Fiction}, the first poisoning attack framework targeting SOTA agentic fact-checking systems. Fact2Fiction employs LLMs to mimic the decomposition strategy and exploit system-generated justifications to craft tailored malicious evidences that compromise sub-claim verification. Extensive experiments demonstrate that Fact2Fiction achieves 8.9\%--21.2\% higher attack success rates than SOTA attacks across various poisoning budgets and exposes security weaknesses in existing fact-checking systems, highlighting the need for defensive countermeasures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。