用大模型生成说服性伪证,让事实核查系统失效
LLM-Based Adversarial Persuasion Attacks on Fact-Checking Systems
- 用大模型结合15种说服技巧重写假消息
- 在FEVER和FEVEROUS上使核查准确率下降超30%
- 适合研究虚假信息防御的学者和安全团队
自动化事实核查(AFC)系统易受对抗攻击,导致虚假声明逃过检测。现有框架多依赖注入噪声或语义修改,但尚未利用说服技巧这一对抗潜力,而说服技巧广泛用于误导性宣传。本文提出一种新型基于大模型的说服性对抗攻击,通过15种归入5类的说服技术重写声明。采用解耦评估策略,研究其对声明验证与证据检索的影响。在FEVER和FEVEROUS基准上的实验表明,此类攻击显著降低验证性能与证据检索效果。分析揭示说服技巧是一类强大的对抗攻击手段,凸显构建更鲁棒事实核查系统的必要性。
原文摘要 · Abstract (English)
Automated fact-checking (AFC) systems are susceptible to adversarial attacks, enabling false claims to evade detection. Existing adversarial frameworks typically rely on injecting noise or altering semantics, yet no existing framework exploits the adversarial potential of persuasion techniques against AFC systems, which are widely used in disinformation campaigns to manipulate audiences. In this paper, we introduce a novel class of persuasive adversarial attacks on AFCs by employing an LLM to rephrase claims using persuasion techniques. Considering $15$ techniques grouped into $5$ categories, we study the effects of persuasion on both claim verification and evidence retrieval using a decoupled evaluation strategy. Experiments on the FEVER and FEVEROUS benchmarks show that persuasion attacks can substantially degrade both verification performance and evidence retrieval. Our analysis identifies persuasion techniques as a potent class of adversarial attacks, highlighting the need for more robust AFC systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。