揭秘伪造事实检测系统的攻击手段,揭示其脆弱性与防御方向
Adversarial Attacks Against Automated Fact-Checking: A Survey
- 系统性梳理针对自动事实核查的对抗攻击方法
- 指出现有模型在误导性内容面前易被欺骗的缺陷
- 适合关注AI安全与信息可信度的研究者阅读
在虚假信息广泛传播的时代,事实核查(FC)在验证声明和促进可靠信息方面发挥着关键作用。尽管自动化事实核查(AFC)已取得显著进展,但现有系统仍易受对抗攻击影响,攻击者可通过操纵或生成声明、证据或声明-证据配对来扭曲真相,误导决策者,最终削弱核查模型的可靠性。尽管学术界对针对AFC系统的对抗攻击兴趣日益增长,但缺乏对核心挑战的全面综合概述,包括理解攻击策略、评估现有模型的鲁棒性以及探索增强鲁棒性的途径。本文首次深入回顾了针对事实核查的对抗攻击,分类总结了现有攻击方法并评估其对AFC系统的影响。同时,我们考察了近期具备对抗意识的防御技术,并指出了亟待进一步研究的开放问题。研究结果强调了构建能够抵御对抗操纵的稳健事实核查框架的紧迫性,以确保高精度验证。
原文摘要 · Abstract (English)
In an era where misinformation spreads freely, fact-checking (FC) plays a crucial role in verifying claims and promoting reliable information. While automated fact-checking (AFC) has advanced significantly, existing systems remain vulnerable to adversarial attacks that manipulate or generate claims, evidence, or claim-evidence pairs. These attacks can distort the truth, mislead decision-makers, and ultimately undermine the reliability of FC models. Despite growing research interest in adversarial attacks against AFC systems, a comprehensive, holistic overview of key challenges remains lacking. These challenges include understanding attack strategies, assessing the resilience of current models, and identifying ways to enhance robustness. This survey provides the first in-depth review of adversarial attacks targeting FC, categorizing existing attack methodologies and evaluating their impact on AFC systems. Additionally, we examine recent advancements in adversary-aware defenses and highlight open research questions that require further exploration. Our findings underscore the urgent need for resilient FC frameworks capable of withstanding adversarial manipulations in pursuit of preserving high verification accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。