arXiv:2505.22604cs.CV2025-05被引 3

无需训练即可抵御对抗攻击,提升AI生成图像检测可靠性

Adversarially Robust AI-Generated Image Detection for Free: An Information Theoretic Perspective

  • 基于信息论分析,发现对抗训练导致特征混淆致检测失效
  • 提出TRIM方法,利用预测熵与KL散度衡量特征偏移,保持原精度
  • 无需重新训练,对多种攻击和数据集均显著优于现有防御

人工智能生成图像(AIGI)的快速发展带来了伪造和虚假信息等滥用风险。为此,众多检测方法被提出,但普遍易受对抗攻击影响,而该领域的防御手段仍十分稀缺。本文首次指出,广泛认为最有效的对抗训练(AT)在AIGI检测中会引发性能崩溃。通过信息论视角分析,我们发现崩溃根源在于特征纠缠,破坏了特征与标签间的互信息。相比之下,标准检测器展现出清晰的特征分离。基于此差异,我们提出无需训练的鲁棒检测方法TRIM(Training-free Robust Detection via Information-theoretic Measures),该方法基于标准检测器,通过预测熵和KL散度量化特征偏移。在多个数据集和攻击场景下的实验验证了其优越性:在ProGAN(GenImage)上比当前最优防御高出33.88%(28.91%),同时有效保持原始检测精度。

原文摘要 · Abstract (English)

Rapid advances in Artificial Intelligence Generated Images (AIGI) have facilitated malicious use, such as forgery and misinformation. Therefore, numerous methods have been proposed to detect fake images. Although such detectors have been proven to be universally vulnerable to adversarial attacks, defenses in this field are scarce. In this paper, we first identify that adversarial training (AT), widely regarded as the most effective defense, suffers from performance collapse in AIGI detection. Through an information-theoretic lens, we further attribute the cause of collapse to feature entanglement, which disrupts the preservation of feature-label mutual information. Instead, standard detectors show clear feature separation. Motivated by this difference, we propose Training-free Robust Detection via Information-theoretic Measures (TRIM), the first training-free adversarial defense for AIGI detection. TRIM builds on standard detectors and quantifies feature shifts using prediction entropy and KL divergence. Extensive experiments across multiple datasets and attacks validate the superiority of our TRIM, e.g., outperforming the state-of-the-art defense by 33.88% (28.91%) on ProGAN (GenImage), while well maintaining original accuracy.

图像检测对抗防御信息论无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。