破解合成图像检测的黑箱机制,发现频域峰值并非关键依据
Beyond Spectral Peaks: Interpreting the Cues Behind Synthetic Image Detection
- 通过移除图像频域峰值,测试多个检测器性能变化
- 多数深度检测器对峰值移除不敏感,说明依赖关系薄弱
- 提出线性可解释基准模型,助力透明化图像取证
多年来,学术界提出了多种基于深度学习的生成图像检测方法以应对生成式AI风险。近期,频域特征(尤其是幅度谱中的周期性峰值)受到广泛关注,常被视为合成图像的强指示信号。然而,当前最先进的检测器多为黑箱模型,其实际是否依赖这些峰值仍不明确,限制了结果的可解释性与可信度。本文开展系统性研究,提出一种移除图像频域峰值的方法,并分析该操作对多个检测器的影响。此外,设计了一种仅依赖频域峰值的简单线性检测器,作为无深度学习干扰的可解释基线。研究发现,多数检测器并不真正依赖频域峰值,挑战了领域内普遍假设,为构建更透明、可靠的图像取证工具指明方向。
原文摘要 · Abstract (English)
Over the years, the forensics community has proposed several deep learning-based detectors to mitigate the risks of generative AI. Recently, frequency-domain artifacts (particularly periodic peaks in the magnitude spectrum), have received significant attention, as they have been often considered a strong indicator of synthetic image generation. However, state-of-the-art detectors are typically used as black-boxes, and it still remains unclear whether they truly rely on these peaks. This limits their interpretability and trust. In this work, we conduct a systematic study to address this question. We propose a strategy to remove spectral peaks from images and analyze the impact of this operation on several detectors. In addition, we introduce a simple linear detector that relies exclusively on frequency peaks, providing a fully interpretable baseline free from the confounding influence of deep learning. Our findings reveal that most detectors are not fundamentally dependent on spectral peaks, challenging a widespread assumption in the field and paving the way for more transparent and reliable forensic tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。