通过双频域与滑动窗口重构,提升对生成图像的通用检测能力
Dual Frequency Branch Framework with Reconstructed Sliding Windows Attention for AI-Generated Image Detection
- 设计滑动窗口重构注意力,捕捉局部区域内部元素关系
- 构建双频域分支框架,融合小波与傅里叶相位信息
- 在65种生成模型上实现2.13%准确率提升,适合对抗新型伪造
生成对抗网络(GANs)和扩散模型的快速发展使得合成图像愈发逼真,带来误导与欺骗等社会风险。因此,检测AI生成图像成为关键挑战。现有方法虽注重细粒度特征提取以增强泛化能力,但常忽略局部区域内元素的重要性与相互依赖性,且仅限于单一频率域,难以捕捉普遍的伪造痕迹。为此,本文首先采用滑动窗口限制注意力范围,并重构窗口内特征以建模局部元素间关系。随后设计包含四个DWT子带与FFT相位部分的双频域分支框架,从多角度丰富局部伪造特征提取。通过双频域特征增强与重构滑动窗口注意力的细粒度提取,本方法在多种生成图像上表现出更优泛化能力。在涵盖65种不同生成模型的数据集上,检测准确率相较当前最优方法提升2.13%。
原文摘要 · Abstract (English)
The rapid advancement of Generative Adversarial Networks (GANs) and diffusion models has enabled the creation of highly realistic synthetic images, presenting significant societal risks, such as misinformation and deception. As a result, detecting AI-generated images has emerged as a critical challenge. Existing researches emphasize extracting fine-grained features to enhance detector generalization, yet they often lack consideration for the importance and interdependencies of internal elements within local regions and are limited to a single frequency domain, hindering the capture of general forgery traces. To overcome the aforementioned limitations, we first utilize a sliding window to restrict the attention mechanism to a local window, and reconstruct the features within the window to model the relationships between neighboring internal elements within the local region. Then, we design a dual frequency domain branch framework consisting of four frequency domain subbands of DWT and the phase part of FFT to enrich the extraction of local forgery features from different perspectives. Through feature enrichment of dual frequency domain branches and fine-grained feature extraction of reconstruction sliding window attention, our method achieves superior generalization detection capabilities on both GAN and diffusion model-based generative images. Evaluated on diverse datasets comprising images from 65 distinct generative models, our approach achieves a 2.13\% improvement in detection accuracy over state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。