arXiv:2503.08484cs.CV2025-03被引 8

利用频谱分形自相似性,实现跨生成模型的AI图像检测。

Generalizable AI-Generated Image Detection Based on Fractal Self-Similarity in the Spectrum

  • 通过频谱中的分形自相似结构捕捉生成图像共性特征。
  • 在16种不同生成器上平均检测准确率达93.93%。
  • 无需依赖特定生成器的伪影,适合检测未知模型生成图像。

随着图像合成技术的快速发展,AI生成图像愈发逼真,其滥用风险上升,亟需可靠检测手段。然而,生成模型多样性使得检测器难以泛化至未见模型。现有方法多依赖特定生成器的伪影,限制了泛化能力。本文研究生成过程本身的结构特性:图像生成本质上是从紧凑表示构建空间丰富内容,并保持不同空间位置的语义结构一致性。我们通过维度增加的平移等变变换形式化这一特性,证明其在傅里叶谱中诱发自相似结构,并在多阶段生成中递归传播,形成分层分形自相似模式。不同频谱子区域间存在由生成过程继承的一致结构对应关系,构成与生成器无关的检测线索。基于此,提出Fractal-CNN,聚焦捕捉频谱自相似性而非特定生成器的频谱值。在多种GAN与扩散模型生成器上的大量实验表明,Fractal-CNN实现强跨生成器泛化,16个测试生成器平均检测准确率为93.93%。

原文摘要 · Abstract (English)

With the rapid development of image synthesis techniques, AI-generated images have become increasingly realistic, which heightens the potential risk associated with their misuse and creates a growing need for reliable detection. However, the growing diversity of generative models makes it increasingly difficult for detectors to generalize to images produced by unseen generators. Most existing methods rely on artifacts associated with specific generators, which limits their generalization to images produced by unseen models. To address this problem, we investigate structural characteristics arising from the image generation process itself. Image generation fundamentally involves constructing spatially rich content from more compact representations, while preserving the semantic identity of structures across different spatial locations. We formalize these properties through dimension-increasing shift-equivariant transformations and show that such transformations induce a self-similar structure in the Fourier spectrum. Across successive generation stages, this structure can propagate recursively and form a hierarchical fractal self-similar pattern. Consequently, different spectral sub-regions exhibit consistent structural correspondences inherited from the generation process, providing a generator-agnostic cue for detection. Based on this observation, we propose Fractal-CNN, which captures spectral self-similarity rather than generator-specific spectral values. Extensive experiments across diverse GAN- and diffusion-based generators demonstrate that Fractal-CNN achieves strong cross-generator generalization, with an average detection accuracy of 93.93% across 16 test generators.

图像检测分形结构频谱分析生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。