arXiv:2511.17766cs.CV2025-11被引 3

用视觉变压器检测AI生成的卫星图像,准确率超95%。

Deepfake Geography: Detecting AI-Generated Satellite Images

  • 对比CNN与视觉变压器,发现ViT更擅长捕捉图像全局结构
  • 在13万张图像上,ViT准确率达95.11%,远超CNN的87.02%
  • 通过注意力热图揭示检测机制,提升模型可信度

生成模型如StyleGAN2和Stable Diffusion的快速发展对卫星影像的真实性构成日益严重的威胁,而其在科学与安全领域的应用愈发关键。尽管深度伪造检测在人脸领域已有广泛研究,但卫星影像面临地形不一致与结构伪影等独特挑战。本研究系统比较了卷积神经网络(CNN)与视觉变压器(ViT)在检测AI生成卫星图像上的表现。基于来自DM-AER和FSI数据集的超过13万张标注RGB图像,结果表明,ViT在准确率(95.11% vs. 87.02%)和整体鲁棒性方面显著优于CNN,归因于其建模长距离依赖与全局语义结构的能力。我们进一步采用针对架构的可解释性方法(如CNN的Grad-CAM与ViT的Chefer注意力归因),揭示了两者不同的检测行为,验证了模型可信度。结果强调了ViT在识别合成图像中结构性不一致与重复纹理模式方面的优势。未来工作将拓展至多光谱与合成孔径雷达(SAR)模态,并引入频域分析以进一步强化检测能力,保障高风险应用场景下的卫星影像完整性。

原文摘要 · Abstract (English)

The rapid advancement of generative models such as StyleGAN2 and Stable Diffusion poses a growing threat to the authenticity of satellite imagery, which is increasingly vital for reliable analysis and decision-making across scientific and security domains. While deepfake detection has been extensively studied in facial contexts, satellite imagery presents distinct challenges, including terrain-level inconsistencies and structural artifacts. In this study, we conduct a comprehensive comparison between Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) for detecting AI-generated satellite images. Using a curated dataset of over 130,000 labeled RGB images from the DM-AER and FSI datasets, we show that ViTs significantly outperform CNNs in both accuracy (95.11 percent vs. 87.02 percent) and overall robustness, owing to their ability to model long-range dependencies and global semantic structures. We further enhance model transparency using architecture-specific interpretability methods, including Grad-CAM for CNNs and Chefer's attention attribution for ViTs, revealing distinct detection behaviors and validating model trustworthiness. Our results highlight the ViT's superior performance in detecting structural inconsistencies and repetitive textural patterns characteristic of synthetic imagery. Future work will extend this research to multispectral and SAR modalities and integrate frequency-domain analysis to further strengthen detection capabilities and safeguard satellite imagery integrity in high-stakes applications.

深度伪造卫星图像视觉变压器可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。