系统梳理生成图像检测新方法,涵盖从空间到多模态的主流技术。
Methods and Trends in Detecting AI-Generated Images: A Comprehensive Review
- 按空间、频域、指纹、补丁等六大框架分类检测方法
- 对比分析在公开数据集上的泛化性与鲁棒性表现
- 推荐融合训练自由与多模态推理的混合框架,提升可解释性
生成模型(如GAN、扩散模型、变分自编码器)的普及推动了高质量多媒体数据的合成,但也引发了对抗攻击、伦理滥用和社 会危害等担忧。为应对这些挑战,研究者们致力于发展高效合成数据检测方法。现有综述多聚焦于深度伪造检测,忽视了近期在合成图像取证中的进展,特别是多模态框架、基于推理的检测和无需训练的方法。本文系统回顾了当前最先进的合成图像检测与分类技术,将核心方法归纳为:空间域、频域、指纹、补丁、无训练及多模态推理六类,并简明阐述其原理。进一步在公开数据集上开展详细对比分析,评估各类方法的泛化能力、鲁棒性与可解释性。最后,指出开放挑战与未来方向,强调结合无训练方法效率与多模态模型语义推理优势的混合框架,有望推动可信且可解释的合成图像取证发展。
原文摘要 · Abstract (English)
The proliferation of generative models, such as Generative Adversarial Networks (GANs), Diffusion Models, and Variational Autoencoders (VAEs), has enabled the synthesis of high-quality multimedia data. However, these advancements have also raised significant concerns regarding adversarial attacks, unethical usage, and societal harm. Recognizing these challenges, researchers have increasingly focused on developing methodologies to detect synthesized data effectively, aiming to mitigate potential risks. Prior reviews have predominantly focused on deepfake detection and often overlook recent advancements in synthetic image forensics, particularly approaches that incorporate multimodal frameworks, reasoning-based detection, and training-free methodologies. To bridge this gap, this survey provides a comprehensive and up-to-date review of state-of-the-art techniques for detecting and classifying synthetic images generated by advanced generative AI models. The review systematically examines core detection paradigms, categorizes them into spatial-domain, frequency-domain, fingerprint-based, patch-based, training-free, and multimodal reasoning-based frameworks, and offers concise descriptions of their underlying principles. We further provide detailed comparative analyses of these methods on publicly available datasets to assess their generalizability, robustness, and interpretability. Finally, the survey highlights open challenges and future directions, emphasizing the potential of hybrid frameworks that combine the efficiency of training-free approaches with the semantic reasoning of multimodal models to advance trustworthy and explainable synthetic image forensics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。