现有合成媒体检测方法泛化能力差,需发展多模态新方案。
Toward Generalized Detection of Synthetic Media: Limitations, Challenges, and the Path to Multimodal Solutions
- 分析24篇论文,发现当前检测方法依赖视觉异常特征。
- 在跨模型、多模态数据上准确率显著下降,泛化能力弱。
- 建议构建多模态深度学习模型,适合安全与可信研究者参考。
过去十年间,人工智能在媒体生成领域快速发展。生成对抗网络(GANs)提升了逼真图像生成质量,扩散模型(Diffusion Models)开启了生成媒体的新纪元。这些进步使真实与合成内容难以区分。深度伪造技术的兴起暴露了其被用于传播虚假信息、政治阴谋、侵犯隐私和欺诈的风险。为此,研究者提出了大量检测模型,主要基于卷积神经网络(CNN)和视觉变换器(ViTs),通过寻找视觉、空间或时间上的异常模式进行识别。然而,这些方法普遍缺乏对未见数据的泛化能力,且在不同生成模型的内容上表现不佳。此外,现有方法在处理多模态数据及高度修改内容时效果有限。本文系统回顾了24项近期关于AI生成媒体检测的研究,逐一分析其贡献与缺陷,总结出当前方法的共性局限与关键挑战。基于此,提出未来应聚焦于多模态深度学习模型的发展方向,这类模型有望实现更鲁棒、更具泛化性的检测能力,为后续研究提供清晰路径。
原文摘要 · Abstract (English)
Artificial intelligence (AI) in media has advanced rapidly over the last decade. The introduction of Generative Adversarial Networks (GANs) improved the quality of photorealistic image generation. Diffusion models later brought a new era of generative media. These advances made it difficult to separate real and synthetic content. The rise of deepfakes demonstrated how these tools could be misused to spread misinformation, political conspiracies, privacy violations, and fraud. For this reason, many detection models have been developed. They often use deep learning methods such as Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs). These models search for visual, spatial, or temporal anomalies. However, such approaches often fail to generalize across unseen data and struggle with content from different models. In addition, existing approaches are ineffective in multimodal data and highly modified content. This study reviews twenty-four recent works on AI-generated media detection. Each study was examined individually to identify its contributions and weaknesses, respectively. The review then summarizes the common limitations and key challenges faced by current approaches. Based on this analysis, a research direction is suggested with a focus on multimodal deep learning models. Such models have the potential to provide more robust and generalized detection. It offers future researchers a clear starting point for building stronger defenses against harmful synthetic media.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。