通过多模态注意力融合视觉文本频域特征,提升跨模型深伪检测能力
CAMME: Adaptive Deepfake Image Detection with Multi-Modal Cross-Attention
- 用多头交叉注意力动态融合视觉、文本和频域特征
- 自然场景与人脸深伪检测准确率分别提升12.56%和13.25%
- 对对抗攻击和图像扰动保持高鲁棒性,适合实际部署
高度逼真的AI生成深伪图像泛滥,严重威胁数字媒体可信度与社会安全。现有检测方法在特定生成领域表现良好,但面对未知生成架构时性能显著下降,这是生成技术快速演进下的根本局限。本文提出CAMME(Cross-Attention Multi-Modal Embeddings)框架,通过多头交叉注意力机制动态融合视觉、文本与频域特征,实现跨域泛化能力。大量实验表明,该框架优于当前最优方法,在自然场景与人脸深伪检测上分别提升12.56%与13.25%。框架在自然图像扰动下仍保持超91%准确率,并在PGD与FGSM对抗攻击下分别达到89.01%与96.14%准确率。结果验证了通过交叉注意力整合互补模态,可更有效重定义决策边界,实现对异构生成架构的可靠深伪检测。
原文摘要 · Abstract (English)
The proliferation of sophisticated AI-generated deepfakes poses critical challenges for digital media authentication and societal security. While existing detection methods perform well within specific generative domains, they exhibit significant performance degradation when applied to manipulations produced by unseen architectures--a fundamental limitation as generative technologies rapidly evolve. We propose CAMME (Cross-Attention Multi-Modal Embeddings), a framework that dynamically integrates visual, textual, and frequency-domain features through a multi-head cross-attention mechanism to establish robust cross-domain generalization. Extensive experiments demonstrate CAMME's superiority over state-of-the-art methods, yielding improvements of 12.56% on natural scenes and 13.25% on facial deepfakes. The framework demonstrates exceptional resilience, maintaining (over 91%) accuracy under natural image perturbations and achieving 89.01% and 96.14% accuracy against PGD and FGSM adversarial attacks, respectively. Our findings validate that integrating complementary modalities through cross-attention enables more effective decision boundary realignment for reliable deepfake detection across heterogeneous generative architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。