arXiv:2506.21711cs.CV2025-06中稿 · and published in K…被引 6

用跨注意力融合时空特征,提升深伪视频检测精度

CAST: Cross-Attentive Spatio-Temporal feature fusion for deepfake detection

  • 引入跨注意力机制,让时间特征动态关注关键空间区域
  • 在多数据集上达99.49% AUC,跨数据集仍保持81.25%以上表现
  • 适合需要高鲁棒性检测的安防与内容审核场景

深度伪造对数字媒体真实性构成重大威胁,亟需能识别细微且随时间变化篡改的先进检测技术。卷积神经网络擅长捕捉空间伪影,而变压器模型在建模时间不一致方面表现优异。然而,现有CNN-Transformer模型常独立处理时空特征,注意力方法也多采用独立机制并简单拼接或平均,限制了时空交互深度。为此,我们提出统一的CAST模型,利用跨注意力实现更深度融合的时空特征融合。该设计使时间特征可动态关注相关空间区域,增强对闪烁眼睛、变形嘴唇等细粒度时变伪影的检测能力,实现更精准定位与深层上下文理解。我们在FaceForensics++、Celeb-DF、DeepfakeDetection和Deepfake Detection Challenge(DFDC)数据集上评估,跨数据集设置下分别取得93.31%和81.25%的AUC,内数据集测试中达99.49% AUC与97.57%准确率,验证了跨注意力融合的有效性。

原文摘要 · Abstract (English)

Deepfakes have emerged as a significant threat to digital media authenticity, increasing the need for advanced detection techniques that can identify subtle and time-dependent manipulations. CNNs are effective at capturing spatial artifacts and Transformers excel at modeling temporal inconsistencies. However, many existing CNN-Transformer models process spatial and temporal features independently. In particular, attention based methods often use independent attention mechanisms for spatial and temporal features and combine them using naive approaches like averaging, addition or concatenation, limiting the depth of spatio-temporal interaction. To address this challenge, we propose a unified CAST model that leverages cross-attention to effectively fuse spatial and temporal features in a more integrated manner. Our approach allows temporal features to dynamically attend to relevant spatial regions, enhancing the model's ability to detect fine-grained, time-evolving artifacts such as flickering eyes or warped lips. This design enables more precise localization and deeper contextual understanding, leading to improved performance across diverse and challenging scenarios. We evaluate the performance of our model using the FaceForensics++, Celeb-DF, DeepfakeDetection and Deepfake Detection Challenge (DFDC) datasets in both intra and cross dataset settings to affirm the superiority of our approach. Our model achieves strong performance with an Area Under the Curve (AUC) of 99.49% and an accuracy of 97.57% in intra-dataset evaluations. In cross-dataset testing, the model achieves AUC scores of 93.31% and 81.25% on the unseen DeepFakeDetection and DFDC datasets, respectively. These results highlight the effectiveness of cross-attention-based feature fusion in enhancing the robustness of deepfake video detection.

深伪检测跨注意力时空融合视频安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。