提出多尺度跨模态注意力机制,提升音视频伪造检测精度
MSCT: Differential Cross-Modal Attention for Deepfake Detection

- 用多尺度自注意力融合邻近特征,增强局部细节提取
- 设计差分跨模态注意力,精准对齐音视频伪造痕迹
- 在FakeAVCeleb数据集上表现优异,适合多媒体安全研究者
音视频深度伪造检测通常采用互补的多模态模型来识别视频中的伪造痕迹。这些方法主要通过音视频对齐来提取伪造痕迹,源于音视频模态间的不一致性。然而,传统多模态伪造检测方法存在特征提取不足和模态对齐偏差问题。为此,我们提出一种多尺度跨模态变换器编码器(MSCT)用于深度伪造检测。该方法包含多尺度自注意力机制,用于整合相邻嵌入特征;以及差分跨模态注意力机制,用于融合多模态特征。实验表明,该结构在FakeAVCeleb数据集上表现具有竞争力,验证了其有效性。
原文摘要 · Abstract (English)
Audio-visual deepfake detection typically employs a complementary multi-modal model to check the forgery traces in the video. These methods primarily extract forgery traces through audio-visual alignment, which results from the inconsistency between audio and video modalities. However, the traditional multi-modal forgery detection method has the problem of insufficient feature extraction and modal alignment deviation. To address this, we propose a multi-scale cross-modal transformer encoder (MSCT) for deepfake detection. Our approach includes a multi-scale self-attention to integrate the features of adjacent embeddings and a differential cross-modal attention to fuse multi-modal features. Our experiments demonstrate competitive performance on the FakeAVCeleb dataset, validating the effectiveness of the proposed structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。