arXiv:2605.19630cs.AI2026-05中稿 · CVPR

用情绪特征提升深度伪造检测泛化能力,跨类型识别更准。

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection

论文配图:EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection
图 1 · 摘自论文原文
  • 融合视觉与音频情绪识别,建模跨模态时序一致性。
  • 在FakeAVCeleb上跨类型检测平均AUC提升2.1%。
  • 适合关注泛化性能的伪造检测研究者使用。

随着生成式AI模型的不断进步,伪造检测面临日益增大的压力。新伪造技术层出不穷,难以针对每种篡改方式收集训练数据,因此模型对未见篡改类型的泛化能力成为当前深度伪造检测的核心挑战。本文提出利用高层次语义线索来辅助低层次特征方法,以提升泛化能力。具体地,我们研究情绪作为高阶语义线索的作用,构建了名为Emo-Boost的多模态深度伪造检测框架。该框架融合现成的基于RGB和声学的深度伪造检测器与我们提出的基于情绪的检测器EmoForensics。EmoForensics通过视觉与音频情绪识别模块,建模音视频流中情绪表征的模态内与跨模态时序一致性。实验发现,EmoForensics与低层次方法捕捉互补信号,二者结合在FakeAVCeleb数据集上使跨篡改类型检测的平均AUC提升2.1%。

原文摘要 · Abstract (English)

With every advancement in generative AI models, forensics is under increasing pressure. The constant emergence of new generation techniques makes it impossible to collect data for each manipulation to train a deepfake detection model. Thus, generalizing to deepfakes unseen during training is one of the major challenges in current deepfake detection research. To tackle this challenge, we employ high-level semantic cues and argue that these cues can support low-level focused approaches in generalizing to unseen types of manipulations. In this work, we study emotions as a high-level semantic cue. We propose Emo-Boost, a multimodal deepfake detection framework that fuses an off-the-shelf RGB- and acoustic-focused deepfake detector with our emotion-based deepfake detector EmoForensics. EmoForensics utilises vision and audio emotion recognition modules and models intra- and inter-modal temporal consistency in emotion representations from an audio-visual stream. We found that EmoForensics and the low-level focused method capture complementary signals. Consequently, combining both signals in EmoBoost enhances the average cross-manipulation generalization AUC by 2.1% on FakeAVCeleb.

伪造检测情绪识别多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。