arXiv:2508.09913cs.CV2025-08NeurIPS被引 16

利用音视频语音特征提升伪造人脸检测泛化能力

SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection

  • 通过自监督掩码预测学习音视频语音表征
  • 跨数据集检测准确率超越现有方法
  • 无需训练时使用假视频,适合真实场景

人脸伪造视频的检测仍是数字取证领域的重大挑战,尤其在面对未见数据集和常见扰动时。本文提出一种基于音视频语音协同的新型表示学习方法,利用音频信号中丰富的语音内容来精确反映面部运动。我们首先在真实视频上通过自监督掩码预测任务学习精准的音视频语音表征,同时编码局部与全局语义信息;随后直接将该模型迁移至伪造检测任务。大量实验表明,该方法在跨数据集泛化性和鲁棒性方面均优于当前最优方法,且训练过程中未使用任何伪造视频。代码已开源。

原文摘要 · Abstract (English)

Detection of face forgery videos remains a formidable challenge in the field of digital forensics, especially the generalization to unseen datasets and common perturbations. In this paper, we tackle this issue by leveraging the synergy between audio and visual speech elements, embarking on a novel approach through audio-visual speech representation learning. Our work is motivated by the finding that audio signals, enriched with speech content, can provide precise information effectively reflecting facial movements. To this end, we first learn precise audio-visual speech representations on real videos via a self-supervised masked prediction task, which encodes both local and global semantic information simultaneously. Then, the derived model is directly transferred to the forgery detection task. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods in terms of cross-dataset generalization and robustness, without the participation of any fake video in model training. Code is available at https://github.com/Eleven4AI/SpeechForensics.

伪造检测音视频融合自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。