解决唱歌场景下音视频伪造检测难题,提升跨场景泛化能力。
From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection

- 用文本引导学习面部真实特征模式,增强跨场景识别能力。
- 在说话和唱歌数据集上均显著优于现有方法,鲁棒性强。
- 构建首个专注歌唱的音视频伪造数据集SHDF,填补领域空白。
随着音视频生成模型的快速发展,可靠的伪造检测愈发关键。现有音视频伪造检测方法多依赖跨模态不一致,但在歌唱场景中,节奏性发声削弱了音视频耦合,引入显著域偏移,大幅降低检测性能。为此,我们利用节奏感知生成模型构建了首个歌唱头像伪造数据集SHDF,填补歌唱场景基准空白。为应对跨场景域偏移,提出文本引导的音视频伪造检测框架T-AVFD,该框架包含面部真实性模式学习器和多模态差异权重学习模块。前者将面部特征与多粒度文本描述对齐,学习可泛化的真实性模式;后者保留内在音视频一致性,并通过差异加权自适应融合真实性模式。在多个说话头像伪造数据集及SHDF上的大量实验表明,T-AVFD持续优于现有基线,在多种扰动下仍具强鲁棒性。
原文摘要 · Abstract (English)
With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio-visual deepfake detection typically rely on cross-modal inconsistencies. In singing, rhythmic vocalization weakens this coupling and introduces a nontrivial domain shift, substantially degrading detection performance. We construct the Singing Head DeepFake (SHDF) dataset using rhythm-aware generative models to fill the gap in singing benchmarks. To cope with cross-scenario domain shifts, we propose a Text-guided Audio-Visual Forgery Detection (T-AVFD) framework that generalizes across both talking and singing scenarios. T-AVFD comprises a facial authenticity pattern learner and a multi-modal differential weight learning module. The pattern learner aligns facial features with multi-granularity textual descriptions to learn generalizable authenticity patterns. The weight learning module preserves intrinsic audio-visual consistency and adaptively integrates it with authenticity patterns via differential weighting. Extensive experiments on multiple talking head deepfake datasets and SHDF show consistent improvements over existing baselines and strong robustness under diverse perturbations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。