用面部动作单元提升音视频伪造检测效果
FauForensics: Boosting Audio-Visual Deepfake Detection with Facial Action Units
- 引入生物不变的面部动作单元作为伪造抗性特征
- 帧级音视频相似度计算,准确率提升4.83%
- 适合关注跨数据集泛化能力的研究者
生成式AI的快速发展加剧了真实音视频伪造的威胁,亟需鲁棒的检测方法。现有方案多针对单模态伪造(音频或视觉),在处理多模态篡改时表现不佳,主要因无法有效处理异构模态特征且跨数据集泛化能力差。为此,我们提出FauForensics框架,引入与情绪生理相关的生物不变面部动作单元(FAUs),作为能减少领域依赖并捕捉合成内容中常被破坏的细微动态的伪造抗性表征。此外,不同于以往比较整段视频的方法,本方法通过专用融合模块结合可学习的跨模态查询,实现细粒度的帧级音视频相似度计算,动态对齐时空唇音关系,缓解多模态特征异质性问题。在FakeAVCeleb和LAV-DF上的实验表明,该方法达到当前最优(SOTA)性能,跨数据集泛化能力显著,平均性能优于现有方法4.83%。
原文摘要 · Abstract (English)
The rapid evolution of generative AI has increased the threat of realistic audio-visual deepfakes, demanding robust detection methods. Existing solutions primarily address unimodal (audio or visual) forgeries but struggle with multimodal manipulations due to inadequate handling of heterogeneous modality features and poor generalization across datasets. To this end, we propose a novel framework called FauForensics by introducing biologically invariant facial action units (FAUs), which is a quantitative descriptor of facial muscle activity linked to emotion physiology. It serves as forgery-resistant representations that reduce domain dependency while capturing subtle dynamics often disrupted in synthetic content. Besides, instead of comparing entire video clips as in prior works, our method computes fine-grained frame-wise audiovisual similarities via a dedicated fusion module augmented with learnable cross-modal queries. It dynamically aligns temporal-spatial lip-audio relationships while mitigating multi-modal feature heterogeneity issues. Experiments on FakeAVCeleb and LAV-DF show state-of-the-art (SOTA) performance and superior cross-dataset generalizability with up to an average of 4.83\% than existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。