arXiv:2505.12966cs.CVcs.AI2025-05中稿 · ed被引 4

多尺度自适应模型提升音视频伪造检测准确率

Multiscale Adaptive Conflict-Balancing Model For Multimedia Deepfake Detection

  • 引入对比学习实现跨模态多层级融合,缓解模态冲突
  • 平均准确率达95.5%,跨数据集测试提升超7.7%
  • 适合需要高鲁棒性伪造检测的系统开发者

计算机视觉与深度学习的发展使深度伪造音视频与真实内容界限模糊,威胁多媒体可信度。现有多模态检测方法仍受限于模态间学习不平衡。为此,我们提出音频-视觉联合学习方法(MACB-DF),通过对比学习辅助多层次、跨模态融合,有效平衡并利用各模态信息。此外,设计正交化多模态帕累托模块,在保留单模态信息的同时,解决音频-视频编码器因损失函数优化目标差异导致的梯度冲突问题。在主流深度伪造数据集上的大量实验与消融研究显示,本模型在关键评估指标上持续提升,平均准确率达到95.5%。尤其在跨数据集泛化能力上表现优异,于训练集为DFDC时,在DefakeAVMiT与FakeAVCeleb数据集上分别取得8.0%和7.7%的绝对性能提升,优于当前最佳方法。

原文摘要 · Abstract (English)

Advances in computer vision and deep learning have blurred the line between deepfakes and authentic media, undermining multimedia credibility through audio-visual forgery. Current multimodal detection methods remain limited by unbalanced learning between modalities. To tackle this issue, we propose an Audio-Visual Joint Learning Method (MACB-DF) to better mitigate modality conflicts and neglect by leveraging contrastive learning to assist in multi-level and cross-modal fusion, thereby fully balancing and exploiting information from each modality. Additionally, we designed an orthogonalization-multimodal pareto module that preserves unimodal information while addressing gradient conflicts in audio-video encoders caused by differing optimization targets of the loss functions. Extensive experiments and ablation studies conducted on mainstream deepfake datasets demonstrate consistent performance gains of our model across key evaluation metrics, achieving an average accuracy of 95.5% across multiple datasets. Notably, our method exhibits superior cross-dataset generalization capabilities, with absolute improvements of 8.0% and 7.7% in ACC scores over the previous best-performing approach when trained on DFDC and tested on DefakeAVMiT and FakeAVCeleb datasets.

伪造检测多模态对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。