用单流框架高效融合音视频特征,实现轻量级深伪检测。
Lightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework
- 设计单流多模态学习框架,通过协同块持续融合音视频特征。
- 仅0.48M参数,对已知和未知类型深伪均表现优异。
- 适合移动端或边缘设备部署,抗音视频错配能力强。
深度伪造是可能被滥用于传播虚假信息的AI合成多媒体数据,其生成同时涉及视觉与音频的篡改。以往方法通常采用独立的子模型分别学习音视频特征,再进行融合,难以充分利用二者内在关联,且易导致冗余网络层,影响效率。为此,本文提出一种轻量级音视频联合检测网络,基于单流多模态学习框架,引入协同音视频学习块,在特征学习过程中实现跨层持续融合。通过迭代使用该模块,网络无需堆叠大量模块即可有效捕捉多模态特征。此外,设计多模态分类模块,增强各模态分类器对内容的依赖性,并提升整体视频分类器对音视频不匹配的鲁棒性。在DF-TIMIT、FakeAVCeleb和DFDC三个基准数据集上实验表明,相比现有先进方法,本方法仅需0.48M参数,却在单模态、多模态及未见类型深伪检测中均取得更优性能。
原文摘要 · Abstract (English)
Deepfakes are AI-synthesized multimedia data that may be abused for spreading misinformation. Deepfake generation involves both visual and audio manipulation. To detect audio-visual deepfakes, previous studies commonly employ two relatively independent sub-models to learn audio and visual features, respectively, and fuse them subsequently for deepfake detection. However, this may underutilize the inherent correlations between audio and visual features. Moreover, utilizing two isolated feature learning sub-models can result in redundant neural layers, making the overall model inefficient and impractical for resource-constrained environments. In this work, we design a lightweight network for audio-visual deepfake detection via a single-stream multi-modal learning framework. Specifically, we introduce a collaborative audio-visual learning block to efficiently integrate multi-modal information while learning the visual and audio features. By iteratively employing this block, our single-stream network achieves a continuous fusion of multi-modal features across its layers. Thus, our network efficiently captures visual and audio features without the need for excessive block stacking, resulting in a lightweight network design. Furthermore, we propose a multi-modal classification module that can boost the dependence of the visual and audio classifiers on modality content. It also enhances the whole resistance of the video classifier against the mismatches between audio and visual modalities. We conduct experiments on the DF-TIMIT, FakeAVCeleb, and DFDC benchmark datasets. Compared to state-of-the-art audio-visual joint detection methods, our method is significantly lightweight with only 0.48M parameters, yet it achieves superiority in both uni-modal and multi-modal deepfakes, as well as in unseen types of deepfakes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。