融合音视频信息,精准识别深度伪造内容。
ERF-BA-TFD+: A Multimodal Model for Audio-Visual Deepfake Detection
- 采用增强感受野机制捕捉音视频长程依赖关系。
- 在DDL-AV数据集上达到最佳准确率与最快处理速度。
- 适合关注多模态伪造检测的从业者和研究者。
深度伪造检测是识别被篡改多媒体内容的关键任务。现实场景中,深度伪造内容常以音视频多模态形式出现。为此,我们提出ERF-BA-TFD+,一种结合增强感受野(ERF)与音视频融合的新模型。该模型同时处理音视频特征,利用其互补信息提升检测准确率与鲁棒性。核心创新在于建模音视频输入中的长程依赖,从而更有效捕捉真实与伪造内容间的细微差异。我们在包含分段与完整视频片段的DDL-AV数据集上评估该模型,该数据集相比以往基准更具全面性和现实性。实验表明,该方法在准确性与处理速度上均超越现有技术,在‘深度伪造检测、定位与可解释性研讨会’音频-视觉检测与定位赛道(DDL-AV)中获第一名。
原文摘要 · Abstract (English)
Deepfake detection is a critical task in identifying manipulated multimedia content. In real-world scenarios, deepfake content can manifest across multiple modalities, including audio and video. To address this challenge, we present ERF-BA-TFD+, a novel multimodal deepfake detection model that combines enhanced receptive field (ERF) and audio-visual fusion. Our model processes both audio and video features simultaneously, leveraging their complementary information to improve detection accuracy and robustness. The key innovation of ERF-BA-TFD+ lies in its ability to model long-range dependencies within the audio-visual input, allowing it to better capture subtle discrepancies between real and fake content. In our experiments, we evaluate ERF-BA-TFD+ on the DDL-AV dataset, which consists of both segmented and full-length video clips. Unlike previous benchmarks, which focused primarily on isolated segments, the DDL-AV dataset allows us to assess the model's performance in a more comprehensive and realistic setting. Our method achieves state-of-the-art results on this dataset, outperforming existing techniques in terms of both accuracy and processing speed. The ERF-BA-TFD+ model demonstrated its effectiveness in the "Workshop on Deepfake Detection, Localization, and Interpretability," Track 2: Audio-Visual Detection and Localization (DDL-AV), and won first place in this competition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。