arXiv:2411.10193cs.CV2024-11被引 17

通过音视频语义差异定位深度伪造,提升小范围篡改检测精度。

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

  • 利用音视频在机器感知中的语义不一致来识别伪造
  • 在AV-Deepfake1M上检测AUC提升30.5,定位[email protected]提高47.88
  • 适合需要精确定位音视频伪造位置的研究与应用

深度伪造技术快速发展,对在线多媒体信息的完整性与可信度构成严重威胁。尽管深度伪造检测已取得显著进展,但音频与视觉模态的同时、局部或细微篡改仍带来巨大挑战。为此,我们提出DiMoDif,一种基于音视频语义一致性假设的多模态检测框架:真实样本中音视频信号在信息层面应一致,而伪造样本则存在不匹配。DiMoDif采用专用于音视频语音识别的深层网络特征,捕捉帧级跨模态不一致,并通过分层跨模态融合网络,结合自适应时间对齐模块和可学习差异映射层,显式建模音视频表示间的细微差异。检测模型使用复合损失函数优化,同时关注帧级判断与伪造片段定位。在极具挑战性的AV-Deepfake1M数据集上,其检测任务AUC较当前最优提升30.5,定位任务[email protected]提升47.88;在FakeAVCeleb与LAV-DF上表现优异。代码已开源。

原文摘要 · Abstract (English)

Deepfake technology has rapidly advanced and poses significant threats to information integrity and trust in online multimedia. While significant progress has been made in detecting deepfakes, the simultaneous manipulation of audio and visual modalities, sometimes at small parts or in subtle ways, presents highly challenging detection scenarios. To address these challenges, we present DiMoDif, an audio-visual deepfake detection framework that leverages the inter-modality differences in machine perception of speech, based on the assumption that in real samples -- in contrast to deepfakes -- visual and audio signals coincide in terms of information. DiMoDif leverages features from deep networks that specialize in visual and audio speech recognition to spot frame-level cross-modal incongruities, and in that way to temporally localize the deepfake forgery. To this end, we devise a hierarchical cross-modal fusion network, integrating adaptive temporal alignment modules and a learned discrepancy mapping layer to explicitly model the subtle differences between visual and audio representations. Then, the detection model is optimized through a composite loss function accounting for frame-level detections and fake intervals localization. DiMoDif outperforms the state-of-the-art on the Deepfake Detection task by 30.5 AUC on the highly challenging AV-Deepfake1M, while it performs exceptionally on FakeAVCeleb and LAV-DF. On the Temporal Forgery Localization task, it outperforms the state-of-the-art by 47.88 [email protected] on AV-Deepfake1M, and performs on-par on LAV-DF. Code available at https://github.com/mever-team/dimodif.

深度伪造检测音视频分析跨模态对齐定位精度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。