通过捕捉音视频局部时间不一致,提升深度伪造检测精度。
Audio-Visual Deepfake Detection With Local Temporal Inconsistencies
- 设计时序距离图与注意力机制,定位音视频时间错位。
- 在DFDC和FakeAVCeleb上优于现有方法,准确率显著提升。
- 适合关注多模态伪造检测的AI安全研究者。
本文提出一种音视频深度伪造检测方法,旨在捕捉音频与视觉模态间的细粒度时间不一致。从架构层面,设计了时序距离图结合注意力机制,以捕捉此类不一致,同时降低无关时间片段的干扰。此外,探索新型伪伪造生成技术,用于合成局部时间不一致。在DFDC和FakeAVCeleb数据集上,该方法相较于现有先进方法表现出更强的检测能力,验证了其有效性。
原文摘要 · Abstract (English)
This paper proposes an audio-visual deepfake detection approach that aims to capture fine-grained temporal inconsistencies between audio and visual modalities. To achieve this, both architectural and data synthesis strategies are introduced. From an architectural perspective, a temporal distance map, coupled with an attention mechanism, is designed to capture these inconsistencies while minimizing the impact of irrelevant temporal subsequences. Moreover, we explore novel pseudo-fake generation techniques to synthesize local inconsistencies. Our approach is evaluated against state-of-the-art methods using the DFDC and FakeAVCeleb datasets, demonstrating its effectiveness in detecting audio-visual deepfakes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。