arXiv:2506.05851cs.MMcs.AI2025-06被引 1

提出新数据集与评估方案,解决音频视频伪造检测中的漏洞问题。

DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection

  • 构建SIMBA基线模型,简化多模态伪造检测设计
  • 发现并缓解音频片段中的静音捷径问题
  • 改进主流数据集的评测方法,提升检测可信度

生成式AI快速发展,使伪造音视频愈发逼真,带来严重的安全与伦理威胁。现有深度伪造检测研究多关注音视频多模态场景,但面临数据集可复现性差、关键缺陷等问题,如广受使用的FakeAVCeleb数据集中存在的静音捷径。本文从数据集、检测方法与评估协议三个核心维度深入分析基准测试难题。针对问题,我们聚焦近期发布的DeepSpeak v1数据集,首次提出评估协议,并用SOTA模型进行基准测试。提出SImple Multimodal BAseline(SIMBA)——一种简洁高效的基线方法,支持多样化设计探索。同时深化对音频捷径现象的理解,提出有效缓解策略。最后,系统分析并优化了FakeAVCeleb数据集的评估方案。研究成果为音视频深度伪造检测提供了更可靠的前进路径。

原文摘要 · Abstract (English)

Generative AI advances rapidly, allowing the creation of very realistic manipulated video and audio. This progress presents a significant security and ethical threat, as malicious users can exploit DeepFake techniques to spread misinformation. Recent DeepFake detection approaches explore the multimodal (audio-video) threat scenario. In particular, there is a lack of reproducibility and critical issues with existing datasets - such as the recently uncovered silence shortcut in the widely used FakeAVCeleb dataset. Considering the importance of this topic, we aim to gain a deeper understanding of the key issues affecting benchmarking in audio-video DeepFake detection. We examine these challenges through the lens of the three core benchmarking pillars: datasets, detection methods, and evaluation protocols. To address these issues, we spotlight the recent DeepSpeak v1 dataset and are the first to propose an evaluation protocol and benchmark it using SOTA models. We introduce SImple Multimodal BAseline (SIMBA), a competitive yet minimalistic approach that enables the exploration of diverse design choices. We also deepen insights into the issue of audio shortcuts and present a promising mitigation strategy. Finally, we analyze and enhance the evaluation scheme on the widely used FakeAVCeleb dataset. Our findings offer a way forward in the complex area of audio-video DeepFake detection.

伪造检测多模态数据集音视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。