arXiv:2410.07436cs.LGcs.SD2024-10被引 12

提升音频深度伪造检测的可解释性,增强真实场景下的可靠性。

Toward Robust Real-World Audio Deepfake Detection: Closing the Explainability Gap

  • 为基于Transformer的音频伪造检测器设计新型可解释方法
  • 构建真实世界泛化能力的新基准数据集
  • 帮助专家理解模型决策,推动公众参与检测

AI生成或篡改的音频深度伪造迅速蔓延,严重威胁媒体真实性与选举安全。现有基于AI的检测方案缺乏可解释性,在真实场景中表现不佳。本文提出针对先进Transformer架构音频伪造检测器的新可解释方法,并开源一个用于评估真实世界泛化能力的新基准。通过缩小基于Transformer的检测器与传统方法之间的可解释性差距,我们的成果不仅增强了人类专家对模型的信任,还为激发公众智能以解决音频伪造检测的可扩展性问题铺平了道路。

原文摘要 · Abstract (English)

The rapid proliferation of AI-manipulated or generated audio deepfakes poses serious challenges to media integrity and election security. Current AI-driven detection solutions lack explainability and underperform in real-world settings. In this paper, we introduce novel explainability methods for state-of-the-art transformer-based audio deepfake detectors and open-source a novel benchmark for real-world generalizability. By narrowing the explainability gap between transformer-based audio deepfake detectors and traditional methods, our results not only build trust with human experts, but also pave the way for unlocking the potential of citizen intelligence to overcome the scalability issue in audio deepfake detection.

音频伪造可解释性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。