arXiv:2509.17162cs.SDeess.AS2025-09被引 1

构建可解释的深度伪造音频检测基准,提升识别与溯源能力

FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection

  • 设计多维度评估体系,覆盖定位、溯源与泛化三方面
  • 涵盖6种篡改类型和12个真实声源,测试模型可靠性
  • 揭示现有方法虽准确但缺乏解释力,推动可信音频认证

生成式音频的快速发展引发伪造数据带来的伦理与安全问题,使深度伪造音频检测成为防范技术滥用的重要防线。尽管已有研究探索该任务,但现有方法多集中于二分类,难以解释篡改机制、追溯来源或泛化至未见声源,限制了检测的可解释性与可靠性。为此,我们提出FakeSound2,一个旨在超越二分类准确率的基准。该基准从定位、溯源与泛化三个维度评估模型性能,涵盖6种篡改类型和12个多样化声源。实验表明,当前系统虽在分类上表现优异,却难以识别伪造模式分布,且无法提供可靠解释。通过揭示这些差距,FakeSound2建立了全面评估体系,明确关键挑战,致力于推动更鲁棒、可解释且泛化的音频认证方法发展。

原文摘要 · Abstract (English)

The rapid development of generative audio raises ethical and security concerns stemming from forged data, making deepfake sound detection an important safeguard against the malicious use of such technologies. Although prior studies have explored this task, existing methods largely focus on binary classification and fall short in explaining how manipulations occur, tracing where the sources originated, or generalizing to unseen sources-thereby limiting the explainability and reliability of detection. To address these limitations, we present FakeSound2, a benchmark designed to advance deepfake sound detection beyond binary accuracy. FakeSound2 evaluates models across three dimensions: localization, traceability, and generalization, covering 6 manipulation types and 12 diverse sources. Experimental results show that although current systems achieve high classification accuracy, they struggle to recognize forged pattern distributions and provide reliable explanations. By highlighting these gaps, FakeSound2 establishes a comprehensive benchmark that reveals key challenges and aims to foster robust, explainable, and generalizable approaches for trustworthy audio authentication.

音频伪造可解释性检测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。