arXiv:2506.03425eess.AScs.AI2025-06中稿 · Interspeech 2025被引 3

用扩散模型定位语音深度伪造的异常区域,效果优于传统解释方法。

A Data-Driven Diffusion-based Approach for Audio Deepfake Explanations

  • 以真实与合成语音的时频差异为标签,训练扩散模型识别伪造痕迹。
  • 在VocV4和LibriSeVoc数据集上,解释准确率显著高于SHAP、LRP等方法。
  • 适合研究语音伪造检测、可解释人工智能或音频安全的学者使用。

评估SHAP、LRP等可解释性技术在语音深度伪造检测中的表现面临挑战,因缺乏清晰的标注真值。即使获得真值,这些方法也难以提供准确解释。本文提出一种新的数据驱动方法,通过对比真实语音与声码器生成语音的时频表示差异,构建真值标签,指导扩散模型识别给定声码化语音中的伪造区域。在VocV4和LibriSeVoc数据集上的实验表明,该方法在定性和定量上均优于传统解释技术。

原文摘要 · Abstract (English)

Evaluating explainability techniques, such as SHAP and LRP, in the context of audio deepfake detection is challenging due to lack of clear ground truth annotations. In the cases when we are able to obtain the ground truth, we find that these methods struggle to provide accurate explanations. In this work, we propose a novel data-driven approach to identify artifact regions in deepfake audio. We consider paired real and vocoded audio, and use the difference in time-frequency representation as the ground-truth explanation. The difference signal then serves as a supervision to train a diffusion model to expose the deepfake artifacts in a given vocoded audio. Experimental results on the VocV4 and LibriSeVoc datasets demonstrate that our method outperforms traditional explainability techniques, both qualitatively and quantitatively.

语音伪造扩散模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。