分析音频伪造检测模型关注的时域特征,揭示其决策依据。
What Does an Audio Deepfake Detector Focus on? A Study in the Time Domain
- 基于相关性设计可解释AI方法,分析模型对音频时序区域的关注度。
- 在大规模数据集上验证,模型更关注语音起止与发音内容。
- 发现小样本研究结论不适用于大样本,结果更具可信度。
为提升音频深度伪造检测(ADD)模型在实际应用中的可信度,本文提出一种基于相关性的可解释AI(XAI)方法,用于分析基于Transformer的ADD模型的预测过程。通过对比标准Grad-CAM与基于SHAP的方法,采用定量忠实性指标及部分伪造测试,全面评估音频不同时间区域的重要性。研究基于大规模数据集,发现现有XAI方法的解释存在差异。所提出的相关性方法在多种指标上表现最优。进一步分析表明,语音/非语音、语音内容以及语音起止点的重要性在大规模数据下与小样本结果不一致,说明先前基于有限语句的研究结论可能不具备泛化性。
原文摘要 · Abstract (English)
Adding explanations to audio deepfake detection (ADD) models will boost their real-world application by providing insight on the decision making process. In this paper, we propose a relevancy-based explainable AI (XAI) method to analyze the predictions of transformer-based ADD models. We compare against standard Grad-CAM and SHAP-based methods, using quantitative faithfulness metrics as well as a partial spoof test, to comprehensively analyze the relative importance of different temporal regions in an audio. We consider large datasets, unlike previous works where only limited utterances are studied, and find that the XAI methods differ in their explanations. The proposed relevancy-based XAI method performs the best overall on a variety of metrics. Further investigation on the relative importance of speech/non-speech, phonetic content, and voice onsets/offsets suggest that the XAI results obtained from analyzing limited utterances don't necessarily hold when evaluated on large datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。