arXiv:2504.15663eess.AScs.AI2025-04中稿 · ICASSP 2025被引 1

用证据深度学习提升假音频检测的不确定性感知能力。

FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning

  • 引入狄利克雷分布建模分类概率,显式捕捉预测不确定性
  • 在ASVspoof2019/2021数据集上显著降低错误率,尤其对未知攻击有效
  • 不确定性与识别错误率高度相关,可为系统提供可信度参考

近期,随着语音合成与语音转换技术的发展,自动说话人验证(ASV)系统面临伪造攻击的威胁,假音频检测受到广泛关注。该任务的核心挑战在于模型对未见的、分布外(OOD)攻击的泛化能力。现有方法因使用softmax分类导致过自信问题,在面对不可预测的伪造尝试时可能产生不可靠预测。为此,本文提出一种新型框架FADEL,通过狄利克雷分布建模类别概率,将模型不确定性融入预测过程,从而在分布外场景下实现更稳健的性能。在ASVspoof2019逻辑访问(LA)和ASVspoof2021 LA数据集上的实验表明,所提方法显著提升了基线模型的表现。此外,我们通过分析平均不确定性与等错误率(EER)之间的强相关性,验证了不确定性估计的有效性。

原文摘要 · Abstract (English)

Recently, fake audio detection has gained significant attention, as advancements in speech synthesis and voice conversion have increased the vulnerability of automatic speaker verification (ASV) systems to spoofing attacks. A key challenge in this task is generalizing models to detect unseen, out-of-distribution (OOD) attacks. Although existing approaches have shown promising results, they inherently suffer from overconfidence issues due to the usage of softmax for classification, which can produce unreliable predictions when encountering unpredictable spoofing attempts. To deal with this limitation, we propose a novel framework called fake audio detection with evidential learning (FADEL). By modeling class probabilities with a Dirichlet distribution, FADEL incorporates model uncertainty into its predictions, thereby leading to more robust performance in OOD scenarios. Experimental results on the ASVspoof2019 Logical Access (LA) and ASVspoof2021 LA datasets indicate that the proposed method significantly improves the performance of baseline models. Furthermore, we demonstrate the validity of uncertainty estimation by analyzing a strong correlation between average uncertainty and equal error rate (EER) across different spoofing algorithms.

假音频检测不确定性深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。