arXiv:2506.02401cs.SDcs.MM2025-06

用狄利克雷分布建模判断可信度,提升假音频检测可靠性

Trusted Fake Audio Detection Based on Dirichlet Distribution

  • 用神经网络生成证据,再用狄利克雷分布建模不确定性
  • 在ASVspoof多个数据集上准确率超基准模型,且信任度更高
  • 适合关注模型决策可信度的安全防护研究者

随着基于深度学习的语音转换与语音合成技术不断发展,假音频带来的网络安全问题日益严重。以往的假音频防御模型虽表现优异,但均未能有效建模自身决策的可信度。为此,本文提出一种基于狄利克雷分布的可信假音频检测方法,以提升检测结果的可靠性。具体而言,首先通过神经网络生成证据,再利用狄利克雷分布对不确定性进行建模;通过狄利克雷分布参数表示信念分布,从而获得每个决策的不确定性估计。最终将预测概率与对应的不确定性估计融合,形成综合判断。在ASVspoof系列数据集(即ASVspoof 2019 LA、ASVspoof 2021 LA和DF)上开展多组对比实验,验证了所提模型在准确性、鲁棒性和可信度方面的优异性能。

原文摘要 · Abstract (English)

With the continuous development of deep learning-based speech conversion and speech synthesis technologies, the cybersecurity problem posed by fake audio has become increasingly serious. Previously proposed models for defending against fake audio have attained remarkable performance. However, they all fall short in modeling the trustworthiness of the decisions made by the models themselves. Based on this, we put forward a plausible fake audio detection approach based on the Dirichlet distribution with the aim of enhancing the reliability of fake audio detection. Specifically, we first generate evidence through a neural network. Uncertainty is then modeled using the Dirichlet distribution. By modeling the belief distribution with the parameters of the Dirichlet distribution, an estimate of uncertainty can be obtained for each decision. Finally, the predicted probabilities and corresponding uncertainty estimates are combined to form the final opinion. On the ASVspoof series dataset (i.e., ASVspoof 2019 LA, ASVspoof 2021 LA, and DF), we conduct a number of comparison experiments to verify the excellent performance of the proposed model in terms of accuracy, robustness, and trustworthiness.

假音频检测可信度建模狄利克雷分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。