攻击语音质量评估模型UTMOS,发现其存在脆弱性。
Attacking UTMOS: Probing the Robustness of a Speech Quality Assessment Model

- 通过优化音频不同表示空间,生成欺骗性输入
- 在保留感知质量的同时降低评分的攻击成功率达78%
- 使用EnCodec隐空间优化最有效,适合安全评估研究
UTMOS已成为语音处理研究中广泛使用的基于深度神经网络的语音质量评估(SQA)指标。本文针对UTMOS开展攻击以探测其鲁棒性。从高质量语音样本出发,在两个方向上优化输入:保持预测得分的降质攻击,以及保持感知质量但降低预测分数的降分攻击。考虑三种输入空间:原始波形、使用HiFi-GAN声码器生成的梅尔频谱,以及神经音频编解码器EnCodec的隐空间。实验结果表明,保持得分的攻击对UTMOS有效;虽然完全保持质量的攻击更难实现,但在EnCodec隐空间的优化提供了最佳成功率。这些结果揭示了UTMOS的失效模式,并强调了对基于DNN的SQA度量进行鲁棒性分析的重要性。
原文摘要 · Abstract (English)
UTMOS has become one of the most commonly used deep neural network-based speech quality assessment (SQA) metrics in speech processing research. In this paper, we attack UTMOS to probe its robustness. Starting from high-quality speech samples, we optimize the input in two directions: a score-preserving attack, which degrades perceived quality while maintaining the predicted score, and a quality-preserving attack, which lowers the predicted score while maintaining perceived quality. We consider three input spaces: raw waveform, mel spectrogram with a HiFi-GAN vocoder, and the latent space of EnCodec, a neural audio codec. Experimental results show that score-preserving attacks are effective against UTMOS. Although perfect quality-preserving attacks are more difficult, optimization in the EnCodec latent space provides the best chance of success. These results reveal failure modes of UTMOS and highlight the importance of robustness analysis for DNN-based SQA metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。