arXiv:2507.11306eess.AS2025-07被引 5

测试生成式语音增强的主观评价方法,发现需结合语音保真度指标才能可靠检测幻觉。

P.808 Multilingual Speech Enhancement Testing: Approach and Results of URGENT 2025 Challenge

  • 基于P.808标准,本地化多语言语音与文本数据进行众包主观评估。
  • 生成式语音增强方法下,主观评分与客观指标存在不一致,需引入语音保真度评估。
  • 适合关注语音增强评测、多语言评估及生成模型可靠性研究者。

在语音增强(SE)系统语音质量评估中,主观听觉测试至今被视为金标准。随着大量生成式或混合方法涌入该领域,现有客观指标的局限性日益凸显。2025年Interspeech URGENT语音增强挑战赛引入非英语数据集,进一步拓展了多语言评估维度。本文简要回顾了ITU-T P.808众包主观绝对类别评级(ACR)测试方法。首次提出对Naderi与Cutler实现的文本到语音(TTS)众包测试中音频与文本成分的本地化流程。并对URGENT挑战赛结果进行了深入分析,揭示了在生成式AI时代,传统(P.808)ACR主观测试作为金标准的可靠性问题。特别发现:对于生成式语音增强方法,主观(ACR MOS)与客观(DNSMOS、NISQA)无参考指标应辅以客观电话保真度指标,方可可靠检测语音内容幻觉。最后,我们即将发布本地化脚本与方法,便于按ITU-T P.808标准开展新的多语言语音增强主观评估。

原文摘要 · Abstract (English)

In speech quality estimation for speech enhancement (SE) systems, subjective listening tests so far are considered as the gold standard. This should be even more true considering the large influx of new generative or hybrid methods into the field, revealing issues of some objective metrics. Efforts such as the Interspeech 2025 URGENT Speech Enhancement Challenge also involving non-English datasets add the aspect of multilinguality to the testing procedure. In this paper, we provide a brief recap of the ITU-T P.808 crowdsourced subjective listening test method. A first novel contribution is our proposed process of localizing both text and audio components of Naderi and Cutler's implementation of crowdsourced subjective absolute category rating (ACR) listening tests involving text-to-speech (TTS). Further, we provide surprising analyses of and insights into URGENT Challenge results, tackling the reliability of (P.808) ACR subjective testing as gold standard in the age of generative AI. Particularly, it seems that for generative SE methods, subjective (ACR MOS) and objective (DNSMOS, NISQA) reference-free metrics should be accompanied by objective phone fidelity metrics to reliably detect hallucinations. Finally, we will soon release our localization scripts and methods for easy deployment for new multilingual speech enhancement subjective evaluations according to ITU-T P.808.

语音增强主观评估生成模型多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。