统一模型同时评估语音自然度、可懂度等多维度质量,更贴近人耳判断。
Uni-VERSA: Versatile Speech Assessment with a Unified Network
- 用一个统一网络同时预测多种语音质量指标。
- 在URGENT24数据集上表现接近人工听感评分。
- 适合语音增强、合成与质量控制场景使用。
主观听感测试仍是语音质量评估的黄金标准,但成本高、结果波动大且难以扩展。现有客观指标如PESQ、F0相关性、DNSMOS通常仅反映语音质量的单一维度。为此,我们提出Uni-VERSA,一种统一网络,可同时预测自然度、可懂度、说话人特征、语调和噪声等多个指标,实现对语音信号的全面评估。我们建立了其框架、评估协议及在语音增强、语音合成和质量控制中的应用。基于URGENT24挑战赛的基准测试表明,Uni-VERSA为单维度评估方法提供了可行替代方案,并与人类听觉感知高度一致,是未来语音质量评估的有前景方向。
原文摘要 · Abstract (English)
Subjective listening tests remain the golden standard for speech quality assessment, but are costly, variable, and difficult to scale. In contrast, existing objective metrics, such as PESQ, F0 correlation, and DNSMOS, typically capture only specific aspects of speech quality. To address these limitations, we introduce Uni-VERSA, a unified network that simultaneously predicts various objective metrics, encompassing naturalness, intelligibility, speaker characteristics, prosody, and noise, for a comprehensive evaluation of speech signals. We formalize its framework, evaluation protocol, and applications in speech enhancement, synthesis, and quality control. A benchmark based on the URGENT24 challenge, along with a baseline leveraging self-supervised representations, demonstrates that Uni-VERSA provides a viable alternative to single-aspect evaluation methods. Moreover, it aligns closely with human perception, making it a promising approach for future speech quality assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。