评测五种语音可懂度指标,发现基于参考的更优且需更多西班牙语数据
An Objective Intelligibility Metric Evaluation on Spanish Speech

- 用新发布的西班牙语数据集评测多种语音可懂度指标
- 基于参考的指标整体优于无参考深度学习方法
- 强调需更多跨语言数据提升无参考模型泛化能力
客观可懂度度量(OIMs)可实现快速低成本的语音可懂度评估,广泛用于语音技术评测。本研究在新发布的西班牙语语音可懂度数据集SpInt上,评估了五种基于参考的OIMs(STOI、ESTOI、STGI、HASPI和SIIB)以及两种基于深度学习的无参考度量(MOSA-Net+和W2V-SIP)。结果表明,基于参考的OIMs始终优于现代数据驱动的无参考方法,后者在训练-测试声学不匹配(如语言不匹配)下性能显著下降。该现象在本研究中尤为明显,因所有被评估指标在开发阶段均未接触过西班牙语语音数据。为促进更鲁棒、更具泛化性的无参考OIMs研究,SpInt数据集已公开发布。
原文摘要 · Abstract (English)
Objective intelligibility metrics (OIMs) enable fast and low-cost evaluation of speech intelligibility and are widely used in speech technology assessment. In this study, we evaluate five reference-based OIMs (STOI, ESTOI, STGI, HASPI, and SIIB) and two deep learning-based no-reference metrics (MOSA-Net+ and W2V-SIP) on SpInt, a new Spanish speech intelligibility dataset. Our results show that reference-based OIMs consistently outperform modern data-driven no-reference approaches, which degrade notably under training-test acoustic mismatches such as language mismatch. This effect is particularly relevant in our scenario, as none of the evaluated metrics were exposed to Spanish speech data during development. Consequently, to foster research on more robust and generalizable no-reference OIMs, SpInt is released publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。