研究发现语音质量评估模型在土耳其语和韩语上表现不稳定,提示需扩展多语言数据集。
Performance of Objective Speech Quality Metrics on Languages Beyond Validation Data: A Study of Turkish and Korean
- 对比英语基准,测试土耳其语与韩语的客观语音质量评估效果
- 土耳其语样本的ViSQOL得分显著更高,男性说话者相关性最强
- 揭示现有评估模型存在语言偏见,适合多语言语音研究者参考
客观语音质量度量广泛用于评估视频会议平台和电信系统的性能,能预测人工评分的语音质量,对体验质量评估至关重要。尽管应用广泛,这些度量工具仅基于有限语言开发,导致在未见过的语言上表现不确定甚至未被研究。本文通过分析土耳其语和韩语中两种客观语音质量度量(PESQ 和 ViSQOL)的表现,提升对此问题的认识。以英语为基准,结果显示土耳其语样本的ViSQOL得分显著更高,且土耳其男性说话者的PESQ与ViSQOL相关性最高。研究强调需探究度量工具在不同语言间的偏差,并建立涵盖多种语言的标注语音质量数据集。
原文摘要 · Abstract (English)
Objective speech quality measures are widely used to assess the performance of video conferencing platforms and telecommunication systems. They predict human-rated speech quality and are crucial for assessing the systems quality of experience. Despite the widespread use, the quality measures are developed on a limited set of languages. This can be problematic since the performance on unseen languages is consequently not guaranteed or even studied. Here we raise awareness to this issue by investigating the performance of two objective speech quality measures (PESQ and ViSQOL) on Turkish and Korean. Using English as baseline, we show that Turkish samples have significantly higher ViSQOL scores and that for Turkish male speakers the correlation between PESQ and ViSQOL is highest. These results highlight the need to explore biases across metrics and to develop a labeled speech quality dataset with a variety of languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。