用SITool工具评估13种语音编解码器的可懂度,发现客观指标仅部分有效。
Benchmarking Neural Speech Codec Intelligibility with SITool
- 开发基于Flask的SITool,支持实验室与众包环境下的听觉测试
- 13种编解码器中神经模型整体可懂度优于传统方法,但效果不一
- 只有STOI和ESTOI与主观结果显著相关,无法捕捉性别和词表差异
语音可懂度评估对神经语音编解码器至关重要,但多数评估仍聚焦整体质量而非可懂度。目前仅有少数公开工具可用于标准化测试,如诊断韵母测试(DRT)和改良韵母测试(MRT)。本文提出用于主观评估的语音可懂度工具包(SITool),一个基于Flask的Web应用,支持实验室与众包场景下的DRT和MRT测试。我们利用SITool对13种神经与传统语音编解码器进行评测,分析音素层面退化情况,并将主观DRT结果与客观可懂度指标对比。结果显示,尽管神经编解码器在主观可懂度上普遍优于传统方法,但仅STOI和ESTOI与主观评价显著相关,而字错误率(WER)未呈现显著相关性;此外,这些客观指标难以捕捉主观评估中观察到的性别与词表特异性差异。
原文摘要 · Abstract (English)
Speech intelligibility assessment is essential for evaluating neural speech codecs, yet most evaluation efforts focus on overall quality rather than intelligibility. Only a few publicly available tools exist for conducting standardized intelligibility tests, like the Diagnostic Rhyme Test (DRT) and Modified Rhyme Test (MRT). We introduce the Speech Intelligibility Toolkit for Subjective Evaluation (SITool), a Flask-based web application for conducting DRT and MRT in laboratory and crowdsourcing settings. We use SITool to benchmark 13 neural and traditional speech codecs, analyzing phoneme-level degradations and comparing subjective DRT results with objective intelligibility metrics. Our findings show that, while neural speech codecs can outperform traditional ones in subjective intelligibility, only STOI and ESTOI - not WER - significantly correlate with subjective results, although they struggle to capture gender and wordlist-specific variations observed in subjective evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。