无需参考语音即可评估发音严重程度,且抗噪能力强。
Reference-free automatic speech severity evaluation using acoustic unit language modelling
- 基于声学单元语言模型构建无参考评分方法
- 在噪声环境下表现优于传统声学特征方法
- 适用于自然口语场景,适合临床与真实应用
随着言语障碍的经济负担日益加重,言语严重程度评估的重要性愈发凸显。现有模型常因学习数据集特异性声学线索而泛化能力差,且多依赖参考语音或转录文本,限制了其在真实场景(如自发言语)中的应用。先前研究发现,自动语音自然度评分与严重程度评分高度相关,由此我们提出无需病理语音数据的无参考方法 SpeechLMScore。同时,我们构建了基于 NKI-CCRT 的 NKI-SpeechRT 数据集,为言语严重程度评估提供更全面的基础。本研究验证了 SpeechLMScore 在性能上优于传统声学特征方法,并评估了无参考与有参考模型之间的差距。此外,利用 NKI-SpeechRT 中的主观噪声评分,我们分析了噪声对模型的影响。结果表明,SpeechLMScore 具有良好的抗噪性,性能显著优于传统方法。
原文摘要 · Abstract (English)
Speech severity evaluation is becoming increasingly important as the economic burden of speech disorders grows. Current speech severity models often struggle with generalization, learning dataset-specific acoustic cues rather than meaningful correlates of speech severity. Furthermore, many models require reference speech or a transcript, limiting their applicability in ecologically valid scenarios, such as spontaneous speech evaluation. Previous research indicated that automatic speech naturalness evaluation scores correlate strongly with severity evaluation scores, leading us to explore a reference-free method, SpeechLMScore, which does not rely on pathological speech data. Additionally, we present the NKI-SpeechRT dataset, based on the NKI-CCRT dataset, to provide a more comprehensive foundation for speech severity evaluation. This study evaluates whether SpeechLMScore outperforms traditional acoustic feature-based approaches and assesses the performance gap between reference-free and reference-based models. Moreover, we examine the impact of noise on these models by utilizing subjective noise ratings in the NKI-SpeechRT dataset. The results demonstrate that SpeechLMScore is robust to noise and offers superior performance compared to traditional approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。