融合声学特征与大模型表示,提升病理性嗓音评估的准确性和鲁棒性。
Towards Robust Assessment of Pathological Voices via Combined Low-Level Descriptors and Foundation Model Representations
- 结合低层声学参数与语音大模型嵌入,用注意力机制提取高阶特征。
- 句子级输入比元音级输入更准确,患者级别评估效果显著提升。
- 在噪声环境下仍保持稳定性能,适合真实场景和远程医疗应用。
嗓音质量的感知评估在诊断和监测嗓音障碍中至关重要。传统方法如CAPE-V和GRBAS依赖专家评分,存在评价者间差异,亟需客观解决方案。本文提出语音质量评估网络VOQANet,利用注意力机制与语音基础模型(SFM)嵌入提取高层特征。进一步提出VOQANet+,融合自监督SFM嵌入与低层声学描述符——抖动、颤音及谐噪比(HNR)。不同于仅基于元音发音的评估(PVQD-A),本模型在元音级与句子级语音(PVQD-S)上均进行测试,以评估泛化能力。实验表明,句子级输入在患者层面表现更优,整体上VOQANet在CAPE-V与GRBAS各维度的均方根误差(RMSE)和皮尔逊相关系数上均优于基线模型,而VOQANet+取得更大提升。此外,该模型在噪声条件下仍保持一致性能,显示其在真实环境与远程医疗中的强鲁棒性。结果表明,结合SFM嵌入与低层特征对实现精准、稳健的病理性嗓音评估具有重要价值。
原文摘要 · Abstract (English)
Perceptual voice quality assessment plays a vital role in diagnosing and monitoring voice disorders. Traditional methods, such as the Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V) and the Grade, Roughness, Breathiness, Asthenia, and Strain (GRBAS) scales, rely on expert raters and are prone to inter-rater variability, emphasizing the need for objective solutions. This study introduces the Voice Quality Assessment Network (VOQANet), a deep learning framework that employs an attention mechanism and Speech Foundation Model (SFM) embeddings to extract high-level features. To further enhance performance, we propose VOQANet+, which integrates self-supervised SFM embeddings with low-level acoustic descriptors-namely jitter, shimmer, and harmonics-to-noise ratio (HNR). Unlike previous approaches that focus solely on vowel-based phonation (PVQD-A), our models are evaluated on both vowel-level and sentence-level speech (PVQD-S) to assess generalizability. Experimental results demonstrate that sentence-based inputs yield higher accuracy, particularly at the patient level. Overall, VOQANet consistently outperforms baseline models in terms of root mean squared error (RMSE) and Pearson correlation coefficient across CAPE-V and GRBAS dimensions, with VOQANet+ achieving even greater performance gains. Additionally, VOQANet+ maintains consistent performance under noisy conditions, suggesting enhanced robustness for real-world and telehealth applications. This work highlights the value of combining SFM embeddings with low-level features for accurate and robust pathological voice assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。