通过模拟听觉损失的多层级影响,提升语音可懂度预测精度。
Modeling Multi-Level Hearing Loss for Speech Intelligibility Prediction
- 基于耳蜗滤波器展宽和调制低通滤波,模拟听力损失的频率与时间分辨率下降。
- 在轻度和中重度听力损失组中,相比HASPI v2,预测误差分别降低16.5%和6.1%。
- 适合研究听觉感知、助听器评估或语音理解建模的科研人员参考。
听力损失的多种感知后果严重阻碍语音交流,但传统临床测听仅关注阈值频率敏感性,未能充分反映频率与时间分辨率的缺陷。为此,本文提出一种语音可懂度预测方法,通过展宽耳蜗滤波器并施加低通调制滤波,显式模拟不同严重程度的听觉退化。随后利用谱时调制(STM)表示分析语音信号,反映听觉分辨率下降对调制结构的影响。同时,采用归一化互相关(NCC)矩阵量化干净语音与噪声中语音的STM表示相似性。这些听觉引导特征用于训练基于视觉变压器的回归模型,融合STM图与NCC嵌入以估计可懂度得分。在Clarity Prediction Challenge语料库上的评估表明,该方法在轻度和中重度听力损失组中均优于听觉辅助语音感知指数v2(HASPI v2),相对均方根误差分别降低16.5%和6.1%。结果表明,显式建模个体化频率与时间分辨率退化对提升语音可懂度预测至关重要,并具备听觉失真可解释性。
原文摘要 · Abstract (English)
The diverse perceptual consequences of hearing loss severely impede speech communication, but standard clinical audiometry, which is focused on threshold-based frequency sensitivity, does not adequately capture deficits in frequency and temporal resolution. To address this limitation, we propose a speech intelligibility prediction method that explicitly simulates auditory degradations according to hearing loss severity by broadening cochlear filters and applying low-pass modulation filtering to temporal envelopes. Speech signals are subsequently analyzed using the spectro-temporal modulation (STM) representations, which reflect how auditory resolution loss alters the underlying modulation structure. In addition, normalized cross-correlation (NCC) matrices quantify the similarity between the STM representations of clean speech and speech in noise. These auditory-informed features are utilized to train a Vision Transformer-based regression model that integrates the STM maps and NCC embeddings to estimate speech intelligibility scores. Evaluations on the Clarity Prediction Challenge corpus show that the proposed method outperforms the Hearing-Aid Speech Perception Index v2 (HASPI v2) in both mild and moderate-to-severe hearing loss groups, with a relative root mean squared error reduction of 16.5% for the mild group and a 6.1% reduction for the moderate-to-severe group. These results highlight the importance of explicitly modeling listener-specific frequency and temporal resolution degradations to improve speech intelligibility prediction and provide interpretability in auditory distortions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。