arXiv:2512.19374cs.SD2025-12

无需参考语音,用深度学习预测听障者语音可懂度

DeepGESI: A Non-Intrusive Objective Evaluation Model for Predicting Speech Intelligibility in Hearing-Impaired Listeners

  • 基于深度学习构建非侵入式模型,不依赖原始语音
  • 在CPC2数据集上预测的GESI与真实值高度相关
  • 速度远超传统方法,适合实际应用场景

语音可懂度评估对诸多语音应用至关重要。然而,现有客观指标多为侵入式,需同时具备清晰参考语音与受损信号才能评估。此外,如STOI等指标主要针对正常听力者设计,对听障者的预测精度有限。尽管GESI可用于听障者可懂度估计,但其仍需参考信号,限制了实际应用。为此,本文提出DeepGESI,一种基于深度学习的非侵入式模型,可在无任何干净参考语音的情况下准确高效预测听障者语音可懂度。实验结果表明,在2nd Clarity Prediction Challenge(CPC2)数据集测试条件下,DeepGESI预测的GESI得分与真实值具有强相关性,且预测速度显著快于传统方法。

原文摘要 · Abstract (English)

Speech intelligibility assessment is essential for many speech-related applications. However, most objective intelligibility metrics are intrusive, as they require clean reference speech in addition to the degraded or processed signal for evaluation. Furthermore, existing metrics such as STOI are primarily designed for normal hearing listeners, and their predictive accuracy for hearing impaired speech intelligibility remains limited. On the other hand, the GESI (Gammachirp Envelope Similarity Index) can be used to estimate intelligibility for hearing-impaired listeners, but it is also intrusive, as it depends on reference signals. This requirement limits its applicability in real-world scenarios. To overcome this limitation, this study proposes DeepGESI, a non-intrusive deep learning-based model capable of accurately and efficiently predicting the speech intelligibility of hearing-impaired listeners without requiring any clean reference speech. Experimental results demonstrate that, under the test conditions of the 2nd Clarity Prediction Challenge(CPC2) dataset, the GESI scores predicted by DeepGESI exhibit a strong correlation with the actual GESI scores. In addition, the proposed model achieves a substantially faster prediction speed compared to conventional methods.

语音可懂度听障评估深度学习非侵入式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。