arXiv:2504.14437eess.AS2025-04中稿 · publication in Spe…被引 3

基于听觉机制的语音可懂度预测模型,更准确评估老年人语音增强效果。

Predicting speech intelligibility in older adults for speech enhancement using the Gammachirp Envelope Similarity Index, GESI

  • 融合听觉生理特征与调制分析,构建从外周到中枢的语音可懂度预测模型。
  • 在日语词和英语句级测试中,对老年听障者可懂度预测精度优于HASPIw2和HASPIv2。
  • 当前时间调制功能建模尚不成熟,需改进测量方法与模型整合方式。

本文提出一种客观语音可懂度衡量指标(OIM)——Gammachirp包络相似性指数(GESI),用于预测老年听障者的语音可懂度(SI)。GESI基于从外周到中枢听觉系统的心理声学知识,利用伽马啁啾滤波组(GCFB)、调制滤波组及扩展余弦相似度计算单一可懂度指标。该模型不仅考虑了听力图反映的听力水平,还纳入了由时间调制传递函数(TMTF)捕捉的时间处理特性。通过在具有不同听力水平的老年受试者上进行语音-噪声环境下理想语音增强的可懂度实验,评估了其性能。结果表明,相较于专为关键词可懂度设计的HASPIw2,GESI能更准确预测主观可懂度评分;在英语句子级可懂度预测中,其表现至少与HASPIv2相当,甚至更优。然而,引入TMTF对模型性能影响不显著,提示目前的TMTF测量与建模尚未成熟,未来可能需要采用带通噪声测量,并改进时间特征的建模方式。

原文摘要 · Abstract (English)

We propose an objective intelligibility measure (OIM), called the Gammachirp Envelope Similarity Index (GESI), that can predict speech intelligibility (SI) in older adults. GESI is a bottom-up model based on psychoacoustic knowledge from the peripheral to the central auditory system. It computes the single SI metric using the gammachirp filterbank (GCFB), the modulation filterbank, and the extended cosine similarity measure. It takes into account not only the hearing level represented in the audiogram, but also the temporal processing characteristics captured by the temporal modulation transfer function (TMTF). To evaluate performance, SI experiments were conducted with older adults of various hearing levels using speech-in-noise with ideal speech enhancement on familiarity-controlled Japanese words. The prediction performance was compared with HASPIw2, which was developed for keyword SI prediction. The results showed that GESI predicted the subjective SI scores more accurately than HASPIw2. GESI was also found to be at least as effective as, if not more effective than, HASPIv2 in predicting English sentence-level SI. The effect of introducing TMTF into the GESI algorithm was insignificant, suggesting that TMTF measurements and models are not yet mature. Therefore, it may be necessary to perform TMTF measurements with bandpass noise and to improve the incorporation of temporal characteristics into the model.

语音可懂度听觉模型老年听障语音增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。