arXiv:2509.16979cs.SDcs.AI2025-09

用语音增强器辅助预测听障者语音可懂度,无需干净参考信号。

Leveraging Multiple Speech Enhancers for Non-Intrusive Intelligibility Prediction for Hearing-Impaired Listeners

  • 利用多个语音增强器构建增强信号路径,实现无参考信号的可懂度预测。
  • 集成强增强器效果最佳,在多数据集上超越基线模型。
  • 通过双片段增强提升跨数据集泛化能力,适合真实场景应用。

针对听障者语音可懂度评估,传统方法依赖听觉测试或需要干净参考信号的侵入式指标(如HASPI),但在真实场景中参考信号常不可得,导致实验室与实际评估脱节。为此,本文提出一种非侵入式可懂度预测框架,借助三个先进语音增强器构建并行增强信号路径,实现无需参考信号的鲁棒预测。实验表明,预测性能依赖于增强器选择,强增强器的集成效果最优。为进一步提升跨数据集泛化能力,引入2片段增强策略以增强听者特异性变异性,显著提高在未见数据集上的鲁棒性。该方法在多个数据集上持续优于非侵入式基线模型CPC2 Champion,验证了增强器引导的非侵入式预测在真实应用中的潜力。

原文摘要 · Abstract (English)

Speech intelligibility evaluation for hearing-impaired (HI) listeners is essential for assessing hearing aid performance, traditionally relying on listening tests or intrusive methods like HASPI. However, these methods require clean reference signals, which are often unavailable in real-world conditions, creating a gap between lab-based and real-world assessments. To address this, we propose a non-intrusive intelligibility prediction framework that leverages speech enhancers to provide a parallel enhanced-signal pathway, enabling robust predictions without reference signals. We evaluate three state-of-the-art enhancers and demonstrate that prediction performance depends on the choice of enhancer, with ensembles of strong enhancers yielding the best results. To improve cross-dataset generalization, we introduce a 2-clips augmentation strategy that enhances listener-specific variability, boosting robustness on unseen datasets. Our approach consistently outperforms the non-intrusive baseline, CPC2 Champion across multiple datasets, highlighting the potential of enhancer-guided non-intrusive intelligibility prediction for real-world applications.

语音增强可懂度预测听障评估非侵入式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。