arXiv:2605.23604eess.AScs.SD2026-05

通过词级对齐的声学融合,提升听力障碍者语音可懂度预测精度。

Word-Level Modeling with Alignment-Aware Acoustic Fusion for Text-Assisted Intelligibility Prediction in Listeners with Hearing Loss

  • 将可懂度预测建模为参考词级正确性判断,利用冻结的Whisper编码器分析受损语音
  • 引入字符级交叉注意力对齐的局部声学分支与全局声学校准分支,提升预测性能
  • 在官方测试集上相关性达0.806,误差降低至24.39,适合听觉康复研究者使用

我们针对听力障碍受试者在CPC3中的文本辅助语音可懂度预测问题展开研究。尽管目标是句子级可懂度百分比,但其实际由参考词识别结果决定。为此,本文提出基于参考词条件的词级正确性建模:采用冻结的Whisper编码器分析退化语音,教师强制解码器以标准转录文作为条件,通过有效参考词的预测正确概率均值得到句子可懂度。为补充转录文条件下的解码器状态,新增基于字符级交叉注意力对齐的词级局部声学分支,以及用于校准的语句级全局声学分支。在官方评估集上,解码器基线达到RMSE 24.92、相关性0.795;联合融合后,错误词F1达0.778,MCC为0.626,相关性提升至0.806,RMSE降至24.39。类似趋势在Whisper medium上亦可见,表明性能提升源于预测粒度细化与对齐感知融合。

原文摘要 · Abstract (English)

We address text-assisted speech intelligibility prediction for hearing-impaired listeners in CPC3. Although the target is a sentence-level percentage, it is determined by reference-word recognition outcomes. We formulate prediction as reference-conditioned word-level correctness modeling: a frozen Whisper encoder analyzes degraded speech, a teacher-forced decoder conditions on the canonical transcript, and sentence intelligibility is obtained by averaging predicted correctness probabilities over valid reference words. To complement transcript-conditioned decoder states, we add a word-aligned local acoustic branch based on character-level cross-attention alignment and an utterance-level global acoustic branch for calibration. On the official evaluation set, the decoder baseline obtains RMSE 24.92 and correlation 0.795, while joint fusion improves to incorrect-word F1 0.778, MCC 0.626, correlation 0.806, and RMSE 24.39. A similar trend with Whisper medium suggests that the gain comes from prediction granularity and alignment-aware fusion.

语音可懂度听力障碍对齐融合Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。