arXiv:2507.05729cs.SDeess.AS2025-07中稿 · INTERSPEECH 2025被引 2

用Mamba替代Transformer,实现低延迟的双耳语音可懂度预测

Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners

  • 采用双向Mamba替代Transformer处理双耳语音时序特征
  • 参数量更少,性能媲美基线模型,适合低功耗设备
  • 特别适合听力障碍者语音理解评估场景

语音可懂度预测(SIP)模型被广泛用于客观评估听力障碍(HI)听众的语音理解能力。在Clarity Prediction Challenge 2(CPC2)中,基于Transformer的非侵入式双耳SIP模型表现出高预测精度。然而,自注意力机制理论上存在高计算与内存开销,成为低延迟、低功耗设备的瓶颈,也可能影响双耳SIP的时序处理。为此,我们提出使用Mamba替代Transformer作为时序处理模块。实验表明,所提SIP模型在保持较少参数的同时,性能与基线相当。分析显示,基于双向Mamba的SIP模型能有效捕捉双耳信号中的上下文与空间语音信息。

原文摘要 · Abstract (English)

Speech intelligibility prediction (SIP) models have been used as objective metrics to assess intelligibility for hearing-impaired (HI) listeners. In the Clarity Prediction Challenge 2 (CPC2), non-intrusive binaural SIP models based on transformers showed high prediction accuracy. However, the self-attention mechanism theoretically incurs high computational and memory costs, making it a bottleneck for low-latency, power-efficient devices. This may also degrade the temporal processing of binaural SIPs. Therefore, we propose Mamba-based SIP models instead of transformers for the temporal processing blocks. Experimental results show that our proposed SIP model achieves competitive performance compared to the baseline while maintaining a relatively small number of parameters. Our analysis suggests that the SIP model based on bidirectional Mamba effectively captures contextual and spatial speech information from binaural signals.

语音可懂度Mamba听力障碍双耳处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。