arXiv:2508.13576eess.AScs.AI2025-08

融合视听信息的端到端模型显著提升人工耳蜗在噪声中的语音可懂度。

End-to-end audio-visual learning for cochlear implant sound coding simulations in noisy environments

  • 构建视听联合训练的端到端人工耳蜗系统,利用视觉线索增强语音信号。
  • 相比传统编码策略,语音信噪比提升7.4666分贝,可懂度显著提高。
  • 适用于听觉障碍者在嘈杂环境下的语音识别研究与临床仿真应用。

人工耳蜗(CI)是帮助重度至极重度听力损失者通过电刺激感知声音的成功生物医学设备,但在噪声环境中聆听仍具挑战性。本文将视听语音增强(AVSE)模块与电极网络编码器(ECS)模型结合,构建端到端的人工耳蜗系统——AVSE-ECS。仿真结果表明,采用联合训练的AVSE-ECS系统在客观语音可懂度方面表现优异,相较于先进组合编码器(ACE)策略,信号-误差比(SER)提升了7.4666 dB。这些发现凸显了基于视听融合的声码编码在人工耳蜗中的巨大潜力。

原文摘要 · Abstract (English)

The cochlear implant (CI) is a successful biomedical device that enables individuals with severe-to-profound hearing loss to perceive sound through electrical stimulation, yet listening in noise remains challenging. Recent deep learning advances offer promising potential for CI sound coding by integrating visual cues. In this study, an audio-visual speech enhancement (AVSE) module is integrated with the ElectrodeNet-CS (ECS) model to form the end-to-end CI system, AVSE-ECS. Simulations show that the AVSE-ECS system with joint training achieves high objective speech intelligibility and improves the signal-to-error ratio (SER) by 7.4666 dB compared to the advanced combination encoder (ACE) strategy. These findings underscore the potential of AVSE-based CI sound coding.

人工耳蜗视听融合语音增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。