用注意力机制提升人工耳蜗用户在噪声中的语音可懂度
DAT-CFTNet: Speech Enhancement for Cochlear Implant Recipients using Attention-based Dual-Path Recurrent Neural Network
- 设计双路径注意力RNN,精准区分语音与噪声的时频区域
- 在复杂噪声下显著提升语音可懂度,优于现有模型
- 特别适合人工耳蜗使用者,抑制非平稳噪声且无音乐噪声
人类听觉系统能动态聚焦语音关键成分,忽略背景噪声或失真。受注意力模型成功启发,本研究在并行语音增强网络瓶颈层引入双路径注意力模块。提出基于注意力的双路径循环神经网络(DAT-RNN),与改进的复数域频域变换网络(CFTNet)结合,构成DAT-CFTNet。该机制可精确区分时频谱图中语音与噪声区域,优化局部与全局上下文信息处理。实验表明,DAT-CFTNet在语音可懂度和质量上持续优于现有模型,包括CFTNet和DCCRN。尤其在人工耳蜗(CI)用户中表现优异,其时频听力恢复受限(>10%)的情况下仍有效提升语音可懂度,能抑制非平稳噪声,避免传统方法常见的音乐噪声问题。模型实现将公开发布。
原文摘要 · Abstract (English)
The human auditory system has the ability to selectively focus on key speech elements in an audio stream while giving secondary attention to less relevant areas such as noise or distortion within the background, dynamically adjusting its attention over time. Inspired by the recent success of attention models, this study introduces a dual-path attention module in the bottleneck layer of a concurrent speech enhancement network. Our study proposes an attention-based dual-path RNN (DAT-RNN), which, when combined with the modified complex-valued frequency transformation network (CFTNet), forms the DAT-CFTNet. This attention mechanism allows for precise differentiation between speech and noise in time-frequency (T-F) regions of spectrograms, optimizing both local and global context information processing in the CFTNet. Our experiments suggest that the DAT-CFTNet leads to consistently improved performance over the existing models, including CFTNet and DCCRN, in terms of speech intelligibility and quality. Moreover, the proposed model exhibits superior performance in enhancing speech intelligibility for cochlear implant (CI) recipients, who are known to have severely limited T-F hearing restoration (e.g., >10%) in CI listener studies in noisy settings show the proposed solution is capable of suppressing non-stationary noise, avoiding the musical artifacts often seen in traditional speech enhancement methods. The implementation of the proposed model will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。