用Mamba模型提升人工耳蜗语音清晰度,降噪效果更优。
TokenSE: a Mamba-based discrete token speech enhancement framework for cochlear implants
- 基于Mamba架构的离散编码器语音增强框架,线性计算复杂度优于Transformer
- 在真实噪声混响环境下显著提升人工耳蜗用户的语音可懂度
- 适合听觉辅助设备如人工耳蜗和助听器的实时语音处理
语音增强(SE)对改善现实环境中语音可懂性和质量至关重要,尤其对于在噪声和混响条件下语音理解严重受损的人工耳蜗(CI)用户。本文提出TokenSE,一种基于离散编码器的语音增强框架,运行于神经音频编解码器空间,利用Mamba模型从退化语音中预测干净语音的编码词元索引。与早期依赖自注意力机制、序列长度增长导致计算复杂度呈二次增长的Transformer不同,Mamba通过输入依赖选择机制实现线性复杂度,是适用于CI和助听器(HA)应用的有力替代方案。客观评估显示,TokenSE在域内和域外数据集上均持续优于基线方法。此外,针对CI用户的主观听觉实验表明,在恶劣噪声和混响环境下,语音可懂度有明显提升。
原文摘要 · Abstract (English)
Speech enhancement (SE) is critical for improving speech intelligibility and quality in real-world environments, particularly for cochlear implant (CI) users who experience severe degradations in speech understanding under noisy and reverberant conditions. In this study, we propose TokenSE, a discrete token-based SE framework operating in the neural audio codec space, which predicts clean codec token indices from degraded speech using a Mamba-based model. Unlike the earlier Transformer architecture, whose self-attention mechanism has a computational complexity that grows quadratically with sequence length, the input-dependent selection mechanism of Mamba achieves linear complexity, making it a compelling alternative to Transformers, especially for CI and hearing-aid (HA) applications. Objective evaluations show that TokenSE consistently outperforms baseline methods on both in-domain and out-of-domain datasets. Moreover, subjective listening experiments with CI users indicate clear benefit in speech intelligibility under adverse noisy and reverberant environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。