arXiv:2601.22260eess.ASeess.SP2026-01

用脑电波引导听觉分离,让人工耳蜗更懂用户注意力。

Brain-Informed Speech Separation for Cochlear Implants

  • 结合脑电注意力信号与音频,动态聚焦目标说话人。
  • 在多人对话中提升信干比,参数量仅略小167k vs. 171k。
  • 适合需要认知适应性处理的听障患者,低延迟易部署。

我们提出一种面向人工耳蜗(CIs)的脑信息引导语音分离方法,利用脑电图(EEG)提取的注意力线索,指导增强过程聚焦于用户关注的说话人。一个轻量级融合网络将音频混合信号与EEG特征结合,生成用于耳蜗刺激的目标语音电极图,同时解决纯音频分离器的标签混淆问题。通过混合课程训练策略提升对劣质注意力线索的鲁棒性,在脑电-语音相关性中等时仍保持稳定性能。在多说话人场景下,该模型的信干比提升优于纯音频基线,且模型规模更小(167k vs. 171k参数)。算法延迟仅为2毫秒,成本相近,表明融合听觉与神经信号在认知自适应耳蜗处理中的巨大潜力。

原文摘要 · Abstract (English)

We propose a brain-informed speech separation method for cochlear implants (CIs) that uses electroencephalography (EEG)-derived attention cues to guide enhancement toward the attended speaker. An attention-guided network fuses audio mixtures with EEG features through a lightweight fusion layer, producing attended-source electrodograms for CI stimulation while resolving the label-permutation ambiguity of audio-only separators. Robustness to degraded attention cues is improved with a mixed curriculum that varies cue quality during training, yielding stable gains even when EEG-speech correlation is moderate. In multi-talker conditions, the model achieves higher signal-to-interference ratio improvements than an audio-only electrodogram baseline while remaining slightly smaller (167k vs. 171k parameters). With 2 ms algorithmic latency and comparable cost, the approach highlights the promise of coupling auditory and neural cues for cognitively adaptive CI processing.

人工耳蜗脑机接口语音分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。