arXiv:2509.17883cs.SDcs.LG2025-09

用脑图调制音频分离,让助听器更懂用户注意力

Brainprint-Modulated Target Speaker Extraction

  • 通过脑电特征学习个性化脑图嵌入,动态调节语音分离
  • 在KUL和鸡尾酒会数据集上显著超越现有方法
  • 适合关注个性化助听与神经接口的科研人员

实现鲁棒且个性化的神经引导目标说话人提取(TSE)是下一代助听器的关键挑战。主要源于两个因素:跨会话的脑电(EEG)信号固有非平稳性,以及高个体间差异导致通用模型效果受限。为此,我们提出脑图调制的目标说话人提取(BM-TSE)框架,实现个性化高保真语音分离。该框架首先采用时空联合的EEG编码器,结合自适应频谱增益(ASG)模块,提取对非平稳性具有鲁棒性的特征。核心在于个性化调制机制:在主体识别(SID)与听觉注意解码(AAD)联合监督下学习统一脑图嵌入,该嵌入同时编码静态用户特征与动态注意力状态,主动优化音频分离过程,动态适配每位用户。在公开的KUL和鸡尾酒会数据集上的评估表明,BM-TSE达到当前最优性能,显著优于已有方法。代码已开源:https://github.com/rosshan-orz/BM-TSE。

原文摘要 · Abstract (English)

Achieving robust and personalized performance in neuro-steered Target Speaker Extraction (TSE) remains a significant challenge for next-generation hearing aids. This is primarily due to two factors: the inherent non-stationarity of EEG signals across sessions, and the high inter-subject variability that limits the efficacy of generalized models. To address these issues, we propose Brainprint-Modulated Target Speaker Extraction (BM-TSE), a novel framework for personalized and high-fidelity extraction. BM-TSE first employs a spatio-temporal EEG encoder with an Adaptive Spectral Gain (ASG) module to extract stable features resilient to non-stationarity. The core of our framework is a personalized modulation mechanism, where a unified brainmap embedding is learned under the joint supervision of subject identification (SID) and auditory attention decoding (AAD) tasks. This learned brainmap, encoding both static user traits and dynamic attentional states, actively refines the audio separation process, dynamically tailoring the output to each user. Evaluations on the public KUL and Cocktail Party datasets demonstrate that BM-TSE achieves state-of-the-art performance, significantly outperforming existing methods. Our code is publicly accessible at: https://github.com/rosshan-orz/BM-TSE.

脑机接口语音分离个性化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。