arXiv:2409.02489cs.SDcs.AI2024-09被引 1

用脑电波引导语音分离,让系统听懂你注意力所在。

NeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention

  • 用脑电图信号编码注意力信息,作为语音分离的唯一参考
  • 跨模态注意力机制提升语音特征表示,生成更精准的分离掩码
  • 在公开数据集上优于现有方法,适合关注人机交互与神经接口的研究者

在听觉注意力研究中发现,被关注的语音与诱发的神经反应之间存在强相关性,可通过脑电图(EEG)测量。因此,可利用听众的脑电信号中的注意力信息,计算上引导从鸡尾酒会场景中提取目标说话人语音。本文提出一种神经引导的语音分离模型 NeuroSpex,仅使用听众的 EEG 信号作为辅助参考,从单声道语音混合中提取关注的语音。我们设计了一种新颖的 EEG 信号编码器以捕捉注意力信息,并引入跨模态注意力(CA)机制增强语音特征表示,生成语音分离掩码。在公开数据集上的实验结果表明,所提模型在多个评估指标上均优于两个基线模型。

原文摘要 · Abstract (English)

In the study of auditory attention, it has been revealed that there exists a robust correlation between attended speech and elicited neural responses, measurable through electroencephalography (EEG). Therefore, it is possible to use the attention information available within EEG signals to guide the extraction of the target speaker in a cocktail party computationally. In this paper, we present a neuro-guided speaker extraction model, i.e. NeuroSpex, using the EEG response of the listener as the sole auxiliary reference cue to extract attended speech from monaural speech mixtures. We propose a novel EEG signal encoder that captures the attention information. Additionally, we propose a cross-attention (CA) mechanism to enhance the speech feature representations, generating a speaker extraction mask. Experimental results on a publicly available dataset demonstrate that our proposed model outperforms two baseline models across various evaluation metrics.

语音分离脑电图跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。