用脑电图增强语音分离,提升目标说话人识别效果。
Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction
- 双路径Mamba处理语音长短序列,KAN网络提取相关脑电特征。
- 在KUL和AVED数据集上相对性能提升达36%和29%(SI-SDR)。
- 适合脑机接口、听觉注意力研究及多源语音分离方向的开发者。
近年来,听觉注意解码(AAD)的快速发展使得利用脑电图(EEG)作为目标说话人分离的辅助信息成为可能。然而,有效建模长序列语音并从脑电信号中解析出目标说话人身份仍是主要挑战。本文提出一种改进的特征提取网络(IFENet),包含采用双路径Mamba的语音编码器和基于柯尔莫戈洛夫-阿诺德网络(KAN)的脑电编码器。我们提出SpeechBiMamba,通过双路径Mamba分别建模语音的局部与全局序列以提取特征;同时提出EEGKAN,有效提取与听觉刺激紧密相关的脑电特征,并利用受试者注意力信息定位目标说话人。在KUL和AVED数据集上的实验表明,IFENet在开放评估条件下相较最先进模型,于尺度无关信干比(SI-SDR)指标上分别取得36%和29%的相对提升。
原文摘要 · Abstract (English)
The recent rapid development of auditory attention decoding (AAD) offers the possibility of using electroencephalography (EEG) as auxiliary information for target speaker extraction. However, effectively modeling long sequences of speech and resolving the identity of the target speaker from EEG signals remains a major challenge. In this paper, an improved feature extraction network (IFENet) is proposed for neuro-oriented target speaker extraction, which mainly consists of a speech encoder with dual-path Mamba and an EEG encoder with Kolmogorov-Arnold Networks (KAN). We propose SpeechBiMamba, which makes use of dual-path Mamba in modeling local and global speech sequences to extract speech features. In addition, we propose EEGKAN to effectively extract EEG features that are closely related to the auditory stimuli and locate the target speaker through the subject's attention information. Experiments on the KUL and AVED datasets show that IFENet outperforms the state-of-the-art model, achieving 36\% and 29\% relative improvements in terms of scale-invariant signal-to-distortion ratio (SI-SDR) under an open evaluation condition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。