用脑电波识别说话人,让大脑直接控制语音提取。
TFGA-Net: Temporal-Frequency Graph Attention Network for Brain-Controlled Speaker Extraction
- 构建时空频图注意力网络,融合脑电与语音的共性特征。
- 在Cocktail Party和KUL数据集上超越现有方法,提升语音分离效果。
- 适合脑机接口、听觉解码与语音增强方向的研究者。
基于脑电图(EEG)信号的听觉注意解码(AAD)快速发展,为脑控目标说话人提取提供了可能。然而,如何有效利用脑电与语音之间的共同信息仍是未解难题。本文提出一种脑控说话人提取模型,通过记录听者的脑电信号来提取目标语音。为有效提取脑电特征,我们提取多尺度时频特征,并引入任务中选择性激活的皮层拓扑结构。为充分利用非欧几里得结构并捕捉全局特征,采用图卷积网络与自注意力机制构建脑电编码器。此外,为充分融合脑电与语音特征并保留全局上下文,同时捕捉语音节奏与语调,引入结合MossFormer与无循环递归结构的MossFormer2作为分离器。在公开数据集Cocktail Party与KUL上的实验表明,本模型在部分客观指标上显著优于当前最优方法。源代码已开源:https://github.com/LaoDa-X/TFGA-NET。
原文摘要 · Abstract (English)
The rapid development of auditory attention decoding (AAD) based on electroencephalography (EEG) signals offers the possibility EEG-driven target speaker extraction. However, how to effectively utilize the target-speaker common information between EEG and speech remains an unresolved problem. In this paper, we propose a model for brain-controlled speaker extraction, which utilizes the EEG recorded from the listener to extract the target speech. In order to effectively extract information from EEG signals, we derive multi-scale time--frequency features and further incorporate cortical topological structures that are selectively engaged during the task. Moreover, to effectively exploit the non-Euclidean structure of EEG signals and capture their global features, the graph convolutional networks and self-attention mechanism are used in the EEG encoder. In addition, to make full use of the fused EEG and speech feature and preserve global context and capture speech rhythm and prosody, we introduce MossFormer2 which combines MossFormer and RNN-Free Recurrent as separator. Experimental results on both the public Cocktail Party and KUL dataset in this paper show that our TFGA-Net model significantly outper-forms the state-of-the-art method in certain objective evaluation metrics. The source code is available at: https://github.com/LaoDa-X/TFGA-NET.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。