arXiv:2501.10402eess.SPcs.SD2025-01中稿 · ICASSP 2025被引 13

用状态空间模型从脑电波重建连续语音频谱,提升解码精度。

SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG

  • 基于Mamba模块建模长序列脑电信号,结合S4-UNet提取局部特征
  • 在SparrKULee数据集上相关系数达0.069,较基线提升38%
  • 适合脑机接口与言语神经机制研究者参考

从脑信号解码语音是一项具有重要意义的研究,有助于理解大脑中的言语处理机制。尽管已有研究利用非侵入式脑电图(EEG)在单词或字母级别成功重建被试感知音频的梅尔频谱,但在连续语音特征的精确重建方面仍存在关键差距,尤其在分钟级层面。为此,本文提出一种基于状态空间模型(SSM)的Mel频谱重建方法——SSM2Mel。该模型引入新型Mamba模块,有效建模想象语音的长序列脑电信号;采用S4-UNet结构增强脑电信号局部特征提取能力,并通过嵌入强度调制器(ESM)模块融入个体特异性信息。实验结果表明,本模型在SparrKULee数据集上的皮尔逊相关系数达到0.069,相比先前基线提升38%。

原文摘要 · Abstract (English)

Decoding speech from brain signals is a challenging research problem that holds significant importance for studying speech processing in the brain. Although breakthroughs have been made in reconstructing the mel spectrograms of audio stimuli perceived by subjects at the word or letter level using noninvasive electroencephalography (EEG), there is still a critical gap in precisely reconstructing continuous speech features, especially at the minute level. To address this issue, this paper proposes a State Space Model (SSM) to reconstruct the mel spectrogram of continuous speech from EEG, named SSM2Mel. This model introduces a novel Mamba module to effectively model the long sequence of EEG signals for imagined speech. In the SSM2Mel model, the S4-UNet structure is used to enhance the extraction of local features of EEG signals, and the Embedding Strength Modulator (ESM) module is used to incorporate subject-specific information. Experimental results show that our model achieves a Pearson correlation of 0.069 on the SparrKULee dataset, which is a 38% improvement over the previous baseline.

脑机接口语音重建状态空间模型EEG解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。