用脑磁图信号重建语音-静默序列,提升神经解码精度
SHINE: Sequential Hierarchical Integration Network for EEG and MEG
- 构建分层序列网络,融合多尺度脑信号特征
- 在50小时数据上实现0.918的F1宏分,优于基线模型
- 适合脑机接口与语音神经机制研究者参考
自然语言在大脑中的表征是认知神经科学的核心挑战,皮层对语音包络的跟随反应在语音解码中起关键作用。本文针对LibriBrain 2025竞赛中的语音检测任务,利用单个受试者聆听LibriVox有声书时采集的超过50小时脑磁图(MEG)信号,提出一种用于脑电与脑磁图的分层序列集成网络(SHINE),以从MEG信号中重建二值语音-静默序列。在扩展赛道中,进一步引入语音包络和梅尔频谱图的辅助重建任务以增强训练。结合基线模型(BrainMagic、AWavNet、ConvConcatNet)的集成方法,在排行榜测试集上分别取得0.9155(标准赛道)和0.9184(扩展赛道)的F1宏分。
原文摘要 · Abstract (English)
How natural speech is represented in the brain constitutes a major challenge for cognitive neuroscience, with cortical envelope-following responses playing a central role in speech decoding. This paper presents our approach to the Speech Detection task in the LibriBrain Competition 2025, utilizing over 50 hours of magnetoencephalography (MEG) signals from a single participant listening to LibriVox audiobooks. We introduce the proposed Sequential Hierarchical Integration Network for EEG and MEG (SHINE) to reconstruct the binary speech-silence sequences from MEG signals. In the Extended Track, we further incorporated auxiliary reconstructions of speech envelopes and Mel spectrograms to enhance training. Ensemble methods combining SHINE with baselines (BrainMagic, AWavNet, ConvConcatNet) achieved F1-macro scores of 0.9155 (Standard Track) and 0.9184 (Extended Track) on the leaderboard test set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。