多尺度神经解码器实现高精度脑磁图语音检测
MEBM-Speech: Multi-scale Enhanced BrainMagic for Robust MEG Speech Detection
- 融合多尺度卷积、双向LSTM与深度可分离卷积,建模长短时语音模式
- 在LibriBrain 2025基准上达到89.3%平均F1宏值,边界稳定性强
- 适合脑科学与临床语音障碍研究,对噪声鲁棒性好
我们提出MEBM-Speech,一种基于BrainMagic骨干网络的多尺度增强神经解码器,用于从非侵入式脑磁图(MEG)信号中进行语音活动检测。该模型整合了三种互补的时间建模机制:用于短期模式提取的多尺度卷积模块、用于长程上下文建模的双向LSTM(BiLSTM),以及用于高效跨尺度特征融合的深度可分离卷积层。结合轻量级时间抖动策略与平均池化,进一步提升起始时刻鲁棒性与边界稳定性。模型实现对MEG信号的连续概率解码,可精细区分语音与静默状态,对认知神经科学和临床应用均具重要意义。在LibriBrain Competition 2025 Track1基准上的全面评估显示,验证集上平均F1宏值达89.3%,测试排行榜结果相当。这些结果表明,多尺度时间表征学习在鲁棒的MEG语音解码中具有显著有效性。
原文摘要 · Abstract (English)
We propose MEBM-Speech, a multi-scale enhanced neural decoder for speech activity detection from non-invasive magnetoencephalography (MEG) signals. Built upon the BrainMagic backbone, MEBM-Speech integrates three complementary temporal modeling mechanisms: a multi-scale convolutional module for short-term pattern extraction, a bidirectional LSTM (BiLSTM) for long-range context modeling, and a depthwise separable convolutional layer for efficient cross-scale feature fusion. A lightweight temporal jittering strategy and average pooling further improve onset robustness and boundary stability. The model performs continuous probabilistic decoding of MEG signals, enabling fine-grained detection of speech versus silence states - an ability crucial for both cognitive neuroscience and clinical applications. Comprehensive evaluations on the LibriBrain Competition 2025 Track1 benchmark demonstrate strong performance, achieving an average F1 macro of 89.3% on the validation set and comparable results on the official test leaderboard. These findings highlight the effectiveness of multi-scale temporal representation learning for robust MEG-based speech decoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。