arXiv:2603.02254cs.SDcs.AI2026-03

用多尺度网络提升脑磁图语音分类准确率

MEBM-Phoneme: Multi-scale Enhanced BrainMagic for End-to-End MEG Phoneme Classification

  • 设计多尺度卷积模块增强脑信号解码能力
  • 在竞赛数据集上达到领先水平的识别准确率
  • 适合研究脑机接口与语音感知的学者参考

我们提出MEBM-Phoneme,一种用于从非侵入式脑磁图(MEG)信号中进行音素分类的多尺度增强神经解码器。基于BrainMagic主干网络,MEBM-Phoneme引入短时多尺度卷积模块,增强原中期编码器,并通过深度可分离卷积实现高效跨尺度融合。采用卷积注意力层动态加权时间依赖关系,优化特征聚合。为应对类别不平衡和会话间分布偏移,引入基于堆叠的局部验证集,结合加权交叉熵损失与随机时间增强。在LibriBrain Competition 2025 Track2上的全面评估显示,模型具备强泛化能力,在验证集与官方测试榜单上均取得具有竞争力的音素解码准确率。结果表明,层次化时间建模与训练稳定性对推进基于MEG的语音感知分析具有重要价值。

原文摘要 · Abstract (English)

We propose MEBM-Phoneme, a multi-scale enhanced neural decoder for phoneme classification from non-invasive magnetoencephalography (MEG) signals. Built upon the BrainMagic backbone, MEBM-Phoneme integrates a short-term multi-scale convolutional module to augment the native mid-term encoder, with fused representations via depthwise separable convolution for efficient cross-scale integration. A convolutional attention layer dynamically weights temporal dependencies to refine feature aggregation. To address class imbalance and session-specific distributional shifts, we introduce a stacking-based local validation set alongside weighted cross-entropy loss and random temporal augmentation. Comprehensive evaluations on LibriBrain Competition 2025 Track2 demonstrate robust generalization, achieving competitive phoneme decoding accuracy on the validation and official test leaderboard. These results underscore the value of hierarchical temporal modeling and training stabilization for advancing MEG-based speech perception analysis.

脑磁图语音识别多尺度建模神经解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。