arXiv:2608.01481cs.LGcs.SD2026-08

用脑磁图解码听觉感知,揭示驱动恢复的关键语音特征与神经源。

Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval

论文配图:Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
图 1 · 摘自论文原文
  • 重构模型前端与解码器,实现脑信号到语音的可解释映射。
  • 在MEG-MASC数据集上达到39.75%准确率,参数量减少20倍。
  • 发现静默、音强、元音和声学起始点是核心驱动因素。

通过深度网络利用非侵入性脑磁图(MEG)记录,可从短段感知语音中重建音频,训练目标为对比wav2vec 2.0音频嵌入。但现有模型权重无法对应电生理量,且驱动重建的语音特性不明确。本文改进高性能的MEG-to-audio重建架构,将空间注意力从平面传感器布局替换为基于三维头盔几何的球谐函数;将受试者特异性表示从270分支缩减至25个,每个分支添加时空匹配的时序滤波器,并简化卷积解码器。训练前去除眼动与心电成分以避免刺激锁定偏差。在MEG-MASC数据集上,模型在6个训练方案下对1005个候选音频的Top-1准确率达39.75%±0.34%,解码器参数量减少约20倍。模型权重成功映射至源空间,恢复出与语音感知网络一致的神经生成器,左侧分支携带更高频节律成分。配对的脑磁图遮蔽实验显示,19个刺激特征中有15个贡献显著,影响最大的为静默、音强、元音及声学起始点。随机词表则相反:用叙事语料的脑信号替换其内容反而提升重建效果,表明无语篇结构的活动信息可恢复性更低。将wav2vec目标压缩至约12个学习特征维度仍保持准确率,而强时间压缩导致明显下降。源空间映射与输入干预共同揭示了重建的核心驱动力。

原文摘要 · Abstract (English)

Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map onto electrophysiological quantities, and it remains unclear which speech properties drive retrieval. We build on a high-performing MEG-to-audio retrieval architecture but redesign both its front end and decoder. Its spatial attention operates on a flattened sensor layout; we replace it with spherical harmonics defined on the three-dimensional MEG helmet geometry. We reduce the subject-specific representation from 270 to 25 branches, add a temporal filter to each branch to match it to a neuronal source in space and time, and make the convolutional decoder shallower. Ocular and cardiac components are removed before training to reduce the risk of stimulus-locked shortcuts. On MEG-MASC, the model reaches 39.75 +/- 0.34% Top-1 accuracy among 1005 candidates across six trained solutions, with about 20 times fewer decoder parameters. Its weights map to source space, recovering generators consistent with the speech-perception network, while left-lateralized branches carry higher-frequency rhythmic components not evident on the right. Paired MEG occlusion shows that 15 of 19 stimulus features contribute, with the largest effects for silence, sound intensity, vowels, and acoustic onsets. Random word lists behave oppositely: substituting narrative MEG into them improves retrieval, indicating that activity without narrative structure carries less recoverable information than activity during coherent speech. The wav2vec target can be reduced to about twelve learned feature dimensions without loss of accuracy, whereas strong temporal compression causes a clear loss. Together, source mapping and input interventions reveal what drives retrieval.

脑机接口语音解码可解释性神经源映射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。