arXiv:2601.15909cs.CLcs.AI2026-01中稿 · IEEE ISBI 2026被引 1

用图像模型解码脑电想象语音,准确率最高达90.4%。

Transfer Learning from ImageNet for MEG-Based Decoding of Imagined Speech

  • 将脑电信号转为图像样式输入,适配预训练视觉模型。
  • 在21人数据上实现最高90.4%的想象语音识别准确率。
  • 首次证明预训练模型能捕捉跨被试共享的神经表征。

非侵入式想象语音解码因信号微弱、分布分散且标注数据有限而困难。本文提出一种基于图像的方法,将磁脑图(MEG)信号转化为与预训练视觉模型兼容的时间-频率表示。来自21名受试者进行想象语音任务的MEG数据,通过可学习的传感器空间卷积投影为三种空间小波混叠表示,生成紧凑的类图像输入,用于ImageNet预训练视觉架构。这些模型优于传统和未预训练模型,在想象语音与静默对比中达到最高90.4%的平衡准确率,与静默阅读对比达81.0%,元音解码达60.6%。跨被试评估表明,预训练模型捕捉到共享神经表征,时间分析定位到与想象锁定的判别信息区间。结果表明,将预训练视觉模型应用于图像化MEG表示,可有效捕获非侵入性神经信号中的想象语音结构。

原文摘要 · Abstract (English)

Non-invasive decoding of imagined speech remains challenging due to weak, distributed signals and limited labeled data. Our paper introduces an image-based approach that transforms magnetoencephalography (MEG) signals into time-frequency representations compatible with pretrained vision models. MEG data from 21 participants performing imagined speech tasks were projected into three spatial scalogram mixtures via a learnable sensor-space convolution, producing compact image-like inputs for ImageNet-pretrained vision architectures. These models outperformed classical and non-pretrained models, achieving up to 90.4% balanced accuracy for imagery vs. silence, 81.0% vs. silent reading, and 60.6% for vowel decoding. Cross-subject evaluation confirmed that pretrained models capture shared neural representations, and temporal analyses localized discriminative information to imagery-locked intervals. These findings show that pretrained vision models applied to image-based MEG representations can effectively capture the structure of imagined speech in non-invasive neural signals.

脑机接口语音解码迁移学习MEG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。