arXiv:2605.08075cs.LGeess.AS2026-05

用听觉脑信号映射想象言语,实现零样本解码

Zero-Shot Imagined Speech Decoding via Imagined-to-Listened MEG Mapping

论文配图:Zero-Shot Imagined Speech Decoding via Imagined-to-Listened MEG Mapping
图 1 · 摘自论文原文
  • 用聆听时的脑信号训练模型,反推想象时的神经活动
  • 在未参与训练的受试者上仍能显著高于随机水平解码想象词
  • 适合脑机接口研究者,尤其关注非侵入式语言解码场景

从非侵入性脑记录中解码想象言语极具挑战,因想象数据稀缺且跨被试、跨会话的时间对齐困难。本文提出一种新方法,利用聆听言语时更丰富且可靠标注的脑信号。研究采集了训练过的音乐家在聆听旋律和口语刺激时的配对聆听与想象MEG数据,借助音乐家的共性提升时间对齐精度。构建三阶段解码流程:首先训练六种线性和神经网络模型,将想象脑信号映射到聆听响应;通过未见被试的零基准验证,确认预测的聆听响应保留刺激特异性信息;第二阶段在聆听数据上训练对比词解码器,评估四种嵌入策略(语义、声学、音素);第三阶段将保留被试的想象脑信号经映射流程处理,生成对应聆听响应并由监听解码器解码。基于秩分析,发现想象词可显著高于随机水平解码。报告了概念验证结果,所有评估均在保留被试上进行。同时证明性能随训练数据量增加而提升,表明该方法具有可扩展性,可直接应用于真实脑机接口场景。

原文摘要 · Abstract (English)

Decoding imagined speech from non-invasive brain recordings is challenging because imagined datasets are scarce and difficult to align temporally across subjects and sessions In this work, we propose a new approach to the decoding of imagined speech that leverages the richer and more reliably labeled recordings during listening to speech. We collected paired listened and imagined MEG recordings to rhythmic melodic and spoken stimuli from trained musicians. Using trained musicians helped improve temporal alignment across conditions. We then developed a three-stage decoding pipeline that revealed consistent and meaningful relationships between neural activity evoked by imagining and listening to the same stimuli. First, we trained six linear and neural models to map imagined MEG responses to listened responses. We evaluated these models against a null baseline from unseen subjects to validate that the predicted-listening responses preserve stimulus-specific information. In the second stage, we trained a contrastive word decoder exclusively on the listened MEG responses, and evaluated it using four embedding strategies including semantic, acoustic, and phonetic representations. In the third stage, we process the imagined MEG responses from held-out subjects through the mapping pipeline to compute the corresponding listening responses that are then decoded by the listened decoder. Using rank-based analysis, we show that the imagined words are decodable significantly above chance. We shall report here the results of a proof-of-concept implementation to decode imagined speech, where all evaluations are performed on held-out subjects. We also demonstrate that performance improves with training data size, suggesting that this approach is scalable and can directly be made applicable to realistic brain-computer interface scenarios.

脑机接口想象言语MEG零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。