arXiv:2606.20696cs.CLcs.AI2026-06

用fMRI解码内心话语,不改语言模型也能生成完整句子。

MindAlign: Decoding Inner Speech from fMRI Signals via Multimodal Embedding Alignment under Limited Data

论文配图:MindAlign: Decoding Inner Speech from fMRI Signals via Multimodal Embedding Alignment under Limited Data
图 1 · 摘自论文原文
  • 分两阶段:先对齐脑信号与语义空间,再用视觉提示生成文本
  • 在静默图像描述任务中性能超越纯fMRI基线
  • 跨被试通用,适合快速适配新参与者

从非侵入性脑信号解码内心话语仍面临根本挑战,原因包括缺乏外显语言输出、训练数据有限以及个体间差异大。现有脑到文本方法通常依赖特定任务的解码器微调,限制了可扩展性并增加新被试适配难度。我们提出MindAlign,一种解耦的两阶段脑到语言框架,可在不修改底层语言模型的前提下实现开放式文本生成。第一阶段学习个体特异的神经-语义对齐,将fMRI活动映射到共享多模态语义空间,提取内部生成句子的潜在语义草图。第二阶段将该草图与视觉上下文结合,提示冻结的多模态语言模型进行自由形式生成。在静默图像描述任务的fMRI数据上实验表明,该方法持续优于fMRI-only和随机基线。进一步证明所学的语义到语言投影可跨被试泛化,配合个体神经对齐后实现有效解码。结果表明神经信号调节语义内容,超出图像驱动先验,支持脑到文本解码的可扩展与模块化方向。

原文摘要 · Abstract (English)

Decoding inner speech from non-invasive brain signals remains a fundamental challenge due to the absence of overt linguistic output, limited training data, and large inter-subject variability. Existing brain-to-text approaches often rely on task-specific decoder fine-tuning, which restricts scalability and complicates adaptation to new participants. We propose MindAlign, a decoupled two-stage brain-to-language framework that enables open-ended text generation from fMRI signals without modifying the underlying language model. The first stage learns a subject-specific neural-semantic alignment that maps fMRI activity into a shared multimodal semantic space, extracting a latent semantic sketch of the internally generated sentence. The second stage integrates this sketch with visual context to prompt a frozen multimodal language model for free-form generation. Experiments on fMRI data collected during silent image description demonstrate that the proposed approach consistently outperforms fMRI-only and random baselines. We further show that the learned semantic-to-language projection can generalize across subjects, enabling effective decoding when paired with subject-specific neural alignment. These results indicate that neural signals modulate semantic content beyond image-driven priors, supporting a scalable and modular direction for brain-to-text decoding.

脑机接口fMRI解码多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。