让大脑活动直接参与语言模型,实现可交互的神经解码
Brain-language fusion enables interactive neural readout and in-silico experimentation
- 将脑电数据嵌入大语言模型隐空间,实现自然语言交互
- 零样本泛化能力超越训练时的语义类别,回答更精准
- 适合神经科学与脑机接口研究者,推动可逆向实验的虚拟脑刺激
大型语言模型(LLMs)已革新人机交互,通过将图像等多模态信息融入共享语言空间得以拓展。然而,神经解码仍受限于静态、非交互的方法。我们提出CorText框架,将功能性磁共振成像(fMRI)数据直接嵌入大语言模型的潜在空间,实现对脑数据的开放式自然语言交互。该模型在观看自然场景时记录的fMRI数据上训练,能生成准确的图像描述,并在仅访问神经数据的情况下,比对照组更好地回答复杂问题。结果显示,CorText具备零样本泛化能力,可超越训练中出现的语义类别。基于虚拟微刺激的反事实提示实验揭示了脑状态与语言输出之间一致且渐进的映射关系。这些进展标志着从被动解码向生成式、灵活的脑-语言接口转变。
原文摘要 · Abstract (English)
Large language models (LLMs) have revolutionized human-machine interaction, and have been extended by embedding diverse modalities such as images into a shared language space. Yet, neural decoding has remained constrained by static, non-interactive methods. We introduce CorText, a framework that integrates neural activity directly into the latent space of an LLM, enabling open-ended, natural language interaction with brain data. Trained on fMRI data recorded during viewing of natural scenes, CorText generates accurate image captions and can answer more detailed questions better than controls, while having access to neural data only. We showcase that CorText achieves zero-shot generalization beyond semantic categories seen during training. In-silico microstimulation experiments, which enable counterfactual prompts on brain activity, reveal a consistent, and graded mapping between brain-state and language output. These advances mark a shift from passive decoding toward generative, flexible interfaces between brain activity and language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。