arXiv:2502.15786q-bio.NCcs.AI2025-02ICML被引 17

用脑科学启发的模型实现跨人种的脑信号转文本,性能提升超20%。

MindLLM: A Subject-Agnostic and Versatile Model for fMRI-to-Text Decoding

  • 用脑科学启发的注意力机制处理不同形状输入,实现跨人泛化。
  • 在未见受试者上提升24.5%,新任务适应能力提高25.0%。
  • 可解释性强,适合脑机接口与神经机制研究者使用。

将功能磁共振成像(fMRI)信号解码为文本是神经科学领域的重要挑战,具有推动脑机接口和揭示大脑工作机制的潜力。然而,现有方法常面临预测性能不足、任务类型有限及跨受试者泛化能力差的问题。为此,我们提出MindLLM,一个面向无特定主题且多功能的fMRI-to-text解码模型。该模型由fMRI编码器与现成大语言模型(LLM)构成。fMRI编码器采用脑科学启发的注意力机制,能适应不同输入形状,实现高性能跨人泛化解码。此外,我们提出脑指令微调(Brain Instruction Tuning, BIT),增强模型从fMRI信号中捕捉多样语义表征的能力,支持更灵活的解码。我们在全面的fMRI-to-text基准上评估了MindLLM,结果表明其优于基线模型,在下游任务上提升12.0%,未见受试者泛化能力提升24.5%,新任务适应能力提升25.0%。同时,模型注意力模式提供了可解释的决策过程洞察。

原文摘要 · Abstract (English)

Decoding functional magnetic resonance imaging (fMRI) signals into text has been a key challenge in the neuroscience community, with the potential to advance brain-computer interfaces and uncover deeper insights into brain mechanisms. However, existing approaches often struggle with suboptimal predictive performance, limited task variety, and poor generalization across subjects. In response to this, we propose MindLLM, a model designed for subject-agnostic and versatile fMRI-to-text decoding. MindLLM consists of an fMRI encoder and an off-the-shelf LLM. The fMRI encoder employs a neuroscience-informed attention mechanism, which is capable of accommodating subjects with varying input shapes and thus achieves high-performance subject-agnostic decoding. Moreover, we introduce Brain Instruction Tuning (BIT), a novel approach that enhances the model's ability to capture diverse semantic representations from fMRI signals, facilitating more versatile decoding. We evaluate MindLLM on comprehensive fMRI-to-text benchmarks. Results demonstrate that our model outperforms the baselines, improving downstream tasks by 12.0%, unseen subject generalization by 24.5%, and novel task adaptation by 25.0%. Furthermore, the attention patterns in MindLLM provide interpretable insights into its decision-making process.

脑机接口fMRI解码大模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。