用语言理解大脑活动,让fMRI能被AI像文本一样解析。
fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding
- 将fMRI信号转为语言一致的离散符号,构建脑-语关联基础
- 零样本与少样本表现强,支持多种任务且可高效微调
- 适合神经科学、脑机接口研究者探索脑认知机制
多模态大模型在图像、音频、视频上实现了统一推理,但脑成像领域的拓展仍处空白。为弥合这一差距,我们提出fMRI-LM,一种通过三阶段框架连接功能磁共振(fMRI)与语言的基础模型。第一阶段学习一个神经分词器,将fMRI映射到语言一致空间的离散符号;第二阶段将预训练语言模型适配为联合建模fMRI符号与文本,将脑活动视为可时序预测和语言描述的序列。为解决自然的fMRI-文本对稀缺问题,我们构建了一个大规模描述语料库,将多样影像特征转化为结构化文本描述,捕捉fMRI信号的低层组织。第三阶段采用多任务、多范式指令微调,赋予fMRI-LM高层次语义理解能力,支持多样化下游应用。在多个基准测试中,fMRI-LM展现出强劲的零样本与少样本性能,并可通过参数高效微调(LoRA)快速适应,建立了一条通往语言对齐、通用的结构性与语义性fMRI理解的可扩展路径。
原文摘要 · Abstract (English)
Recent advances in multimodal large language models (LLMs) have enabled unified reasoning across images, audio, and video, but extending such capability to brain imaging remains largely unexplored. Bridging this gap is essential to link neural activity with semantic cognition and to develop cross-modal brain representations. To this end, we present fMRI-LM, a foundational model that bridges functional MRI (fMRI) and language through a three-stage framework. In Stage 1, we learn a neural tokenizer that maps fMRI into discrete tokens embedded in a language-consistent space. In Stage 2, a pretrained LLM is adapted to jointly model fMRI tokens and text, treating brain activity as a sequence that can be temporally predicted and linguistically described. To overcome the lack of natural fMRI-text pairs, we construct a large descriptive corpus that translates diverse imaging-based features into structured textual descriptors, capturing the low-level organization of fMRI signals. In Stage 3, we perform multi-task, multi-paradigm instruction tuning to endow fMRI-LM with high-level semantic understanding, supporting diverse downstream applications. Across various benchmarks, fMRI-LM achieves strong zero-shot and few-shot performance, and adapts efficiently with parameter-efficient tuning (LoRA), establishing a scalable pathway toward a language-aligned, universal model for structural and semantic understanding of fMRI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。