用少量数据快速适配,精准预测大脑对语音的响应。
RABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brain

- 通过脑部微调构建紧凑的音-脑映射模型
- 零样本预测准确率超越现有最先进模型
- 适合需要高效个体适配的脑科学研究
大脑的语言理解具有情境依赖性,因刺激内容和个体差异而异,导致计算模型难以跨人和跨任务泛化。为此,我们提出RABBiT(基于脑部微调的快速自适应BOLD基础模型),一种针对自然语音诱发脑活动的轻量级音-功能磁共振成像(fMRI)编码器。在324名参与者、多个未见的fMRI数据集上评估显示,RABBiT实现了对听觉与语言选择性脑区中语音诱发响应的高精度零样本预测,优于现有的最强基线模型和群体平均模型。仅需10分钟个体化数据,通过参数高效的微调,性能进一步提升,显著优于个性化线性模型。其性能源于两项关键创新:(1) 学习的区域特异性注意力机制;(2) 将脑响应分解为共享与个体特异成分,并结合脑部微调的语音主干网络。该模型生成的结构化、区域特异性表征具备可解释性。无需大量个体数据和建模,支持可扩展的大规模人群语言神经机制分析。代码已开源:https://github.com/bridge-ai-neuro/rabbit。
原文摘要 · Abstract (English)
Language understanding in the brain is context-dependent, varying across experimental stimuli and individuals, which makes it difficult to build computational models that generalize across both. This calls for a foundation model of language-evoked brain activity that can capture shared structure while adapting efficiently to new participants and inputs. We introduce RABBiT (Rapidly Adaptive BOLD foundation model via BraIn-Tuning), a compact audio-to-fMRI encoder designed for accurate zero- and few-shot prediction. A comprehensive evaluation on 324 participants across multiple unseen fMRI datasets shows that RABBiT enables accurate zero-shot prediction of fMRI responses to natural speech across auditory and language-selective regions, surpassing the SOTA foundation model for fMRI and predictions based on group averages. With as little as 10 minutes of participant-specific data, RABBiT further improves performance via parameter-efficient tuning, substantially outperforming per-participant linear models. RABBiT's performance is driven by two key innovations: (1) learned region-specific attention, and (2) a decomposition of brain responses into shared and subject-specific components, combined with a brain-tuned speech backbone. In addition to supporting strong predictive accuracy, the structured, region-specific representations that RABBiT learns enable interpretability. By eliminating the need for extensive per-participant data and model fitting, RABBiT enables scalable population-level analyses of language in the human brain. We make the code available at https://github.com/bridge-ai-neuro/rabbit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。