arXiv:2510.21520cs.CL2025-10NeurIPS被引 4

用多人脑数据微调语音模型,提升泛化与效率。

Brain-tuning Improves Generalizability and Efficiency of Brain Alignment in Speech Models

  • 用多参与者fMRI数据联合微调预训练语音模型。
  • 所需数据量减少5倍,脑对齐度提升50%。
  • 适合跨被试研究与神经科学-人工智能交叉应用。

预训练语言模型在对齐人类大脑对自然语言的响应方面表现优异,是研究语言加工的理想模型。然而,现有方法依赖特定被试且受个体数据量限制,阻碍了跨被试泛化与群体分析。本文提出一种可扩展、通用的脑微调方法,通过联合预测多个被试的fMRI响应来微调预训练语音模型。结果表明,该方法不仅保持强个体对齐,还实现跨被试泛化:1)预测新被试脑数据所需fMRI数据量减少5倍;2)整体脑对齐度最高提升50%;3)在未见数据集上仍具强泛化能力。此外,多被试脑微调还提升了下游语义任务性能,表明利用多被试脑数据可学习更通用的语义表征。这些发现体现了神经科学与AI的双向互益,有助于弥合两领域鸿沟。代码与模型公开于https://github.com/bridge-ai-neuro/multi-brain-tuning。

原文摘要 · Abstract (English)

Pretrained language models are remarkably effective in aligning with human brain responses elicited by natural language stimuli, positioning them as promising model organisms for studying language processing in the brain. However, existing approaches for both estimating and improving this brain alignment are participant-dependent and highly affected by the amount of data available per participant, hindering both generalization to new participants and population-level analyses. In this work, we address these limitations by introducing a scalable, generalizable brain-tuning method, in which we fine-tune pretrained speech language models to jointly predict fMRI responses from multiple participants. We demonstrate that the resulting brain-tuned models exhibit strong individual brain alignment while generalizing across participants. Specifically, our method leads to 1) a 5-fold decrease in the amount of fMRI data needed to predict brain data from new participants, 2) up to a 50% increase in the overall brain alignment, and 3) strong generalization to new unseen datasets. Furthermore, this multi-participant brain-tuning additionally improves downstream performance on semantic tasks, suggesting that training using brain data from multiple participants leads to more generalizable semantic representations. Taken together, these findings demonstrate a bidirectional benefit between neuroscience and AI, helping bridge the gap between the two fields. We make our code and models publicly available at https://github.com/bridge-ai-neuro/multi-brain-tuning.

脑科学语音模型多被试fMRI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。