arXiv:2410.09230cs.CLcs.AI2024-10ICLR被引 38

用大脑活动数据训练语音模型,让其理解更接近人类脑机制。

Improving Semantic Understanding in Speech Language Models via Brain-tuning

  • 用听故事时的fMRI数据微调模型,引入脑相关语义偏差。
  • 模型对新脑数据的语义区域匹配度提升,且减少对语音特征依赖。
  • 提升下游任务表现,表征空间更具语义偏好,适合神经语言研究。

语音语言模型与人类大脑对自然语言的响应具有惊人的一致性。然而,当前模型严重依赖低层语音特征,缺乏与大脑相关的语义理解,限制了其作为大脑语义处理模型生物体的潜力。本文通过使用人们听自然故事时的fMRI记录进行微调,直接将脑相关偏差注入模型,该方法称为脑调优(brain-tuning)。在三个预训练模型族上测试后发现,脑调优不仅提升了模型对新脑记录在语义语言区域的整体一致性,还降低了对低层语音特征的依赖。此外,脑调优显著改善了多种下游任务的表现,并使表示空间具备更强的语义偏好。这是首次提供一致证据,表明将脑信号融入语言模型训练可有效提升其语义理解能力。

原文摘要 · Abstract (English)

Speech language models align with human brain responses to natural language to an impressive degree. However, current models rely heavily on low-level speech features, indicating they lack brain-relevant semantics which limits their utility as model organisms of semantic processing in the brain. In this work, we address this limitation by inducing brain-relevant bias directly into the models via fine-tuning with fMRI recordings of people listening to natural stories, a process we name brain-tuning. After testing it on 3 different pretrained model families, we show that brain-tuning not only improves overall alignment with new brain recordings in semantic language regions, but also reduces the reliance on low-level speech features for this alignment. Excitingly, we further show that brain-tuning leads to 1) consistent improvements in performance on a range of downstream tasks and 2) a representational space with increased semantic preference. Our results provide converging evidence, for the first time, that incorporating brain signals into the training of language models improves the models' semantic understanding.

脑机接口语义理解fMRI语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。