arXiv:2502.12771cs.CLq-bio.NC2025-02

用非线性多模态模型提升脑电预测语言理解的能力

Mind the Gap: Aligning the Brain with Language Models Requires a Nonlinear and Multimodal Approach

  • 融合音频与语言特征,采用非线性映射建模大脑响应
  • 预测性能比传统线性模型提升17.2%至17.9%
  • 揭示听觉与语义信息在多个脑区的融合机制

自监督语言与音频模型能有效预测大脑对语音的反应。然而,传统预测模型依赖单模态特征的线性映射,忽略了听觉信号与语言、语义信息在广泛脑网络中复杂整合的过程。本文提出一种非线性、多模态预测模型,结合预训练模型(如LLAMA、Whisper)提取的音频与语言特征。该方法在未归一化和归一化相关性上分别较传统单模态线性模型提升17.2%和17.9%,较先前最先进模型提升7.7%和14.4%。这些改进为未来虚拟脑机接口测试和解码性能提升奠定了基础,并揭示了听觉与语义信息在运动皮层、体感皮层及高级语义区域的融合方式,与现有神经语言学理论一致。研究强调了非线性与多模态方法在脑建模中的潜力,推动自然情境下的神经语言学研究迈向新阶段。

原文摘要 · Abstract (English)

Self-supervised language and audio models effectively predict brain responses to speech. However, traditional prediction models rely on linear mappings from unimodal features, despite the complex integration of auditory signals with linguistic and semantic information across widespread brain networks during speech comprehension. Here, we introduce a nonlinear, multimodal prediction model that combines audio and linguistic features from pre-trained models (e.g., LLAMA, Whisper). Our approach achieves a 17.2% and 17.9% improvement in prediction performance (unnormalized and normalized correlation) over traditional unimodal linear models, as well as a 7.7% and 14.4% improvement, respectively, over prior state-of-the-art models. These improvements represent a major step towards future robust in-silico testing and improved decoding performance. They also reveal how auditory and semantic information are fused in motor, somatosensory, and higher-level semantic regions, aligning with existing neurolinguistic theories. Overall, our work highlights the often neglected potential of nonlinear and multimodal approaches to brain modeling, paving the way for future studies to embrace these strategies in naturalistic neurolinguistics research.

脑机接口多模态模型非线性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。