arXiv:2501.03246q-bio.NCcs.CL2025-01中稿 · ICLR

用脑磁图揭示听觉与语言处理的神经路径差异。

Bridging Auditory Perception and Language Comprehension through MEG-Driven Encoding Models

  • 构建音频与文本到脑活动的编码模型,预测脑响应
  • 文本模型比音频模型相关性更高,达显著提升
  • 听觉信息走侧颞叶直接通道,语言信息激活额叶高级区

理解听觉与语言加工的神经机制对认知神经科学至关重要。本研究利用脑磁图(MEG)数据分析大脑对口语刺激的反应,构建两类编码模型:基于时间-频率分解(TFD)和wav2vec2隐空间表示的音频→MEG编码器,以及基于CLIP和GPT-2嵌入的文本→MEG编码器。两者均能有效预测神经活动,且文本→MEG模型表现更优,相关性显著更高。空间上,音频嵌入主要激活侧颞叶,负责初级听觉处理与信号整合;而文本嵌入则主要激活额叶皮层,尤其是布罗卡区,在8–30 Hz频段尤为明显,与语义整合及语言生成相关。结果表明,听觉信息通过直接感官通路处理,而语言信息则依赖融合意义与认知控制的网络进行编码。该研究揭示了听觉与语言处理的不同神经路径,提升了复杂语言刺激下神经响应建模的准确性。

原文摘要 · Abstract (English)

Understanding the neural mechanisms behind auditory and linguistic processing is key to advancing cognitive neuroscience. In this study, we use Magnetoencephalography (MEG) data to analyze brain responses to spoken language stimuli. We develop two distinct encoding models: an audio-to-MEG encoder, which uses time-frequency decompositions (TFD) and wav2vec2 latent space representations, and a text-to-MEG encoder, which leverages CLIP and GPT-2 embeddings. Both models successfully predict neural activity, demonstrating significant correlations between estimated and observed MEG signals. However, the text-to-MEG model outperforms the audio-based model, achieving higher Pearson Correlation (PC) score. Spatially, we identify that auditory-based embeddings (TFD and wav2vec2) predominantly activate lateral temporal regions, which are responsible for primary auditory processing and the integration of auditory signals. In contrast, textual embeddings (CLIP and GPT-2) primarily engage the frontal cortex, particularly Broca's area, which is associated with higher-order language processing, including semantic integration and language production, especially in the 8-30 Hz frequency range. The strong involvement of these regions suggests that auditory stimuli are processed through more direct sensory pathways, while linguistic information is encoded via networks that integrate meaning and cognitive control. Our results reveal distinct neural pathways for auditory and linguistic information processing, with higher encoding accuracy for text representations in the frontal regions. These insights refine our understanding of the brain's functional architecture in processing auditory and textual information, offering quantitative advancements in the modelling of neural responses to complex language stimuli.

脑机接口语言神经机制MEG建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。