arXiv:2606.26880cs.CLcs.LG2026-06

语言模型能有效预测自然语言理解中的神经活动。

Heterogeneous Neural Predictivity from Language Models During Naturalistic Comprehension

  • 用冻结语言模型分析脑电数据,提取可预测神经信号的特征。
  • 432个有效数据行中67个达到严格预测标准,特征消融影响预测效果。
  • 适合研究神经语言机制或模型与大脑对应关系的学者。

语言模型表示为自然语言刺激提供了结构化的高维注释,可作为理解过程中的有效神经预测因子。我们分析了来自Brain Treebank、MEG-MASC和Podcast ECoG的锁定衍生数据,使用八个冻结的语言模型、分块编码模型,并设置时间、干扰项和表征容量的匹配控制。在源级摘要中,广泛存在对保留预测的正向结果,且优于低层基线。在Brain Treebank和Podcast ECoG中,432个可评估行中有67行满足受控的仅预测标准;模型侧特征消融在多数可评估源行中改变了预测分数。基于大脑来源、时间关联、声学及植入信号的控制验证了分析流程的组件级敏感性。这些发现表明,语言模型导出的量可标注自然语音与文本理解中的神经活动。参与者级别匹配控制的优势是局部而非全局的,响应模式与特征特异性对比限制了表征或计算解释。完全共索引的整合解释需未来联合索引覆盖。整体而言,分析确认语言模型特征具有神经预测价值,并将预测效用与共享神经组织或语言处理计算的主张区分开来。

原文摘要 · Abstract (English)

Language-model representations provide structured, high-dimensional annotations of naturalistic language stimuli and can serve as informative neural predictors during comprehension. We analyzed locked derived data from Brain Treebank, MEG-MASC, and Podcast ECoG with eight frozen language models, blocked encoding models, and matched temporal, nuisance, and representation-capacity controls. Positive held-out prediction and gains over low-level baselines were widespread in source-level summaries. Across Brain Treebank and Podcast ECoG, 67 of 432 evaluable rows met a controlled predictive-only criterion, and model-side feature ablations changed prediction scores in most evaluable source rows. Brain-derived, timing-linked, acoustic, and implanted-signal controls confirmed component-level sensitivity of the analysis pipeline. These findings show that language-model-derived quantities can annotate neural activity during natural speech and text comprehension. Participant-level matched-control advantages were localized rather than uniform, response-profile and feature-specificity contrasts bounded representational or computational interpretations, and complete co-indexed integrated interpretation will require future jointly indexed coverage. Together, the analyses identify language-model features as useful neural predictors and separate predictive usefulness from claims about shared neural organization or language-processing computations.

神经语言语言模型预测建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。