arXiv:2510.12355cs.CL2025-10被引 4

通过词级归因分析,发现脑-模型对齐与语言预测依赖不同词汇。

Fine-grained Analysis of Brain-LLM Alignment through Input Attribution

  • 提出细粒度输入归因方法,定位影响脑-模型对齐的关键词语。
  • 发现脑对齐关注语义和语篇,而语言预测偏重句法和近期词。
  • 方法可推广至其他语言任务,揭示模型预测的认知相关性。

理解大语言模型(LLMs)与人类大脑活动之间的对齐关系,有助于揭示语言处理的计算原理。本文提出一种细粒度输入归因方法,识别对脑-模型对齐最重要的具体词语,并用于探讨一个有争议的问题:脑对齐(BA)与下一词预测(NWP)之间的关系。研究发现,BA与NWP依赖的词集差异显著:NWP表现出近期效应和首词偏好,侧重句法;而BA则更关注语义和语篇层面信息,且仅有较弱的近期效应。该工作深化了我们对大模型与人类语言处理关系的理解,揭示了二者在特征依赖上的本质差异。此外,本方法可广泛应用于探索各类语言任务中模型预测的认知相关性。

原文摘要 · Abstract (English)

Understanding the alignment between large language models (LLMs) and human brain activity can reveal computational principles underlying language processing. We introduce a fine-grained input attribution method to identify the specific words most important for brain-LLM alignment, and leverage it to study a contentious research question about brain-LLM alignment: the relationship between brain alignment (BA) and next-word prediction (NWP). Our findings reveal that BA and NWP rely on largely distinct word subsets: NWP exhibits recency and primacy biases with a focus on syntax, while BA prioritizes semantic and discourse-level information with a more targeted recency effect. This work advances our understanding of how LLMs relate to human language processing and highlights differences in feature reliance between BA and NWP. Beyond this study, our attribution method can be broadly applied to explore the cognitive relevance of model predictions in diverse language processing tasks.

脑机对齐语言模型归因分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。