arXiv:2603.23091cs.CL2026-03中稿 · ICLR被引 1

让大模型'大脑错位',发现脑对齐影响语言理解能力

When Language Models Lose Their Mind: The Consequences of Brain Misalignment

  • 故意训练模型预测脑活动差,但保持语言能力高
  • 200多个任务测试显示,脑错位模型表现显著下降
  • 适合关注模型认知机制与安全性的研究者

尽管脑对齐的大语言模型(LLMs)因其作为认知模型的潜力以及提升AI安全性与可信度的前景而受到关注,但这种脑对齐对语言能力的影响仍不明确。本文通过引入脑错位模型——即有意训练为预测脑活动表现差但仍保持高水平语言建模能力的模型——来探究脑对齐的功能影响。我们在超过200个涵盖语义、句法、话语、推理和形态学等多元语言领域的下游任务上评估这些模型,并将其与匹配良好的脑对齐模型进行对比,以分离脑对齐对语言理解的具体作用。实验结果表明,脑错位会显著损害下游性能,凸显了脑对齐在实现稳健语言能力中的关键作用。这些发现强调了脑对齐在大语言模型中的重要性,并为神经表征与语言处理之间的关系提供了新见解。

原文摘要 · Abstract (English)

While brain-aligned large language models (LLMs) have garnered attention for their potential as cognitive models and for potential for enhanced safety and trustworthiness in AI, the role of this brain alignment for linguistic competence remains uncertain. In this work, we investigate the functional implications of brain alignment by introducing brain-misaligned models--LLMs intentionally trained to predict brain activity poorly while maintaining high language modeling performance. We evaluate these models on over 200 downstream tasks encompassing diverse linguistic domains, including semantics, syntax, discourse, reasoning, and morphology. By comparing brain-misaligned models with well-matched brain-aligned counterparts, we isolate the specific impact of brain alignment on language understanding. Our experiments reveal that brain misalignment substantially impairs downstream performance, highlighting the critical role of brain alignment in achieving robust linguistic competence. These findings underscore the importance of brain alignment in LLMs and offer novel insights into the relationship between neural representations and linguistic processing.

大模型脑对齐语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。