arXiv:2510.22860cs.CLq-bio.NC2025-10NeurIPS被引 5

分离语言模型中的推理成分,揭示大脑深层思维的神经机制。

Far from the Shallow: Brain-Predictive Reasoning Embedding through Residual Disentanglement

  • 通过残差解耦方法,从语言模型中分离出词法、句法、语义和推理四类独立表征。
  • 推理表征能预测其他语言特征无法解释的大脑活动,且激活时间晚于350毫秒。
  • 适用于研究语言理解中深层认知过程的神经基础,尤其适合关注推理机制的研究者。

理解人类大脑如何从处理简单语言输入演进到高阶推理,是神经科学的核心挑战。现代大语言模型(LLMs)虽被用于建模语言刺激下的神经反应,但其内部表征高度“纠缠”,混合了词汇、句法、语义和推理信息。这种纠缠导致传统脑编码分析偏向语言浅层特征(如词汇和句法),难以剥离深层认知过程的神经基础。本文提出一种残差解耦方法,通过探测语言模型中特定特征的层级,迭代回归低层次表示,生成词法、句法、语义与关键的推理四类近正交嵌入。我们使用这些解耦嵌入建模神经外科患者听自然语言时的颅内皮层脑电图(ECoG)数据。结果表明:1)孤立的推理嵌入具有独特预测能力,可解释其他语言特征未涵盖的神经活动变异,甚至延伸至经典语言区外的视觉区域;2)推理的神经信号在时间上明显延迟,峰值出现在约350-400毫秒,符合其在加工层级中的顶层位置;3)标准非解耦的LLM嵌入可能具有误导性,其预测成功主要源于浅层语言特征,掩盖了深层认知处理的细微贡献。

原文摘要 · Abstract (English)

Understanding how the human brain progresses from processing simple linguistic inputs to performing high-level reasoning is a fundamental challenge in neuroscience. While modern large language models (LLMs) are increasingly used to model neural responses to language, their internal representations are highly "entangled," mixing information about lexicon, syntax, meaning, and reasoning. This entanglement biases conventional brain encoding analyses toward linguistically shallow features (e.g., lexicon and syntax), making it difficult to isolate the neural substrates of cognitively deeper processes. Here, we introduce a residual disentanglement method that computationally isolates these components. By first probing an LM to identify feature-specific layers, our method iteratively regresses out lower-level representations to produce four nearly orthogonal embeddings for lexicon, syntax, meaning, and, critically, reasoning. We used these disentangled embeddings to model intracranial (ECoG) brain recordings from neurosurgical patients listening to natural speech. We show that: 1) This isolated reasoning embedding exhibits unique predictive power, accounting for variance in neural activity not explained by other linguistic features and even extending to the recruitment of visual regions beyond classical language areas. 2) The neural signature for reasoning is temporally distinct, peaking later (~350-400ms) than signals related to lexicon, syntax, and meaning, consistent with its position atop a processing hierarchy. 3) Standard, non-disentangled LLM embeddings can be misleading, as their predictive success is primarily attributable to linguistically shallow features, masking the more subtle contributions of deeper cognitive processing.

脑机接口语言模型推理机制神经解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。