arXiv:2502.01615cs.CL2025-02Transactions of th…被引 42

大模型内部表征更像人脑,早期层对应快速眼动,晚期层对应慢速认知信号。

Large Language Models Are Human-Like Internally

  • 从模型内部层分析,大模型预测下一词概率更贴近人类阅读行为。
  • 在自定步速阅读、眼动时长、MAZE任务和N400脑电波中均表现良好匹配。
  • 发现模型早期层对应快速眼动,后期层对应慢速认知反应,启发跨学科研究。

近期认知建模研究指出,更大的语言模型(LMs)与人类阅读行为的拟合度较差(Oh and Schuler, 2023b;Shain et al., 2024;Kuribayashi et al., 2024),从而质疑其认知合理性。本文通过机制可解释性视角重新审视该论点,指出先前结论因仅关注模型末层而产生偏差。我们的分析表明,更大语言模型内部层生成的下一个词概率,与人类句子处理数据的拟合程度不亚于甚至优于小模型。这一一致性在行为(自定步速阅读时间、注视时长、MAZE任务处理时间)和神经生理(N400脑电位)指标上均成立,挑战了早期混杂结果,提示大模型的认知合理性被低估。此外,我们首次揭示模型层与人类测量间的有趣关系:早期层更接近快速注视时长,后期层则与较慢的信号如N400电位和MAZE处理时间更匹配。本工作为机制可解释性与认知建模的交叉研究开辟新路径。

原文摘要 · Abstract (English)

Recent cognitive modeling studies have reported that larger language models (LMs) exhibit a poorer fit to human reading behavior (Oh and Schuler, 2023b; Shain et al., 2024; Kuribayashi et al., 2024), leading to claims of their cognitive implausibility. In this paper, we revisit this argument through the lens of mechanistic interpretability and argue that prior conclusions were skewed by an exclusive focus on the final layers of LMs. Our analysis reveals that next-word probabilities derived from internal layers of larger LMs align with human sentence processing data as well as, or better than, those from smaller LMs. This alignment holds consistently across behavioral (self-paced reading times, gaze durations, MAZE task processing times) and neurophysiological (N400 brain potentials) measures, challenging earlier mixed results and suggesting that the cognitive plausibility of larger LMs has been underestimated. Furthermore, we first identify an intriguing relationship between LM layers and human measures: earlier layers correspond more closely with fast gaze durations, while later layers better align with relatively slower signals such as N400 potentials and MAZE processing times. Our work opens new avenues for interdisciplinary research at the intersection of mechanistic interpretability and cognitive modeling.

认知建模机制可解释大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。