arXiv:2504.14107cs.AIcs.CL2025-04被引 4

用Transformer前向传播动态揭示人类认知处理的相似机制

Signatures of human-like processing in Transformer forward passes

  • 通过分析模型各层计算动态,发现其存在类似人类的竞争干扰现象
  • 动态过程指标比最终输出更能预测人类认知行为,提升预测精度
  • 大模型未必更像人,为认知研究提供新视角

现代AI模型正被用于研究人类认知。传统方法仅比较模型输出与人类行为,而本文关注前向传播中各层的动态过程是否反映人类真实认知机制。我们测试了20个开源模型在6个领域中的表现,发现当人类出现决策冲突时,模型初期也会倾向于错误选项,表现出竞争干扰特征。进一步分析表明,模型计算过程的动态特征比最终层的静态输出更能预测人类处理模式。值得注意的是,模型规模增大并未带来更强的人类相似性。该研究提出,AI模型不仅是输入到输出的黑箱,更可作为显式认知过程的模拟工具。

原文摘要 · Abstract (English)

Modern AI models are increasingly being used as theoretical tools to study human cognition. One dominant approach is to evaluate whether human-derived measures are predicted by a model's output: that is, the end-product of a forward pass. However, recent advances in mechanistic interpretability have begun to reveal the internal processes that give rise to model outputs, raising the question of whether models might use human-like processing strategies. Here, we investigate the relationship between real-time processing in humans and layer-time dynamics of computation in Transformers, testing 20 open-source models in 6 domains. We first explore whether forward passes show mechanistic signatures of competitor interference, taking high-level inspiration from cognitive theories. We find that models indeed appear to initially favor a competing incorrect answer in the cases where we would expect decision conflict in humans. We then systematically test whether forward-pass dynamics predict signatures of processing in humans, above and beyond properties of the model's output probability distribution. We find that dynamic measures improve prediction of human processing measures relative to static final-layer measures. Moreover, across our experiments, larger models do not always show more human-like processing patterns. Our work suggests a new way of using AI models to study human cognition: not just as a black box mapping stimuli to responses, but potentially also as explicit processing models.

认知建模Transformer机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。