arXiv:2608.08159cs.AI2026-08被引 1

检验大模型概念可操控性时,测量方法影响结果,需谨慎对待所谓‘类脑’现象。

When Is a Steerable Concept Representation Real? Measurement Confounds in a Cross-Family Audit of Neuroscience Parallels in LLMs

  • 通过统一测量协议审计5大模型家族17个版本,发现结果受方法选择显著影响。
  • 原始激活单元与读出指标不校准会导致规模相关假象,修正后趋势消失。
  • 虽在多模型中发现地理地图、数字编码等现象,但跨语言结构方向会因方法改变而反转。

大语言模型常被报道具备类人神经与认知特征,如概念细胞、心理数轴和认知地图。这些结论多基于单模型的线性探测与激活操控,但两类方法对测量选择极为敏感。报告的类比可能源于模型本身、测量过程或两者共同作用。本文对四个神经科学启发范式在17个来自五个模型家族(参数量0.6B至72B)的模型上进行跨家族审计。主实验考察概念方向的因果可控性:使用原始激活单元、固定层与系数时,操控效果随模型规模上升,似具涌现特性;但该趋势实由未校准流程导致,非文献中公认的结论。调整任一因素——原始单元、读出指标或操作点——即消除该趋势。采用残差归一化可比干预与预留操作点后,各规模下概念操控仍显著,但通义千问3系列无显著趋势,仅置信区间不排除中等正斜率。其余结果混杂:地理世界图在所有测试检查点(最高72B)均能一致解码;数字大小编码强烈存在,但单神经元形状(钟形或单调)取决于选择标准;语言特异性结构可定位,但跨语言不对称方向在不同归因方法下反转。结果表明,当前人工智能神经科学研究的主要瓶颈并非缺乏现象,而是缺乏可比测量与充分控制。代码、协议与刺激数据已公开。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly reported to exhibit human-like neural and cognitive signatures, including concept cells, mental number lines, and cognitive maps. These claims often rely on linear probing and activation steering applied to a single model, yet both methods are highly sensitive to measurement choices. A reported parallel may therefore reflect the model, the measurement procedure, or both. We audit four representative neuroscience-inspired paradigms across 17 models from five families, spanning $0.6$B to $72$B parameters. Our main experiment examines the causal steerability of concept directions. With raw activation units and a fixed layer and coefficient, steerability appears to increase with model scale, resembling an emergent capability. However, this pattern is produced by an uncalibrated pipeline rather than by a claim established in the steering literature. The trend depends jointly on raw units, the readout metric, and the operating point; correcting any one of these removes it. With residual-norm-comparable interventions and held-out operating-point selection, concept steering remains significant at every scale, but shows no significant trend across the Qwen3 series, although the confidence interval does not rule out a moderate positive slope. The remaining results are mixed. A linear geographic world map is consistently decodable in every tested checkpoint up to $72$B. Number magnitude is strongly encoded, but whether individual neurons appear bell-shaped or monotonic depends on the selection criterion. Language-specific structure is localizable, but the direction of the cross-lingual asymmetry reverses under a different attribution method. These results suggest that the main constraint on AI neuroscience is not a lack of phenomena, but a lack of comparable measurements and adequate controls. We release the protocol, stimuli, and code.

大模型类脑机制可解释性测量偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。