arXiv:2606.02907cs.CLcs.AI2026-06中稿 · ACL被引 7

线性探测显示模型区分推理类型,实则受任务格式干扰

Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States

  • 用线性探测分析大模型隐藏状态的推理模式
  • 高准确率分离源于任务格式而非真实推理结构
  • 适合关注模型可解释性与格式偏差的研究者

对大语言模型(LLM)隐藏状态进行线性探测,常被用来声称模型为不同推理类型学习了独立表征。我们以 Qwen3-14B 模型在三个基准测试上验证该结论:LogiQA 2.0(演绎)、ARC-Challenge(归纳)和 $α$NLI(溯因)。在 40 层中的第 32 层,线性探测达到 100% 交叉验证准确率,且几何结构明显分离(内在维度分别为 20.6、28.5、33.6;凸包污染 ≤1.5%)。然而,这种分离完全由任务格式混杂因素驱动。去除源身份、选项数量和响应长度后,准确率降至随机水平。轨迹锚点相似性表明任务间推理高度共享(42.5% 相似度,随机为 33.3%),因果操控实验(n=20)也未发现几何结构与推理模式间的功能关联(p=0.286)。因此,高探测准确率反映的是任务格式,而非计算结构,提示机制可解释性研究应常规去除格式混淆。

原文摘要 · Abstract (English)

Linear probing of large language model (LLM) hidden states is widely used to claim that models learn distinct representations for different reasoning types. We test this by probing Qwen3-14B on three benchmarks spanning the classical trichotomy: LogiQA 2.0 (deductive), ARC-Challenge (inductive), and $α$NLI (abductive). At layer 32 of 40, linear probes achieve 100\% cross-validated accuracy with well-separated geometry (intrinsic dimensionalities: 20.6, 28.5, 33.6; convex hull contamination $\leq$1.5\%). However, this separation is entirely driven by format confounds. Residualizing source identity, option count, and response length reduces accuracy to chance. Trace-anchor similarity indicates largely shared reasoning across tasks (42.5\% agreement vs.\ 33.3\% chance), and causal steering with random controls ($n=20$) shows no functional link between geometry and reasoning mode ($p=0.286$). Thus, high probe accuracy reflects task format rather than computational structure, motivating routine format deconfounding in mechanistic interpretability.

模型可解释性线性探测任务格式推理模式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。