arXiv:2606.12818cs.CLcs.AI2026-06

揭示大模型在数字推理中受无关数字干扰的内在路径机制

Localizing Anchoring Pathways in Language Models

  • 通过对比正确答案与锚点答案的逻辑差,定位模型内部敏感信号
  • 边缘级方法比节点级更准确还原锚定效应信号,且高低锚点路径高度共享
  • 基础模型与指令微调模型间路径转移弱,说明训练阶段改变关键路径

提示词中的无关数字会改变语言模型的判断,引发数值推理中的锚定效应。我们通过带有共享选项的可控多选任务,研究该锚定敏感信号在模型内部的传递路径。定义了对正确答案与锚点对应答案之间的逻辑差度量,并验证其能有效追踪行为上的锚定效应。基于7B–8B规模的Qwen和Llama基础模型及指令微调模型,采用基于归因的电路定位方法发现,边缘级方法比节点级方法更忠实还原该信号。低锚点与高锚点的电路在模型内具有强可转移性,表明锚定路径存在共通结构;但基础模型与指令微调版本间的稀疏转移性较差,说明后训练过程改变了哪些路径更重要。整体结果为锚定相关决策信号在语言模型内部的承载机制提供了可解释的机理解释。

原文摘要 · Abstract (English)

Irrelevant numbers in a prompt can shift language model judgments, producing anchoring effects in numerical reasoning. We study where this anchor-sensitive signal is carried inside language models using a controlled multiple-choice setup with shared answer options. We define a logit-difference metric comparing the correct answer option with the answer option corresponding to the anchor, and validate that it tracks behavioral anchoring. Using attribution-based circuit localization on 7B--8B Qwen and Llama base and instruction-tuned models, we find that edge-level methods recover this signal more faithfully than node-level methods. Low- and high-anchor circuits transfer strongly within a model, suggesting shared pathway structure across anchor direction. However, sparse transfer across base and instruction-tuned variants is less reliable, indicating that post-training changes which pathways matter most. Overall, our results provide a mechanistic account of how anchoring-related decision signals are carried inside language models.

模型可解释性锚定效应路径定位语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。