arXiv:2609.04463cs.CLcs.AI2026-09

通过分析模型内部电路,预测大模型在算术题中跨格式泛化能力。

Shared circuits predict whether LLMs generalize across formats in arithmetic reasoning

论文配图:Shared circuits predict whether LLMs generalize across formats in arithmetic reasoning
图 1 · 摘自论文原文
  • 用归因修补法定位模型处理数字和文字算术题的独立神经电路。
  • 模型对文字算术题的正确率与数字电路重叠度强相关。
  • 无需标注数据即可预测泛化表现,适合研究模型可解释性的人。

在多种推理任务中,包括算术推理,人类能轻松应对输入格式的表面变化:会解2+5的人自然也会解‘two plus five’。而大语言模型对提示的表面变化则更为脆弱:虽然对数值算术问题几乎全对,但在文字形式的问题上准确率显著下降。本文探讨能否从模型内部结构预测这种跨格式泛化能力。我们采用归因修补法,分别定位模型在三种语言中处理数值算术(2+5)与文字算术(英语‘two plus five’,西班牙语‘dos m'as cinco’,意大利语‘due più cinque’)所依赖的神经电路;随后检验模型自身数值电路的重叠程度是否能预测其对文字格式的泛化能力。结果表明,三层次证据支持该假设:电路重叠度可解释三种文字格式的相对难度,以及模型在哪些题目上表现更好,其预测效果媲美有监督探测器,且无需标注数据。

原文摘要 · Abstract (English)

In many forms of reasoning, including arithmetic reasoning, generalizing across superficial changes in input format is effortless for humans: anyone who can solve 2+5 can also solve 'two plus five'. In contrast, LLMs are more brittle to surface variations of the prompts: for example, they solve numeric arithmetic problems almost perfectly but are substantially less accurate on verbal renditions of the same problems. Here, we ask whether generalization across formats can be predicted from the models' internals. Using attribution patching, we first independently localize the circuit that each model recruits to solve numeric arithmetic problems (2+5) vs. verbal ones, in three languages: English ('two plus five'), Spanish ('dos m\'as cinco'), and Italian ('due pi\`u cinque'); then, we test whether overlap with the model's own numeric circuit predicts its generalization to the verbal formats. Indeed, we find support for this idea at three levels: circuit overlap accounts for the relative difficulty of the three verbal formats, for which models generalize best, and for which items are solved correctly, rivaling supervised probes while requiring no labeled data.

大模型可解释性算术推理神经电路

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。