arXiv:2507.12379cs.CLcs.AI2025-07EMNLP被引 17

用内部状态检测语言模型的算术错误,准确率超90%。

Probing for Arithmetic Errors in Language Models

  • 通过简单探测器从隐藏状态还原预测与正确答案
  • 在三位加法任务中模型正确性预测准确率超90%
  • 可定位错误推理步骤并引导重提示,提升任务准确率

我们研究语言模型内部激活是否可用于检测算术错误。以三位数加法为控制场景,发现简单探测器能准确从隐藏状态解码出模型输出和正确答案,无论模型输出是否正确。基于此,我们训练出轻量级错误检测器,对模型正确性的预测准确率超过90%。进一步分析仅含加法的GSM8K链式思维轨迹,发现基于简单算术训练的探测器在更复杂场景中仍具良好泛化能力,揭示出一致的内部表征。最后,我们证明这些探测器可引导对错误推理步骤的定向重提示,在不干扰正确输出的前提下提升任务准确率。结果表明,仅凭内部激活即可预判算术错误,简单探测器为轻量级模型自纠错提供了可行路径。

原文摘要 · Abstract (English)

We investigate whether internal activations in language models can be used to detect arithmetic errors. Starting with a controlled setting of 3-digit addition, we show that simple probes can accurately decode both the model's predicted output and the correct answer from hidden states, regardless of whether the model's output is correct. Building on this, we train lightweight error detectors that predict model correctness with over 90% accuracy. We then extend our analysis to structured chain-of-thought traces on addition-only GSM8K problems and find that probes trained on simple arithmetic generalize well to this more complex setting, revealing consistent internal representations. Finally, we demonstrate that these probes can guide selective re-prompting of erroneous reasoning steps, improving task accuracy with minimal disruption to correct outputs. Our findings suggest that arithmetic errors can be anticipated from internal activations alone, and that simple probes offer a viable path toward lightweight model self-correction.

模型纠错算术推理内部探测轻量级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。