arXiv:2506.08966cs.CLcs.LG2025-06EMNLP被引 3

发现预训练模型能精准表示数字,可借此修复算术错误

Pre-trained Language Models Learn Remarkably Accurate Representations of Numbers

  • 设计新探测方法,从嵌入中高精度还原数字
  • 实测多款开源模型均能准确表示数字
  • 揭示算术错误根源并提出修正方案

预训练语言模型常出现算术错误。现有研究对模型表示中的数值探测效果有限,认为这是分布学习嵌入难以精确表达数量所致。但本文发现,先前探测方法无法捕捉到数字嵌入中涌现的正弦模式结构。为此,我们提出一种新型探测技术,能在多种开源语言模型中近乎完美地解码出输入嵌入中的数值。这证明仅经过预训练,语言模型便能以极高的精度表示数字。最后,我们发现嵌入精度(由探测准确率衡量)可解释模型在基础算术中的大部分错误,并证实将嵌入对齐至探测发现的模式,能有效缓解这些错误。

原文摘要 · Abstract (English)

Pretrained language models (LMs) are prone to arithmetic errors. Existing work showed limited success in probing numeric values from models' representations, indicating that these errors can be attributed to the inherent unreliability of distributionally learned embeddings in representing exact quantities. However, we observe that previous probing methods are inadequate for the emergent structure of learned number embeddings with sinusoidal patterns. In response, we propose a novel probing technique that decodes numeric values from input embeddings with near-perfect accuracy across a range of open-source LMs. This proves that after the sole pre-training, LMs represent numbers with remarkable precision. Finally, we find that the embeddings' precision, judged by our probe's accuracy, explains a large portion of LM's errors in elementary arithmetic, and show that aligning the embeddings with the pattern our probes discover can mitigate these errors.

语言模型数字表示算术错误嵌入探测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。