arXiv:2510.08009cs.AIcs.LG2025-10被引 5

语言模型不连续地表示数字,高精度时噪声显著增加。

Language Models Do Not Embed Numbers Continuously

  • 用线性重建和主成分分析验证数字嵌入的连续性
  • 高精度数字下解释方差下降,$R^2 \> 0.95$但主成分贡献小
  • 适用于需高精度数值处理的场景,如金融、科学计算

近期研究关注大语言模型在特定算术任务中对整数的处理方式,以及其对数值的表征机制。已有工作表明模型嵌入可重构原始值,但未检验其是否将连续数值作为连续空间建模。本文基于嵌入空间的期望性质(如线性重建与主成分分析)发现,语言模型不仅以非连续方式表征数值,还引入显著噪声。使用OpenAI、Google Gemini与Voyage AI三类主流模型验证,尽管重建精度高($R^2 \geq 0.95$),但主成分仅解释嵌入空间中少量变异。说明多数嵌入维度与简单数值输入空间正交。且随着小数精度提升,线性重建与解释方差均下降,尽管输入的序数关系保持不变。该结果对需要高数值精度、大数值范围或混合符号值的应用具有重要影响。

原文摘要 · Abstract (English)

Recent research has extensively studied how large language models manipulate integers in specific arithmetic tasks, and on a more fundamental level, how they represent numeric values. These previous works have found that language model embeddings can be used to reconstruct the original values, however, they do not evaluate whether language models actually model continuous values as continuous. Using expected properties of the embedding space, including linear reconstruction and principal component analysis, we show that language models not only represent numeric spaces as non-continuous but also introduce significant noise. Using models from three major providers (OpenAI, Google Gemini and Voyage AI), we show that while reconstruction is possible with high fidelity ($R^2 \geq 0.95$), principal components only explain a minor share of variation within the embedding space. This indicates that many components within the embedding space are orthogonal to the simple numeric input space. Further, both linear reconstruction and explained variance suffer with increasing decimal precision, despite the ordinal nature of the input space being fundamentally unchanged. The findings of this work therefore have implications for the many areas where embedding models are used, in-particular where high numerical precision, large magnitudes or mixed-sign values are common.

语言模型数值表征嵌入空间精度问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。