arXiv:2601.15251cs.CL2026-01被引 1

不同数字写法让大模型算术能力下降,提示词可有效弥补

The Effect of Scripts and Formats on LLM Numeracy

  • 测试多种数字书写形式,发现模型准确率显著下降
  • 使用少量示例提示可使准确率提升近50%以上
  • 适合多语言数字处理、跨格式数据生成的场景

大型语言模型在基础算术任务上已达到接近人类水平的表现。然而,当数值表达方式偏离其训练语料库中的主流惯例时,模型表现却鲜受关注。本文研究了在多种数字符号与格式下的数值推理能力,发现即使数学逻辑相同,使用非主流书写形式时模型准确率会大幅下降。进一步实验表明,通过少样本提示和显式数值映射等策略,可显著缩小这一差距。研究揭示了多语言数值推理中的潜在挑战,并为可靠地跨不同数字符号与格式进行数字理解、操作与生成提供了可操作的建议。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved impressive proficiency in basic arithmetic, rivaling human-level performance on standard numerical tasks. However, little attention has been given to how these models perform when numerical expressions deviate from the prevailing conventions present in their training corpora. In this work, we investigate numerical reasoning across a wide range of numeral scripts and formats. We show that LLM accuracy drops substantially when numerical inputs are rendered in underrepresented scripts or formats, despite the underlying mathematical reasoning being identical. We further demonstrate that targeted prompting strategies, such as few-shot prompting and explicit numeral mapping, can greatly narrow this gap. Our findings highlight an overlooked challenge in multilingual numerical reasoning and provide actionable insights for working with LLMs to reliably interpret, manipulate, and generate numbers across diverse numeral scripts and formatting styles.

算术推理多语言提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。