arXiv:2602.18776cs.CLcs.AI2026-02

评测大模型读阿拉伯数字能力,发现准确率差距大且结构化输出难。

ArabicNumBench: Evaluating Arabic Number Reading in Large Language Models

  • 构建涵盖210个任务的阿拉伯数字阅读评测基准
  • 少样本思维链提示使准确率提升至80.06%,远超零样本
  • 高准确率模型常不生成结构化输出,需额外提取

我们提出ArabicNumBench,一个全面评估大语言模型在阿拉伯数字阅读任务上的基准,覆盖东阿拉伯-印度数字(0-9阿拉伯书写)和西阿拉伯数字(0-9)。使用四种提示策略(零样本、零样本思维链、少样本、少样本思维链)对10家厂商的71个模型进行评估,测试任务共210个,涵盖纯数字、地址、日期、数量和价格六类场景。评估包含59,010个独立测试用例,并追踪输出提取方法以衡量结构化输出生成能力。结果显示性能差异显著,准确率在14.29%至99.05%之间。少样本思维链提示的准确率(80.06%)是零样本(28.76%)的2.8倍。一个显著发现是:达到98%-99%高准确率的模型大多生成非结构化输出,多数响应缺乏阿拉伯语思维链标记。仅有6个模型在所有测试中始终生成结构化输出,多数即使数值准确也需依赖备用提取方法。对281个模型-策略组合的全面评估表明,数值准确性和指令遵循是两个独立能力,为阿拉伯语数字理解建立基线,并为生产级阿拉伯语NLP系统提供选型指导。

原文摘要 · Abstract (English)

We present ArabicNumBench, a comprehensive benchmark for evaluating large language models on Arabic number reading tasks across Eastern Arabic-Indic numerals (0-9 in Arabic script) and Western Arabic numerals (0-9). We evaluate 71 models from 10 providers using four prompting strategies (zero-shot, zero-shot CoT, few-shot, few-shot CoT) on 210 number reading tasks spanning six contextual categories: pure numerals, addresses, dates, quantities, and prices. Our evaluation comprises 59,010 individual test cases and tracks extraction methods to measure structured output generation. Evaluation reveals substantial performance variation, with accuracy ranging from 14.29\% to 99.05\% across models and strategies. Few-shot Chain-of-Thought prompting achieves 2.8x higher accuracy than zero-shot approaches (80.06\% vs 28.76\%). A striking finding emerges: models achieving elite accuracy (98-99\%) often produce predominantly unstructured output, with most responses lacking Arabic CoT markers. Only 6 models consistently generate structured output across all test cases, while the majority require fallback extraction methods despite high numerical accuracy. Comprehensive evaluation of 281 model-strategy combinations demonstrates that numerical accuracy and instruction-following represent distinct capabilities, establishing baselines for Arabic number comprehension and providing actionable guidance for model selection in production Arabic NLP systems.

大模型评测阿拉伯语数字识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。