arXiv:2505.14852cs.CLcs.AI2025-05

测试小模型在13类数学题上的零样本推理能力

EasyMath: A 0-shot Math Benchmark for SLMs

  • 构建13类数学任务的紧凑基准,覆盖基础运算到应用题
  • 40亿参数模型零样本准确率超75%,规模越大表现越好
  • 适合评估轻量级模型的数学推理能力,无需微调

EasyMath 是一个针对小型语言模型实用数学推理能力的紧凑型基准。涵盖十三个类别,从基础算术、运算顺序到应用题、代数表达式及边界情况,不包含专业领域内容。我们对23个模型(参数量1400万至40亿)在零样本设置下进行了测试,采用精确匹配、数值接近和符号一致三种方式评估自由回答。结果显示,模型性能随规模与训练程度提升,链式思考带来小幅增益,且一致性随规模增大而改善。

原文摘要 · Abstract (English)

EasyMath is a compact benchmark for practical math reasoning in small language models. It covers thirteen categories, from basic arithmetic and order of operations to word problems, algebraic expressions, edge cases, and omits specialist topics. We tested 23 models (14M to 4B parameters) using exact, numerical, and symbolic checks on free-form answers in a zero-shot setting. Accuracy rises with size and training, chain-of-thought adds modest gains, and consistency improves at scale.

数学推理小模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。