arXiv:2409.11074cs.CLcs.AI2024-09被引 5

构建罗马尼亚语数学推理数据集,推动低资源语言AI发展

RoMath: A Mathematical Reasoning Benchmark in Romanian

  • 针对罗马尼亚语设计三类数学题集,覆盖高考、竞赛与合成数据
  • 现有开源模型在罗马尼亚语数学题上表现较差,平均准确率不足40%
  • 填补非英语数学推理资源空白,适合多语言AI研究者使用

数学长期以来以自然语言形式呈现,主要用于人类理解。随着机械化数学和证明助手的发展,理解非正式数学文本的需求日益增长,但现有基准大多仅关注英语,忽视其他语言。本文提出RoMath,一个面向罗马尼亚语的数学推理基准套件,包含三个子集:Baccalaureate(高考)、Competitions(竞赛)和Synthetic(合成数据),涵盖多种数学领域和难度级别,旨在提升非英语语言模型能力,促进多语言AI发展。通过聚焦罗马尼亚语这一低资源语言及其独特的语言特征,RoMath弥补了以英语为中心的模型局限性,强调了为非主流语言构建专用资源的重要性,而非依赖简单翻译。我们对多个开源权重模型进行了基准测试,凸显为欠代表语言建立资源的紧迫性。代码与数据集将公开。

原文摘要 · Abstract (English)

Mathematics has long been conveyed through natural language, primarily for human understanding. With the rise of mechanized mathematics and proof assistants, there is a growing need to understand informal mathematical text, yet most existing benchmarks focus solely on English, overlooking other languages. This paper introduces RoMath, a Romanian mathematical reasoning benchmark suite comprising three subsets: Baccalaureate, Competitions and Synthetic, which cover a range of mathematical domains and difficulty levels, aiming to improve non-English language models and promote multilingual AI development. By focusing on Romanian, a low-resource language with unique linguistic features, RoMath addresses the limitations of Anglo-centric models and emphasizes the need for dedicated resources beyond simple automatic translation. We benchmark several open-weight language models, highlighting the importance of creating resources for underrepresented languages. Code and datasets are be made available.

数学推理多语言AI低资源语言基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。