arXiv:2603.10767cs.CL2026-03被引 2

构建14种语言的高质量数学题数据集,支持多语言强化学习训练。

mAceReason-Math: A Dataset of High-Quality Multilingual Math Problems Ready For RLVR

  • 基于专为强化学习设计的语料库,翻译并清洗高难度数学题。
  • 每种语言超1万道题,覆盖14种语言,难度适配当前大模型能力。
  • 专为多语言强化学习验证奖励(RLVR)设计,助力跨语言推理研究。

强化学习结合可验证奖励(RLVR)已显著提升预训练大模型在数学与逻辑问题上的能力,但现有研究与数据集仍以英语为主。尽管已有多种语言的训练数据和评测基准,但它们未针对RLVR及当前模型能力进行设计,且难度普遍过低,难以提供有效训练信号。为此,我们推出了mAceReason-Math,一个源自专为RLVR优化的语料库(AceReason-Math)的高质量多语言数学题数据集。我们对翻译进行了严格清洗与改进,覆盖14种语言,每种语言包含超过10,000个样本。该数据集将开放发布,以推动多语言RLVR研究与基准测试的发展。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) has been successfully applied to significantly boost the capabilities of pretrained large language models, especially in the math and logic problem domains. However, current research and available training datasets remain English-centric. While multilingual training data and benchmarks have been created in the past, they were not created with RLVR and current model capability in mind, and their level of difficulty is often too low to provide appropriate training signals for current models. To address this gap, we provide mAceReason-Math, a dataset of high-quality translations of challenging math problems sourced from a corpus specifically curated for RLVR (AceReason-Math). We further take specific care to clean and improve our translations, resulting in a coverage of 14 languages with more than 10,000 samples per language. We release the dataset to facilitate multilingual RLVR research and benchmarking in the research community.

多语言数学推理强化学习数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。