构建首个葡语数学推理数据集,填补非英语数学评测空白
MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese
- 从葡语母语源收集1729道数学题,涵盖奥数与考试
- 顶尖模型在选择题上表现好,但图文和开放题准确率下降
- 适合研究多语言数学推理或葡萄牙语AI的学者使用
大型语言模型在复杂数学推理领域的研究进展迅速,但现有评测数据集普遍存在语言偏倚,绝大多数仅限于英语或英文翻译。本文提出 { csc Math-PT},一个包含1729道用欧洲和巴西葡语编写的数学问题的新数据集,内容源自葡萄牙和巴西的数学奥赛、竞赛及考试等高质量本地资源。我们对当前最先进的大模型在 { sc Math-PT} 上进行了全面评估,发现前沿推理模型在选择题上表现优异,优于开源模型,但在涉及图表或开放回答的问题上性能显著下降。为促进后续研究,我们公开发布该数据集及模型输出结果。
原文摘要 · Abstract (English)
The use of large language models (LLMs) for complex mathematical reasoning is an emergent area of research, with fast progress in methods, models, and benchmark datasets. However, most mathematical reasoning evaluations exhibit a significant linguistic bias, with the vast majority of benchmark datasets being exclusively in English or (at best) translated from English. We address this limitation by introducing {\sc Math-PT}, a novel dataset comprising 1,729 mathematical problems written in European and Brazilian Portuguese. {\sc Math-PT} is curated from a variety of high-quality native sources, including mathematical Olympiads, competitions, and exams from Portugal and Brazil. We present a comprehensive benchmark of current state-of-the-art LLMs on {\sc Math-PT}, revealing that frontier reasoning models achieve strong performance in multiple choice questions compared to open weight models, but that their performance decreases for questions with figures or open-ended questions. To facilitate future research, we release the benchmark dataset and model outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。