arXiv:2601.16218cs.CLcs.AI2026-01被引 2

首个多语言多模态数学推理数据集,评测视觉语言模型在跨语言数学题上的表现。

M3Kang: Evaluating Multilingual Multimodal Mathematical Reasoning in Vision-Language Models

  • 基于全球最大数学竞赛构建,覆盖108种语言和1747道带图数学题。
  • 模型表现随语言数量和模型规模提升,但难随题目难度增长,且弱于人类学生。
  • 首次证明多语言技术可有效扩展至多模态场景,适合评估多语言AI能力者参考。

尽管当前先进的视觉语言模型(VLMs)展现出强大的推理能力,其在多语言数学推理方面的表现仍缺乏深入研究,尤其与人类表现相比差距明显。为此,我们提出M3Kang,这是首个大规模多语言、多模态数学推理数据集,源自全球最大的数学竞赛——Kangaroo Math Competition,该竞赛每年吸引超过六百万18岁以下青少年参与,覆盖90多个国家。M3Kang包含1,747道独特的选择题,按年级难度分级,涵盖108种文化多样化的语言,部分题目包含解题必需的图表。我们利用该数据集对闭源与开源的SOTA模型进行了广泛基准测试。结果显示,尽管近期有进展,模型在基础数学和图表推理上仍表现不佳,性能随语言数量和模型规模提升,但不随题目难度层级变化。同时发现,多语言技术可有效拓展至多模态场景,显著优于基线方法。我们的分析还整合了逾6.8万名学生的成绩数据,实现与人类表现的直接对比。我们已开源M3Kang,包括仅含英语的子集M2Kang,以及构建数据集的框架与代码库。

原文摘要 · Abstract (English)

Despite state-of-the-art vision-language models (VLMs) have demonstrated strong reasoning capabilities, their performance in multilingual mathematical reasoning remains underexplored, particularly when compared to human performance. To bridge this gap, we introduce M3Kang, the first massively multilingual, multimodal mathematical reasoning dataset for VLMs. It is derived from the Kangaroo Math Competition, the world's largest mathematics contest, which annually engages over six million participants under the age of 18 across more than 90 countries. M3Kang includes 1,747 unique multiple-choice problems organized by grade-level difficulty, with translations into 108 culturally diverse languages, some of them including diagrams essential for solving them. Using this dataset, we conduct extensive benchmarking on both closed- and open-source SOTA models. We observe that, despite recent advances, models still struggle with basic math and diagram-based reasoning, with performance scaling with language presence and model size, but not with grade level. We also find that multilingual techniques can be effectively extended to the multimodal setting, resulting in significant improvements over baseline approaches. Our analysis also incorporates performance data from over 68,000 students, enabling direct comparison with human performance. We are open-sourcing M3Kang, including the English-only subset M2Kang, along with the framework and codebase used to construct the dataset.

多模态数学推理多语言评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。