测试43种语言下大模型推理能力,发现机器翻译题也能用。
mmPISA-bench: Do LLMs Reason Equally Well Across 43 Languages?

- 基于PISA考试设计43语言推理题库,含25题共2150数据点。
- 大模型在各语言中表现接近人类水平,机器翻译不影响准确率。
- 部分语言推理更耗资源且更不准确,适合多语言评测研究者。
我们提出mmPISA-bench,一个源自OECD国际学生评估项目(PISA)的紧凑高质量多语言推理基准。该基准包含25道需推理才能正确回答的多项选择题,每题提供43种语言的官方人工翻译,并附带机器翻译版本(总计2150个数据点)。我们在多种语言、推理难度层级和翻译类型下评估了两种主流商用大模型的答题准确率。结果显示,现代大模型在所有测试语言中均能有效推理,准确率接近人类测试者水平,尽管存在轻微语言间差异。进一步发现,机器翻译题目并未降低准确率,表明高质量机器翻译(合成数据)在官方译本不可用时可能已足够用于大规模多语言推理评估。最后分析了令牌使用量与推理成本,发现某些语言下的大模型使用既更昂贵又更不准确。
原文摘要 · Abstract (English)
We introduce mmPISA-bench, a compact high-quality multilingual reasoning benchmark derived from the OECD Programme for International Student Assessment (PISA). The benchmark consists of 25 multiple-choice questions that require reasoning in order to be answered correctly. Each question is provided in official human translations to 43 languages and complemented with machine-translated counterparts (i.e., 2,150 data points in total). We evaluate two mainstream proprietary LLMs across languages, reasoning effort levels, and translation types in terms of their ability to answer the questions correctly. Our results show that modern LLMs can reason effectively across all evaluated languages, achieve accuracy comparable to human test-takers, with some performance variations across covered languages. We further find that machine-translated questions do not degrade accuracy relative to official human translations which suggests that high-quality machine translation (synthetic data) might often be adequate for large-scale multilingual reasoning evaluations where official translations are not available. Finally, we analyze token usage and related inference cost and find that LLMs usage in some languages is simultaneously more expensive and less accurate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。