arXiv:2507.17476cs.CLcs.AI2025-07被引 13

构建多语言原生推理基准,揭示大模型跨文化推理短板

MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs

  • 用法语/西班牙语/中文母语者原创题目,覆盖四类推理任务
  • 14个主流模型在多语言上平均得分不足50%,数学题差10%以上
  • 首次提供跨语言同题对比,适合评估模型跨文化能力

尽管近期大语言模型在英语推理基准上进展迅速,但对其跨多种语言与文化背景的多语言推理能力的评估仍有限。现有多语言推理基准通常通过翻译英语基准构建,偏向于英语语境下的推理问题。本文提出多语言原生推理挑战(MultiNRC),包含超过1,000道由法语、西班牙语和中文母语者撰写的原生语言与文化相关推理题。该基准涵盖四大核心推理类别:语言特异性语言推理、文字游戏与谜题、文化/传统推理,以及具有文化关联性的数学推理。对于后两类,我们还由精通英语的母语者人工翻译了英文对应题,实现跨语言同题直接比较。系统评估了当前14个主流大模型在MultiNRC及其英文等价集上的表现。结果表明:(1)当前模型在原生多语言推理上仍表现不佳,无一模型得分超过50%;(2)模型在语言、文化与逻辑推理任务中呈现不同优劣势;(3)多数模型在英语数学题上表现优于原语言版本(+10%),凸显其对文化语境知识的掌握仍存挑战。

原文摘要 · Abstract (English)

Although recent Large Language Models (LLMs) have shown rapid improvement on reasoning benchmarks in English, the evaluation of such LLMs' multilingual reasoning capability across diverse languages and cultural contexts remains limited. Existing multilingual reasoning benchmarks are typically constructed by translating existing English reasoning benchmarks, biasing these benchmarks towards reasoning problems with context in English language/cultures. In this work, we introduce the Multilingual Native Reasoning Challenge (MultiNRC), a benchmark designed to assess LLMs on more than 1,000 native, linguistic and culturally grounded reasoning questions written by native speakers in French, Spanish, and Chinese. MultiNRC covers four core reasoning categories: language-specific linguistic reasoning, wordplay & riddles, cultural/tradition reasoning, and math reasoning with cultural relevance. For cultural/tradition reasoning and math reasoning with cultural relevance, we also provide English equivalent translations of the multilingual questions by manual translation from native speakers fluent in English. This set of English equivalents can provide a direct comparison of LLM reasoning capacity in other languages vs. English on the same reasoning questions. We systematically evaluate current 14 leading LLMs covering most LLM families on MultiNRC and its English equivalent set. The results show that (1) current LLMs are still not good at native multilingual reasoning, with none scoring above 50% on MultiNRC; (2) LLMs exhibit distinct strengths and weaknesses in handling linguistic, cultural, and logical reasoning tasks; (3) Most models perform substantially better in math reasoning in English compared to in original languages (+10%), indicating persistent challenges with culturally grounded knowledge.

多语言推理大模型评测文化认知语言差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。