arXiv:2602.10732cs.CL2026-02ACL

构建多语言文化推理基准,分离思维类型与文化背景

Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling

  • 用100个语言无关模板,分拆推理类型与文化要素
  • 覆盖20国文化、20种语言,零样本测试中模型平均准确率80.8%
  • 文化相关数学题最难,本地语言下开源模型性能显著下降

现有多语言评测常忽略文化语境下的推理能力:翻译数据集延续英语中心范式,而文化主导数据集又难以控制推理类型。我们提出Macaron,一种基于模板的基准,将推理类型与文化维度解耦。采用100个语言无关模板,覆盖7类推理和22个文化维度,由母语标注者生成符合场景的英文与本地语言多项选择题,并系统推导真假判断题。数据集包含11,862个实例,覆盖20个国家/文化背景、10种书写系统及20种语言与方言(含阿姆哈拉语、约鲁巴语、祖鲁语、吉尔吉斯语及部分阿拉伯方言等低资源语言)。在对21个多语言大模型的零样本评估中,推理模式模型表现最佳(总体准确率80.8%),英语与本地语言间差距极小;而开放权重模型在本地语言中性能大幅下降,真/假任务常接近随机水平。文化相关的数学与计数类模板始终最困难。数据可在此获取:https://huggingface.co/datasets/AlaaAhmed2444/Macaron。

原文摘要 · Abstract (English)

Multilingual benchmarks rarely test reasoning over culturally grounded premises: translated datasets keep English-centric scenarios, while culture-first datasets often lack control over the reasoning required. We propose Macaron, a template-first benchmark that factorizes reasoning type and cultural aspect across question languages. Using 100 language-agnostic templates that cover 7 reasoning types, 22 cultural aspects, native annotators create scenario-aligned English and local-language multiple-choice questions, and systematically derived True/False questions. Macaron contains 11,862 instances spanning 20 countries/cultural contexts, 10 scripts, and 20 languages and dialects (including low-resource ones like Amharic, Yoruba, Zulu, Kyrgyz, and some Arabic dialects). In zero-shot evaluation of 21 multilingual LLMs, reasoning-mode models achieve the strongest performance (80.8% overall) and near-parity between English and local languages, while open-weight models degrade substantially in local languages and often approach chance on T/F tasks. Culture-grounded mathematical and counting templates are consistently the hardest. The data can be accessed here https://huggingface.co/datasets/AlaaAhmed2444/Macaron.

多语言文化推理评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。