arXiv:2604.19262cs.CLcs.AI2026-04被引 1

评测大模型在真实场景中的多语言多文化推理能力

CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks

论文配图:CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks
图 1 · 摘自论文原文
  • 通过人机协作构建多语言多文化真实任务评测集
  • 14种语言覆盖51个地区,16类主题共2610个难题
  • 顶尖模型准确率仅44.48%,凸显能力短板

大型语言模型(LLMs)已在全球部署,催生了大量评估其多语言与多文化能力的基准。然而,现有基准多聚焦通用语言理解或浅层文化常识,对需在真实、上下文丰富的场景中进行推理的“具身任务”评估严重不足。为此,我们提出CulturALL——一个全面且具有挑战性的基准,用于评估大模型在具身任务中的多语言与多文化能力。CulturALL采用人机协同框架构建:专家标注确保难度与事实准确性,大模型辅助降低人工负担。通过融合多样化来源,实现场景全覆盖。每个题目精心设计,具备高难度。CulturALL包含2,610个样本,涵盖14种语言、51个地区,分布在16个主题中,全面覆盖具身任务。实验表明,表现最佳的模型在CulturALL上准确率为44.48%,凸显仍有巨大提升空间。

原文摘要 · Abstract (English)

Large language models (LLMs) are now deployed worldwide, inspiring a surge of benchmarks that measure their multilingual and multicultural abilities. However, these benchmarks prioritize generic language understanding or superficial cultural trivia, leaving the evaluation of grounded tasks -- where models must reason within real-world, context-rich scenarios -- largely unaddressed. To fill this gap, we present CulturALL, a comprehensive and challenging benchmark to assess LLMs' multilingual and multicultural competence on grounded tasks. CulturALL is built via a human--AI collaborative framework: expert annotators ensure appropriate difficulty and factual accuracy, while LLMs lighten the manual workload. By incorporating diverse sources, CulturALL ensures comprehensive scenario coverage. Each item is carefully designed to present a high level of difficulty, making CulturALL challenging. CulturALL contains 2,610 samples in 14 languages from 51 regions, distributed across 16 topics to capture the full breadth of grounded tasks. Experiments show that the best LLM achieves 44.48% accuracy on CulturALL, underscoring substantial room for improvement.

多语言多文化评测基准具身推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。