arXiv:2604.20225cs.CL2026-04ACL被引 1

构建多语言多文化评测基准,揭示大模型在全球化应用中的能力短板。

The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models

论文配图:The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models
图 1 · 摘自论文原文
  • 按文化层与认知层设计三类九层评测框架,覆盖26语言51文化
  • 通过专家本地化实现19语言原生数据,跨文化测试集覆盖34文化
  • 诊断20+主流模型表现,发现地域差异与任务间显著差距

评估大语言模型的多语言与多文化能力对其实现全球应用至关重要。现有评测存在三大缺陷:(1)评价维度分散,常忽视深层文化细节;(2)主观任务语言覆盖不足,依赖低质量机器翻译;(3)分析浅显,缺乏诊断深度。为此,我们提出GaoYao基准,包含182.3万样本、26种语言和51个国家/地区。首先,GaoYao构建统一框架,将评测任务分为三类文化层级(通用多语言、跨文化、单一文化)与九个认知子层级。其次,通过专家严格本地化,将主观评测扩展至19种语言,并合成34个文化的跨文化测试集,覆盖范围超越以往达111%。第三,对20余个主流及小型大模型进行深度诊断分析。结果揭示显著的地理性能差异与任务间断层,为未来研究提供可靠路线图。基准已开源(https://github.com/lunyiliu/GaoYao)。

原文摘要 · Abstract (English)

Evaluating the multilingual and multicultural capabilities of Large Language Models (LLMs) is essential for their global utility. However, current benchmarks face three critical limitations: (1) fragmented evaluation dimensions that often neglect deep cultural nuances; (2) insufficient language coverage in subjective tasks relying on low-quality machine translation; and (3) shallow analysis that lacks diagnostic depth beyond simple rankings. To address these, we introduce GaoYao, a comprehensive benchmark with 182.3k samples, 26 languages and 51 nations/areas. First, GaoYao proposes a unified framework categorizing evaluation tasks into three cultural layers (General Multilingual, Cross-cultural, Monocultural) and nine cognitive sub-layers. Second, we achieve native-quality expansion by leveraging experts to rigorously localize subjective benchmarks into 19 languages and synthesizing cross-cultural test sets for 34 cultures, surpassing prior coverage by up to 111%. Third, we conduct an in-depth diagnostic analysis on 20+ flagship and compact LLMs. Our findings reveal significant geographical performance disparities and distinct gaps between tasks, offering a reliable map for future work. We release the benchmark (https://github.com/lunyiliu/GaoYao).

多语言评测文化理解LLM评估跨文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。