arXiv:2506.19468cs.CLcs.AI2025-06ACL被引 13

评测61种语言的LLM多语言能力,发现实际表现与宣称差距大。

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages

  • 构建覆盖61语言的跨语言对齐评测集MuBench
  • 发现英语与低资源语言间性能差距显著
  • 提出多语言一致性指标,辅助模型优化

多语言大语言模型发展迅速,新模型不断宣称支持更多语言。然而现有评估数据集有限且缺乏跨语言对齐,导致多语言能力评估在语言和技能覆盖上碎片化。为此,我们提出MuBench,一个涵盖61种语言并评估多种能力的基准。我们评估了几种前沿多语言LLM,发现宣称与实际语言覆盖存在明显差距,尤其在英语与低资源语言之间存在持续的性能差异。利用MuBench的对齐特性,我们提出多语言一致性(MLC)作为准确率的补充指标,用于分析性能瓶颈并指导模型改进。最后,我们基于500B token的英文和中文数据,预训练了一系列1.2B参数模型,通过调整语言比例和并行数据比例,研究跨语言迁移动态。

原文摘要 · Abstract (English)

Multilingual large language models (LLMs) are advancing rapidly, with new models frequently claiming support for an increasing number of languages. However, existing evaluation datasets are limited and lack cross-lingual alignment, leaving assessments of multilingual capabilities fragmented in both language and skill coverage. To address this, we introduce MuBench, a benchmark covering 61 languages and evaluating a broad range of capabilities. We evaluate several state-of-the-art multilingual LLMs and find notable gaps between claimed and actual language coverage, particularly a persistent performance disparity between English and low-resource languages. Leveraging MuBench's alignment, we propose Multilingual Consistency (MLC) as a complementary metric to accuracy for analyzing performance bottlenecks and guiding model improvement. Finally, we pretrain a suite of 1.2B-parameter models on English and Chinese with 500B tokens, varying language ratios and parallel data proportions to investigate cross-lingual transfer dynamics.

多语言评测基准LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。