2000多个多语言基准揭示:翻译不如本地化,英语仍占主导。
The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks
- 分析2000+多语言基准,发现多数依赖英文原生数据
- 翻译基准与本地人类判断相关性仅0.47,本地化达0.68
- 建议构建文化语境适配的本土化评测,而非简单翻译
随着大语言模型在语言能力上的持续进步,稳健的多语言评估已成为推动技术公平发展的关键。本文分析了2021至2024年间来自148个国家的2000多个非英语多语言基准,评估多语言评测的过去、现状与未来实践。研究发现,尽管投入超千万美元,英语在这些基准中仍严重过载;多数基准使用原始语言内容而非翻译,且主要来自中国、印度、德国、英国和美国等高资源国家。对比基准表现与人类判断,发现科技类任务相关性较高(0.70至0.85),而传统NLP任务如问答(如XQuAD)相关性极低(0.11至0.30)。将英文基准翻译为其他语言效果有限,本地化基准与本地人类判断的相关性(0.68)显著高于翻译版本(0.47)。这表明应重视文化与语言定制化的评测建设。本文总结当前多语言评估的六大局限,提出指导原则,并提出五个关键研究方向,呼吁全球协作建立以人类对齐、面向真实应用的多语言评测体系。
原文摘要 · Abstract (English)
As large language models (LLMs) continue to advance in linguistic capabilities, robust multilingual evaluation has become essential for promoting equitable technological progress. This position paper examines over 2,000 multilingual (non-English) benchmarks from 148 countries, published between 2021 and 2024, to evaluate past, present, and future practices in multilingual benchmarking. Our findings reveal that, despite significant investments amounting to tens of millions of dollars, English remains significantly overrepresented in these benchmarks. Additionally, most benchmarks rely on original language content rather than translations, with the majority sourced from high-resource countries such as China, India, Germany, the UK, and the USA. Furthermore, a comparison of benchmark performance with human judgments highlights notable disparities. STEM-related tasks exhibit strong correlations with human evaluations (0.70 to 0.85), while traditional NLP tasks like question answering (e.g., XQuAD) show much weaker correlations (0.11 to 0.30). Moreover, translating English benchmarks into other languages proves insufficient, as localized benchmarks demonstrate significantly higher alignment with local human judgments (0.68) than their translated counterparts (0.47). This underscores the importance of creating culturally and linguistically tailored benchmarks rather than relying solely on translations. Through this comprehensive analysis, we highlight six key limitations in current multilingual evaluation practices, propose the guiding principles accordingly for effective multilingual benchmarking, and outline five critical research directions to drive progress in the field. Finally, we call for a global collaborative effort to develop human-aligned benchmarks that prioritize real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。