测试30语言对话性能,发现开源模型全面落后于商业模型。
Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+

- 用自对弈多轮对话游戏评估30种语言的LLM表现。
- 开源模型在欧盟24种语言中平均分低于40,商业模型领先超30%。
- 非英语语言运行成本高31%,得分低10%,服务不平等。
我们评估大语言模型(LLMs)作为语言代理,在30种语言(欧盟24种官方语言及6种其他语言)中进行目标导向的自对弈对话游戏。该范式为多轮、无参考、程序化评分,且游戏机制与语言无关,新增语言仅需本地化固定提示和词表文件。评估九个开放权重与商用LLM发现:无一开源模型能覆盖欧盟24种语言;在每种官方语言中,商用系统均优于所有开源模型,两个最弱模型在欧盟24国平均分低于40。即使在公共网络文本少四个数量级的语言中,商用模型仍保持领先,表明语言对等可实现,但仅靠公开爬取数据不足。模型母语地区带来优势:中文在两款中国开发模型中最强,但最佳中文成绩仍由美国商用系统取得。覆盖率并非服务对等:跨模型与语言汇总,非英语语言平均运行成本高出31%,得分低10%。
原文摘要 · Abstract (English)
We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU languages plus six others. Unlike static or preference-based evaluation, this paradigm is multi-turn, reference-free and programmatically scored, and because the game mechanics are language-agnostic it extends to a new language by localising a fixed set of prompt and word-list files. Evaluating nine open-weight and commercial LLMs, we find that no open-weight model covers the EU-24 well: in every official language both commercial systems outscore every open-weight model, and the two weakest average below 40 points across the EU-24. The commercial systems stay ahead even in languages with four orders of magnitude less public web text, showing that linguistic parity is achievable, but not from public crawls alone. A model's home region lifts it without closing the gap: Chinese is the strongest of all 30 languages for two Chinese-developed models, yet the best Chinese score of any model belongs to a US commercial system. Coverage is also not parity of service. Pooled over models and languages, the median non-English language costs 31% more to run than English, and scores 10% lower.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。