中国大模型语言能力与西方模型相似,仅中文表现更好。
Do Chinese models speak Chinese languages?
- 对比21种语言变体,分析中西开源大模型多语言能力差异。
- 中文模型在普通话上优于西方模型,但对少数族裔语言支持不足。
- 全球基准测试和共享训练资源导致模型语言能力趋同。
顶尖开源大模型的发布确立了中国在人工智能领域的领先地位。这些模型是否支持中国境内的语言?还是与欧美开发的模型一样?比较多语言能力对两个方面至关重要:一是揭示预训练数据的构建方式,反映资源分配与研发重点;二是帮助中国模型开发者平衡服务国内多元语言人口与优化以英语为主的全球基准。我们通过对比中国与西方开发的开源大模型,在21种语言变体(包括亚洲区域、汉语及欧洲语言)上的表现,研究其语言优先级。信息公平性与阅读理解实验显示,中国模型在这些语言上的性能与西方模型高度相关(r=0.93),唯一例外是普通话表现更优。中国模型在法语、德语上表现良好,但在识别维吾尔语、哈萨克语等少数民族语言时存在困难。所有研究模型表现出相似的多语言性能分布,尽管其开发背景语言文化各异。这种同质化现象可归因于全球基准测试实践和共享训练资源的影响。我们的结果表明,当前语言支持并非必然,而是权衡取舍的结果,对开发者、政策制定者和用户具有重要意义。
原文摘要 · Abstract (English)
The release of top-performing open-weight LLMs has cemented China's role as a leading force in AI development. Do these models support languages spoken in China? Or do they support the same languages as models developed in the United States or in Europe? Comparing multilingual capabilities is important for two reasons. First, language ability provides insights into pre-training data curation, and thus into resource allocation and development priorities. Second, Chinese model developers need to navigate the tension between serving a linguistically diverse population domestically, and optimizing for globally visible benchmarks that are predominantly English. We investigate Chinese model developers' priorities through a comparative study of Chinese-developed and Western-developed open-weight LLMs, on 21 language variants including Asian regional, Chinese, and European languages. Our experiments on Information Parity and reading comprehension show Chinese models' performance across these languages correlates strongly (r=0.93) with their Western counterparts, with the sole exception being better Mandarin. Chinese-developed models are good at French and German, but they sometimes cannot identify languages spoken by Chinese minorities such as Kazakh and Uyghur. Overall, all open-weight LLMs we study have a similar multilingual performance profile, despite the diverse linguistic and cultural contexts the model developers operated within. We interpret the homogenization as consistent with the influence of global benchmarking practices and shared training resources. Rather than treating current language support as inevitable, our results highlight multilingual development as a space of prioritization and trade-offs, with implications for model developers, policymakers, and users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。