扩充WMT24至55种语言,验证大模型在多语种翻译中表现最佳
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects
- 新增46种语言/方言的人工参考译文与校对稿,覆盖4个领域
- 在全部55种语言上,大模型翻译性能超越传统MT系统
- 适合关注多语言评估与大模型跨语言能力的研究者
随着大语言模型在非英语语言上的能力提升,构建基准数据集以评估其多语言性能(如机器翻译)变得愈发重要。本文将WMT24数据集扩展至涵盖55种语言,新增46种语言及方言的人工撰写参考译文和后编辑内容,并对原WMT24中9种语言中的8种进行了参考译文的后编辑。数据集覆盖文学、新闻、社交和语音四个领域。我们使用自动指标对多种MT服务和大模型进行评测,结果显示大模型在所有55种语言上均表现最佳。该结果需通过人工评估进一步验证,留待未来工作完成。
原文摘要 · Abstract (English)
As large language models (LLM) become more and more capable in languages other than English, it is important to collect benchmark datasets in order to evaluate their multilingual performance, including on tasks like machine translation (MT). In this work, we extend the WMT24 dataset to cover 55 languages by collecting new human-written references and post-edits for 46 new languages and dialects in addition to post-edits of the references in 8 out of 9 languages in the original WMT24 dataset. The dataset covers four domains: literary, news, social, and speech. We benchmark a variety of MT providers and LLMs on the collected dataset using automatic metrics and find that LLMs are the best-performing MT systems in all 55 languages. These results should be confirmed using a human-based evaluation, which we leave for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。