开源多语言大模型评估工具,提升非英语文本评价能力
M-Prometheus: A Suite of Open Multilingual LLM Judges
- 构建3B至14B参数的多语言评估模型,支持直接评分与成对比较
- 在20+语言奖励基准上超越现有开源模型,4个语对翻译评价表现优异
- 使用合成多语言数据训练更有效,适合多语言模型研发者使用
利用大语言模型自动评估长文本(LLM-as-a-judge)日益普遍,但现有模型大多仅针对英语优化,缺乏对多语言评估能力的系统研究,导致非英语语言的自动评估质量落后,制约多语言模型发展。为此,我们提出M-Prometheus,一套涵盖3B至14B参数的开源多语言大模型评估工具,可对多语言输出进行直接评分和成对比较。该模型在超过20种语言的多语言奖励基准上优于当前最优开源模型,在4个语对的文学机器翻译评估中表现突出。此外,其在解码阶段能显著提升三种语言生成内容质量,验证了其对多语言模型开发的实际价值。通过大量消融实验,我们发现骨干模型选择及使用合成多语言反馈数据而非翻译数据是关键因素。相关模型、训练数据与代码均已开源。
原文摘要 · Abstract (English)
The use of language models for automatically evaluating long-form text (LLM-as-a-judge) is becoming increasingly common, yet most LLM judges are optimized exclusively for English, with strategies for enhancing their multilingual evaluation capabilities remaining largely unexplored in the current literature. This has created a disparity in the quality of automatic evaluation methods for non-English languages, ultimately hindering the development of models with better multilingual capabilities. To bridge this gap, we introduce M-Prometheus, a suite of open-weight LLM judges ranging from 3B to 14B parameters that can provide both direct assessment and pairwise comparison feedback on multilingual outputs. M-Prometheus models outperform state-of-the-art open LLM judges on multilingual reward benchmarks spanning more than 20 languages, as well as on literary machine translation (MT) evaluation covering 4 language pairs. Furthermore, M-Prometheus models can be leveraged at decoding time to significantly improve generated outputs across all 3 tested languages, showcasing their utility for the development of better multilingual models. Lastly, through extensive ablations, we identify the key factors for obtaining an effective multilingual judge, including backbone model selection and training on synthetic multilingual feedback data instead of translated data. We release our models, training dataset, and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。