构建47种语言的母语级评估框架,提升多语言大模型质量评测与对齐。
MENLO: From Preferences to Proficiency -- Evaluating and Modeling Native-like Quality Across 47 Languages
- 基于受众设计机制,建立可操作的母语级响应质量评估方法。
- 构建6423组标注数据,覆盖47种语言,四维度评估一致性高。
- 强化学习微调后,模型评分逼近人类,可用于提升多语言生成能力。
确保大语言模型在多种语言中生成母语级质量回应是一项挑战。为此,我们提出MENLO框架,基于受众设计启发的机制,实现母语级响应质量的可操作评估。利用MENLO,我们创建了一个包含6,423个由人类标注的提示-回应偏好对的数据集,涵盖四种质量维度,且在47种语言变体中具有高标注者间一致性。评估显示,零样本大模型评判器在成对评估和结构化标注指南下表现显著提升,但仍逊于人类标注者。通过强化学习、奖励塑造和多任务学习微调,我们实现了显著改进。此外,经强化学习训练的评判器可作为生成式奖励模型,提升大模型的多语言熟练度,尽管仍存在与人类判断的差异。研究结果为可扩展的多语言评估与偏好对齐指明了前景。我们已公开数据集与评估框架,以支持后续研究(https://huggingface.co/datasets/facebook/menlo)。
原文摘要 · Abstract (English)
Ensuring native-like quality of large language model (LLM) responses across many languages is challenging. To address this, we introduce MENLO, a framework that operationalizes the evaluation of native-like response quality based on audience design-inspired mechanisms. Using MENLO, we create a dataset of 6,423 human-annotated prompt-response preference pairs covering four quality dimensions with high inter-annotator agreement in 47 language varieties. Our evaluation reveals that zero-shot LLM judges benefit significantly from pairwise evaluation and our structured annotation rubrics, yet they still underperform human annotators on our dataset. We demonstrate substantial improvements through fine-tuning with reinforcement learning, reward shaping, and multi-task learning approaches. Additionally, we show that RL-trained judges can serve as generative reward models to enhance LLMs' multilingual proficiency, though discrepancies with human judgment remain. Our findings suggest promising directions for scalable multilingual evaluation and preference alignment. We release our dataset and evaluation framework to support further research in multilingual LLM evaluation (https://huggingface.co/datasets/facebook/menlo).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。