对比不同方法在多语言文本评价中的表现,为低资源语言提供实用指南。
Towards Reliable Multilingual LLMs-as-a-Judge: An Empirical Study

- 对比指令翻译、单语与多语监督等策略,评估模型表现
- 小模型微调后性能接近商用大模型,大模型零样本更适配跨域场景
- 低资源语言如巴斯克语仍可有效评估,适合多语言评测研究者
大型语言模型(LLMs)在自动文本评价中应用日益广泛,但多数研究集中于英语。随着对多语言评价需求上升,将基于LLM的评价器扩展至多语言环境仍具挑战,尤其在低资源语言和缺乏领域数据时。本文系统研究了在有无领域数据条件下,构建多语言LLM评价器的策略,涵盖英语、西班牙语和巴斯克语(高、中、低资源语言),考察指令翻译、单语与多语监督、模型规模等因素。我们扩展了两个现有元评价数据集至巴斯克语和西班牙语。结果表明:有领域数据时,微调的小模型性能可媲美商用模型;无领域数据时,大模型零样本评估更优。此外,使用非领域数据微调反而会降低性能。这些发现为构建高效可靠的多语言评估流程提供了实践指导。数据与代码已公开于 hitz-zentroa/mJudge。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used for the automatic evaluation of generated text, yet most prior work focuses on English. Despite the growing demand for multilingual evaluation, extending LLM-based evaluators to multilingual settings remains challenging, particularly for low-resource languages and scenarios where in-domain data is scarce. This work explores several strategies for developing multilingual LLMs-as-a-judge, considering whether in-domain data is available for fine-tuning or not. We systematically analyze English, Spanish, and Basque, representing high-, mid-, and low-resource languages, considering instruction translation, monolingual versus multilingual supervision, and model size. For evaluation, we extend two existing meta-evaluation datasets to Basque and Spanish. Our results reveal key trade-offs: When in-domain data is available, fine-tuned smaller models can achieve performance comparable to proprietary models, whereas zero-shot evaluation with larger models proves more effective in out-of-domain settings. We also observe that fine-tuning on out-of-domain data can adversely affect model performance. These findings provide practical guidance for building efficient, reliable multilingual evaluation pipelines. The data and code are publicly available at hitz-zentroa/mJudge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。