多语言大模型评卷可靠性差,平均一致性仅0.3,低资源语言更差。
How Reliable is Multilingual LLM-as-a-Judge?
- 用五种模型在25种语言上评估五项任务,分析多语言评卷一致性。
- 平均Fleiss' Kappa仅约0.3,部分模型表现更差,低资源语言尤其不稳。
- 模型规模和多语言训练不能提升一致性,提出集成策略改善实际应用。
LLM-as-a-Judge已成为流行的评估策略,利用先进大语言模型根据人类指令评估生成结果。尽管该方法可替代人工标注,其在多语言评估中的可靠性仍不确定。为此,我们对多语言LLM-as-a-Judge进行了全面分析,评估了来自不同模型家族的五种模型,在涉及25种语言的五项多样化任务上的表现。结果显示,这些模型在跨语言判断中难以保持一致,平均Fleiss' Kappa约为0.3,部分模型表现更差。进一步分析发现,一致性在不同语言间差异显著,低资源语言表现尤其不佳。此外,无论是多语言训练还是模型规模扩大,并不能直接提升判断一致性。研究提示当前大模型尚不可靠用于多语言生成评估。最后,我们提出一种集成策略,有效提升多语言评判在实际应用中的稳定性。
原文摘要 · Abstract (English)
LLM-as-a-Judge has emerged as a popular evaluation strategy, where advanced large language models assess generation results in alignment with human instructions. While these models serve as a promising alternative to human annotators, their reliability in multilingual evaluation remains uncertain. To bridge this gap, we conduct a comprehensive analysis of multilingual LLM-as-a-Judge. Specifically, we evaluate five models from different model families across five diverse tasks involving 25 languages. Our findings reveal that LLMs struggle to achieve consistent judgment results across languages, with an average Fleiss' Kappa of approximately 0.3, and some models performing even worse. To investigate the cause of inconsistency, we analyze various influencing factors. We observe that consistency varies significantly across languages, with particularly poor performance in low-resource languages. Additionally, we find that neither training on multilingual data nor increasing model scale directly improves judgment consistency. These findings suggest that LLMs are not yet reliable for evaluating multilingual predictions. We finally propose an ensemble strategy which improves the consistency of the multilingual judge in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。