对比10个大模型在多语言文本生成评估中的表现,发现高资源语言更准确,微调能提升整体性能。
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
- 用大模型作评估器时,去掉参考答案可提升跨语言表现
- 高资源语言的评估相关性高于低资源语言,差距明显
- 针对特定语言微调后,模型在多种语言上都更可靠
以往研究显示大语言模型在多语言自然语言生成评估中具有潜力,但尚未充分探索其在不同语言间的评估能力差异。本研究对10个近期大模型进行了全面分析,涵盖高资源与低资源语言,通过相关性分析、扰动攻击和微调实验发现:1)在提示中不包含参考答案,并使用参数量大的大模型作为评估器,可在多种语言上获得更好表现;2)多数大模型在高资源语言上的评估结果与人类判断的相关性高于低资源语言;3)对扰动攻击最敏感的语言,通常也是与人类判断相关性最高的语言;4)针对某一语言进行微调后,模型在多种语言上的评估性能均得到广泛提升。研究揭示了大模型评估能力在不同语言间存在不平衡,提示低资源语言场景需更多关注。
原文摘要 · Abstract (English)
Previous research has shown that LLMs have potential in multilingual NLG evaluation tasks. However, existing research has not fully explored the differences in the evaluation capabilities of LLMs across different languages. To this end, this study provides a comprehensive analysis of the multilingual evaluation performance of 10 recent LLMs, spanning high-resource and low-resource languages through correlation analysis, perturbation attacks, and fine-tuning. We found that 1) excluding the reference answer from the prompt and using large-parameter LLM-based evaluators leads to better performance across various languages; 2) most LLM-based evaluators show a higher correlation with human judgments in high-resource languages than in low-resource languages; 3) in the languages where they are most sensitive to such attacks, they also tend to exhibit the highest correlation with human judgments; and 4) fine-tuning with data from a particular language yields a broadly consistent enhancement in the model's evaluation performance across diverse languages. Our findings highlight the imbalance in LLMs'evaluation capabilities across different languages and suggest that low-resource language scenarios deserve more attention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。