arXiv:2507.06774cs.CL2025-07被引 2

用检查清单提升多语言大模型评分能力,无需训练

Checklist Engineering Empowers Multilingual LLM Judges

  • 基于检查清单设计评分逻辑,不依赖训练数据
  • 在多语言基准上性能接近GPT-4o,超越多数基线
  • 适合资源有限但需多语言评估的研究者

自动化文本评估是自然语言处理的核心挑战。近年来,使用大语言模型作为评估者(即LLM-as-a-Judge)成为主流,但多语言场景下的研究仍不足。现有方法常依赖专有模型或大量微调数据,成本高、效率低。本文提出无需训练的Checklist Engineering-based LLM-as-a-Judge(CE-Judge)框架,利用检查清单机制实现多语言评估。在多种语言及三个基准数据集上,无论点对点还是成对评估设置,该方法均显著优于基线,性能与GPT-4o相当。

原文摘要 · Abstract (English)

Automated text evaluation has long been a central issue in Natural Language Processing (NLP). Recently, the field has shifted toward using Large Language Models (LLMs) as evaluators-a trend known as the LLM-as-a-Judge paradigm. While promising and easily adaptable across tasks, this approach has seen limited exploration in multilingual contexts. Existing multilingual studies often rely on proprietary models or require extensive training data for fine-tuning, raising concerns about cost, time, and efficiency. In this paper, we propose Checklist Engineering based LLM-as-a-Judge (CE-Judge), a training-free framework that uses checklist intuition for multilingual evaluation with an open-source model. Experiments across multiple languages and three benchmark datasets, under both pointwise and pairwise settings, show that our method generally surpasses the baselines and performs on par with the GPT-4o model.

多语言大模型评估零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。