arXiv:2410.07069cs.CLcs.AI2024-10NAACL被引 24

全面评估指令遵循的评测模型,发现最佳评测组合。

ReIFE: Re-evaluating Instruction-Following Evaluation

  • 用25个基础模型+15种评测协议,系统测试LLM评测器性能。
  • 低能力模型在评测协议优化后提升更明显,表现更稳定。
  • 需多数据集验证,单一数据集结果可能不具代表性。

指令遵循的自动评估通常依赖大语言模型(LLM)判断响应质量,但现有评估缺乏对基础模型与评测协议的全面考察。为此,我们开展了一项大规模元评估,涵盖25个基础LLM和15种近期提出的评测协议,在4个经人工标注的数据集上评估了LLM评测器的准确率。结果表明:(1)基础模型的性能排名在不同评测协议下基本一致,低能力模型在协议优化后提升更显著;(2)评测协议的有效性高度依赖基础模型能力,需使用多层级能力模型以确保稳健性;(3)不同数据集上的评测结果并不总一致,因此严谨评估需覆盖具有差异特征的多个数据集。我们发布了名为ReIFE的元评估套件,包含超过500种LLM评测配置的代码与结果,支持未来指令遵循评估研究。

原文摘要 · Abstract (English)

The automatic evaluation of instruction following typically involves using large language models (LLMs) to assess response quality. However, there is a lack of comprehensive evaluation of these LLM-based evaluators across two dimensions: the base LLMs and the evaluation protocols. Therefore, we present a thorough meta-evaluation of instruction following, including 25 base LLMs and 15 recently proposed evaluation protocols, on 4 human-annotated datasets, assessing the evaluation accuracy of the LLM-evaluators. Our evaluation allows us to identify the best-performing base LLMs and evaluation protocols with a high degree of robustness. Moreover, our large-scale evaluation reveals: (1) Base LLM performance ranking remains largely consistent across evaluation protocols, with less capable LLMs showing greater improvement from protocol enhancements; (2) Robust evaluation of evaluation protocols requires many base LLMs with varying capability levels, as protocol effectiveness can depend on the base LLM used; (3) Evaluation results on different datasets are not always consistent, so a rigorous evaluation requires multiple datasets with distinctive features. We release our meta-evaluation suite ReIFE, which provides the codebase and evaluation result collection for more than 500 LLM-evaluator configurations, to support future research in instruction-following evaluation.

评测评估LLM指令遵循

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。