全面评估指令遵循的评测模型,发现最佳评测组合。
ReIFE: Re-evaluating Instruction-Following Evaluation
- 用25个基础模型+15种评测协议,系统测试LLM评测器性能。
- 低能力模型在评测协议优化后提升更明显,表现更稳定。
- 需多数据集验证,单一数据集结果可能不具代表性。
指令遵循的自动评估通常依赖大语言模型(LLM)判断响应质量,但现有评估缺乏对基础模型与评测协议的全面考察。为此,我们开展了一项大规模元评估,涵盖25个基础LLM和15种近期提出的评测协议,在4个经人工标注的数据集上评估了LLM评测器的准确率。结果表明:(1)基础模型的性能排名在不同评测协议下基本一致,低能力模型在协议优化后提升更显著;(2)评测协议的有效性高度依赖基础模型能力,需使用多层级能力模型以确保稳健性;(3)不同数据集上的评测结果并不总一致,因此严谨评估需覆盖具有差异特征的多个数据集。我们发布了名为ReIFE的元评估套件,包含超过500种LLM评测配置的代码与结果,支持未来指令遵循评估研究。
原文摘要 · Abstract (English)
The automatic evaluation of instruction following typically involves using large language models (LLMs) to assess response quality. However, there is a lack of comprehensive evaluation of these LLM-based evaluators across two dimensions: the base LLMs and the evaluation protocols. Therefore, we present a thorough meta-evaluation of instruction following, including 25 base LLMs and 15 recently proposed evaluation protocols, on 4 human-annotated datasets, assessing the evaluation accuracy of the LLM-evaluators. Our evaluation allows us to identify the best-performing base LLMs and evaluation protocols with a high degree of robustness. Moreover, our large-scale evaluation reveals: (1) Base LLM performance ranking remains largely consistent across evaluation protocols, with less capable LLMs showing greater improvement from protocol enhancements; (2) Robust evaluation of evaluation protocols requires many base LLMs with varying capability levels, as protocol effectiveness can depend on the base LLM used; (3) Evaluation results on different datasets are not always consistent, so a rigorous evaluation requires multiple datasets with distinctive features. We release our meta-evaluation suite ReIFE, which provides the codebase and evaluation result collection for more than 500 LLM-evaluator configurations, to support future research in instruction-following evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。