arXiv:2412.15524cs.CLcs.AI2024-12被引 8

用真人回复提升大模型指令遵循评估可靠性,最高提升3.2%。

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models

  • 采用人类撰写回复作为评估依据,减少模型自评偏差。
  • 在11类任务中实现与人类判断最高3.2%的吻合度提升。
  • 适合关注评估公平性与真实表现的研究者和开发者。

大语言模型(LLMs)指令遵循能力的评估长期依赖强大语言模型作为评判者,引入了未解决的偏差,使结果偏离人类判断。本文重新评估多种自动评估方法在广泛指令跟随任务中的表现,发现使用人类撰写回复能显著提升评估可靠性,使与人类判断的一致性最高提升3.2%。我们还发现,人类回复提供了与模型生成回复正交的评估视角,应在比较模型输出时作为额外上下文。基于此,我们构建了新基准HREF(Human Response-Guided Evaluation of Instruction Following),包含4,258个样本,覆盖11个任务类别,采用按类别选择最优评估方法的复合设置。HREF强调各任务个体性能,且无数据污染。此外,我们研究了评估集规模、评判模型、基线模型及提示模板等关键设计的影响,并上线实时排行榜,对私有评估集进行持续评测。

原文摘要 · Abstract (English)

Evaluating the capability of Large Language Models (LLMs) in following instructions has heavily relied on a powerful LLM as the judge, introducing unresolved biases that deviate the judgments from human judges. In this work, we reevaluate various choices for automatic evaluation on a wide range of instruction-following tasks. We experiment with methods that leverage human-written responses and observe that they enhance the reliability of automatic evaluations across a wide range of tasks, resulting in up to a 3.2% improvement in agreement with human judges. We also discovered that human-written responses offer an orthogonal perspective to model-generated responses in following instructions and should be used as an additional context when comparing model responses. Based on these observations, we develop a new evaluation benchmark, Human Response-Guided Evaluation of Instruction Following (HREF), comprising 4,258 samples across 11 task categories with a composite evaluation setup, employing a composite evaluation setup that selects the most reliable method for each category. In addition to providing reliable evaluation, HREF emphasizes individual task performance and is free from contamination. Finally, we study the impact of key design choices in HREF, including the size of the evaluation set, the judge model, the baseline model, and the prompt template. We host a live leaderboard that evaluates LLMs on the private evaluation set of HREF.

指令遵循评估基准人类反馈大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。