arXiv:2501.00560cs.CLcs.AI2025-01NAACL被引 24

评测大模型时,自动评分系统在相近模型间表现骤降,需谨慎选择评估组件。

Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference

  • 通过控制实验优化自动评测框架的四个核心组件组合。
  • 当模型性能接近时,自动评测准确率显著下降。
  • 需单独评估评测模型在系统中的表现,而非仅看单个任务精度。

评估不同大模型的能力与人类偏好对齐程度至关重要。由于人工评估成本高、耗时长,自动大模型评测框架(LLM bencher)成为必需。该框架包含输入集、评估模型、评估类型和聚合方法四个部分。然而,以往研究未充分探讨各组件的选择及其组合对结果的影响。本文通过受控实验,提出组件选择建议,并发现:当模型性能相近时,自动评测系统表现急剧下降,暴露了当前方法的局限性。此外,评估模型在个体任务中的表现(如输出选择准确性)与其作为评测组件的整体有效性并不总一致,强调了对评测系统进行专门的体系级评估的重要性。

原文摘要 · Abstract (English)

Evaluating and ranking the capabilities of different LLMs is crucial for understanding their performance and alignment with human preferences. Due to the high cost and time-consuming nature of human evaluations, an automatic LLM bencher (i.e., an automatic evaluation framework that aims to rank LLMs based on their alignment with human preferences) is indispensable. An automatic LLM bencher consists of four components: the input set (e.g., a user instruction), the evaluation model (e.g., an LLM), the evaluation type (e.g., pairwise comparison), and the aggregation method (e.g., the ELO rating system). However, previous work has not thoroughly explored how to select these components or how their different combinations influence the results. In this work, through controlled experiments, we provide a series of recommendations on how to choose each component to better automate the evaluation of LLMs. Furthermore, we discovered that when evaluating LLMs with similar performance, the performance of the automatic LLM bencher declines sharply, underscoring the limitations of current benchers and calling for future work. Lastly, we found that the evaluation models' performance at the instance level (e.g., the accuracy of selecting the best output) does not always align with their effectiveness when used as a component of a bencher, highlighting the importance of dedicated system-level evaluation of benchers.

大模型评测自动评估人类偏好对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。