arXiv:2502.11454cs.CL2025-02ICLR

统一优化多个目标,让大模型评估更准更快省资源。

UniCBE: An Uniformity-driven Comparing Based Evaluation Framework with Unified Multi-Objective Optimization

  • 构建三个解耦采样矩阵,同时解决偏差、不确定性和更新问题。
  • 在AlpacaEval上节省超17%成本,相关性达0.995以上。
  • 适合持续迭代模型的高效评估场景,尤其节省长期评测开支。

人类偏好在衡量大语言模型并引导其对齐人类价值观中起关键作用。然而,当前基于比较的评估(CBE)方法通常仅聚焦单一优化目标,难以有效利用稀缺但宝贵的偏好信号。为解决此问题,我们深入分析提升CBE准确率、收敛速度和可扩展性的关键因素:抑制采样偏差、平衡不确定性下降过程以及缓解更新不确定性。基于这些指导原则,我们提出UniCBE——一种统一的均匀性驱动型CBE框架,通过构建并集成三个解耦的采样概率矩阵,分别确保特定方面的均匀性,从而同步优化核心目标。我们进一步对最优元组采样与偏好聚合策略进行消融实验,实现高效评估。在AlpacaEval基准上,UniCBE在评估预算减少超过17%的情况下,仍实现与真实标签皮尔逊相关系数超过0.995,展现优异准确性与收敛性;在新模型持续引入的场景中,评估成本甚至可节省超50%,凸显其卓越可扩展性。

原文摘要 · Abstract (English)

Human preference plays a significant role in measuring large language models and guiding them to align with human values. Unfortunately, current comparing-based evaluation (CBE) methods typically focus on a single optimization objective, failing to effectively utilize scarce yet valuable preference signals. To address this, we delve into key factors that can enhance the accuracy, convergence, and scalability of CBE: suppressing sampling bias, balancing descending process of uncertainty, and mitigating updating uncertainty. Following the derived guidelines, we propose UniCBE, a unified uniformity-driven CBE framework which simultaneously optimize these core objectives by constructing and integrating three decoupled sampling probability matrices, each designed to ensure uniformity in specific aspects. We further ablate the optimal tuple sampling and preference aggregation strategies to achieve efficient CBE. On the AlpacaEval benchmark, UniCBE saves over 17% of evaluation budgets while achieving a Pearson correlation with ground truth exceeding 0.995, demonstrating excellent accuracy and convergence. In scenarios where new models are continuously introduced, UniCBE can even save over 50% of evaluation costs, highlighting its improved scalability.

大模型评估偏好学习高效评估优化框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。