arXiv:2602.02898cs.AIcs.CL2026-02

用实际偏好自动优化评测基准,让模型排名更贴近真实使用效果。

Aligning Language Model Benchmarks with Pairwise Preferences

  • 基于模型在任务中的表现和部署时的对比排名,自动调整评测题权重。
  • 新基准能准确预测未见过的模型在人类偏好上的相对排序。
  • 适合关注模型真实效能、追求可解释性评测的研究者使用。

语言模型评测基准广泛使用且计算高效,但许多研究发现其难以预测实际应用中的表现。为弥合这一差距,本文提出‘评测对齐’(benchmark alignment),利用少量模型性能信息自动更新离线评测基准,生成能预测特定测试场景下模型成对偏好的新静态基准。我们提出首个解决方案 BenchAlign,通过结合模型在各题目上的表现以及部署期间收集到的模型成对排名,学习出与偏好一致的评测题权重,使新基准能够根据人类偏好对未见过的模型进行排序。实验表明,对齐后的基准不仅能准确排序不同规模的模型,且保持可解释性。本工作揭示了将评测基准与实际人类偏好对齐的潜力,有望加速模型朝真实应用价值演进。

原文摘要 · Abstract (English)

Language model benchmarks are pervasive and computationally-efficient proxies for real-world performance. However, many recent works find that benchmarks often fail to predict real utility. Towards bridging this gap, we introduce benchmark alignment, where we use limited amounts of information about model performance to automatically update offline benchmarks, aiming to produce new static benchmarks that predict model pairwise preferences in given test settings. We then propose BenchAlign, the first solution to this problem, which learns preference-aligned weightings for benchmark questions using the question-level performance of language models alongside ranked pairs of models that could be collected during deployment, producing new benchmarks that rank previously unseen models according to these preferences. Our experiments show that our aligned benchmarks can accurately rank unseen models according to models of human preferences, even across different sizes, while remaining interpretable. Overall, our work provides insights into the limits of aligning benchmarks with practical human preferences, which stands to accelerate model development towards real utility.

评测对齐人类偏好模型排序可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。