arXiv:2602.07593cs.LGcs.GT2026-02被引 3

破解多指标评测难题,让模型排名更稳定可靠

Beyond Arrow: From Impossibility to Possibilities in Multi-Criteria Benchmarking

  • 将评测视为社会选择问题,用投票机制整合多维度指标
  • 发现特定偏好结构下可避免悖论,实现合理模型排序
  • 验证主流评测集满足条件,适用于实际模型对比

现代评测基准如 HELM、MMLU 考察准确率、鲁棒性与效率等多个指标。当试图将这些指标合并为单一排名时,常见聚合方法可能出现不一致或对模型集合变化敏感的问题。本文将此聚合过程形式化为社会选择问题:每个指标在每项数据集上诱导出模型的偏好排序,而基准算子则跨指标聚合这些排序。尽管已有研究关注阿罗不可能性定理,但本文认为该不可能性常源于病态例子,并识别出使此类问题消失的充分条件。具体而言,本文研究了三种排名组合的限制:单峰偏好、组可分偏好和距离受限偏好,在这三类结构下,基准算子能构造出行为良好的模型排名。实证上,本文考察了 HELM、MMLU 等多个现代基准套件,验证了不同评测任务中上述结构性条件的实际成立情况。

原文摘要 · Abstract (English)

Modern benchmarks such as HELM MMLU account for multiple metrics like accuracy, robustness and efficiency. When trying to turn these metrics into a single ranking, natural aggregation procedures can become incoherent or unstable to changes in the model set. We formalize this aggregation as a social choice problem where each metric induces a preference ranking over models on each dataset, and a benchmark operator aggregates these votes across metrics. While prior work has focused on Arrow's impossibility result, we argue that the impossibility often originates from pathological examples and identify sufficient conditions under which these disappear, and meaningful multi-criteria benchmarking becomes possible. In particular, we deal with three restrictions on the combinations of rankings and prove that on single-peaked, group-separable and distance-restricted preferences, the benchmark operator allows for the construction of well-behaved rankings of the involved models. Empirically, we investigate several modern benchmark suites like HELM MMLU and verify which structural conditions are fulfilled on which benchmark problems.

评测基准多指标排序社会选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。