arXiv:2506.09813cs.LGcs.GT2025-06NeurIPS被引 2

用社会选择理论设计更公平的模型评估指标子集筛选方法

Metritocracy: Representative Metrics for Lite Benchmarks

  • 引入位置代表性与比例代表性,量化评估指标的公平覆盖程度
  • 证明最坏情况下保证性质所需的最少指标数量上限与下限
  • 在大模型和医院质量评估中验证理论实用性,适合评估系统设计者

大语言模型评估中常需从完整指标集选取子集以提升效率或可解释性,但‘代表性’定义模糊。本文借鉴社会选择理论,形式化提出两种代表性概念:位置代表性(确保每个候选在各排名截断处均有充分覆盖)与位置比例性(确保任意候选在任何排名位置上不过度或不足代表,误差小于小量)。我们证明了在最坏情况下保证任一性质所需最小指标数的上下界。还研究了允许额外输入需优先代表的指标组的广义形式。最后通过大模型评估与医院质量评估的实际案例,将理论与实践结合。

原文摘要 · Abstract (English)

A common problem in LLM evaluation is how to choose a subset of metrics from a full suite of possible metrics. Subset selection is usually done for efficiency or interpretability reasons, and the goal is often to select a ``representative'' subset of metrics. However, ``representative'' is rarely clearly defined. In this work, we use ideas from social choice theory to formalize two notions of representation for the selection of a subset of evaluation metrics. We first introduce positional representation, which guarantees every alternative is sufficiently represented at every position cutoff. We then introduce positional proportionality, which guarantees no alternative is proportionally over- or under-represented by more than a small error at any position. We prove upper and lower bounds on the smallest number of metrics needed to guarantee either of these properties in the worst case. We also study a generalized form of each property that allows for additional input on groups of metrics that must be represented. Finally, we tie theory to practice through real-world case studies on both LLM evaluation and hospital quality evaluation.

模型评估社会选择指标筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。