用鲁棒彩票提升语言模型评估,避免偏好分歧导致的偏差。
Robust AI Evaluation through Maximal Lotteries
- 引入鲁棒彩票机制,优化最差情况下的表现
- 在大规模偏好数据集上提升胜率保障的可靠性
- 适合关注公平评估与多元用户需求的研究者
当前主观任务的语言模型评估普遍采用成对比较:标注者从两个回复中选择更优者。排行榜将这些比较聚合为单一的布拉德利-特里(BT)排名,强制异质偏好形成全序,违背社会选择理论的基本要求。相比之下,社会选择理论中的最大彩票方法可在不假设偏好结构的前提下聚合成对偏好。然而我们发现,最大彩票对偏好异质性极为敏感,可能青睐在特定任务或用户群体上严重表现不佳的模型。为此,本文提出鲁棒彩票,在合理偏好数据变动下优化最差情况性能。在大规模偏好数据集上,鲁棒彩票提供了更可靠的胜率保证,并恢复出稳定表现的顶级模型集合。通过从排名转向多赢家集合,鲁棒彩票为支持多样化人类偏好的互补型AI生态提供了原则性路径。
原文摘要 · Abstract (English)
The standard way to evaluate language models on subjective tasks is through pairwise comparisons: an annotator chooses the "better" of two responses to a prompt. Leaderboards aggregate these comparisons into a single Bradley-Terry (BT) ranking, forcing heterogeneous preferences into a total order and violating basic social-choice desiderata. In contrast, social choice theory provides an alternative approach called maximal lotteries, which aggregates pairwise preferences without imposing any assumptions on their structure. However, we show that maximal lotteries are highly sensitive to preference heterogeneity and can favor models that severely underperform on specific tasks or user subpopulations. We introduce robust lotteries that optimize worst-case performance under plausible shifts in the preference data. On large-scale preference datasets, robust lotteries provide more reliable win rate guarantees across the annotator distribution and recover a stable set of top-performing models. By moving from rankings to pluralistic sets of winners, robust lotteries offer a principled step toward an ecosystem of complementary AI systems that serve the full spectrum of human preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。