arXiv:2411.00119cs.MAcs.LG2024-11被引 8

用投票机制优化通用智能体排名,更准更鲁棒。

Soft Condorcet Optimization for Ranking of General Agents

  • 基于软康多塞原则,从偏好数据中推断最优排名
  • 在865组数据上平均误差仅0.043(归一化Kendall-tau)
  • 适合评估多任务智能体,尤其适用于缺失数据场景

推动AI模型与智能体发展需在标准基准上比较其性能;对于通用智能体,需在多种任务上聚合个体表现。本文提出一种受社会选择理论启发的新型排名方法——软康多塞优化(SCO),以计算使预测错误最少的最优排名。该排名是将评估数据视为噪声样本时的极大似然估计,满足康多塞原始投票准则。当存在康多塞胜者时,SCO评分能将其最大化,而经典Elo系统不保证此性质。我们提出三种优化算法并评估其性能:在865个来自PrefLib公开档案的偏好配置中,SCO排名与最优排名的归一化Kendall-tau距离平均为0至0.043;在模拟的噪声锦标赛中,当超过59%偏好数据缺失时,仍优于多个基线;在包含52,958名玩家、31,049场游戏的经典七人版《同盟》游戏中,基于留出测试集的评估显示,SCO对最优排名的逼近效果最佳。

原文摘要 · Abstract (English)

Driving progress of AI models and agents requires comparing their performance on standardized benchmarks; for general agents, individual performances must be aggregated across a potentially wide variety of different tasks. In this paper, we describe a novel ranking scheme inspired by social choice frameworks, called Soft Condorcet Optimization (SCO), to compute the optimal ranking of agents: the one that makes the fewest mistakes in predicting the agent comparisons in the evaluation data. This optimal ranking is the maximum likelihood estimate when evaluation data (which we view as votes) are interpreted as noisy samples from a ground truth ranking, a solution to Condorcet's original voting system criteria. SCO ratings are maximal for Condorcet winners when they exist, which we show is not necessarily true for the classical rating system Elo. We propose three optimization algorithms to compute SCO ratings and evaluate their empirical performance. When serving as an approximation to the Kemeny-Young voting method, SCO rankings are on average 0 to 0.043 away from the optimal ranking in normalized Kendall-tau distance across 865 preference profiles from the PrefLib open ranking archive. In a simulated noisy tournament setting, SCO achieves accurate approximations to the ground truth ranking and the best among several baselines when 59\% or more of the preference data is missing. Finally, SCO ranking provides the best approximation to the optimal ranking, measured on held-out test sets, in a problem containing 52,958 human players across 31,049 games of the classic seven-player game of Diplomacy.

智能体评估排序优化投票机制多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。