arXiv:2601.07651cs.AIcs.GT2026-01被引 2

提出主动评估框架,用动态选任务降低智能体评估成本。

Active Evaluation of General Agents: Problem Definition and Comparison of Baseline Algorithms

  • 在线选择最有效任务和智能体进行评分,动态优化评估效率。
  • 真实数据中软康多塞优化比埃洛系统更优,合成数据上性能相当。
  • 按比例代表原则选任务,在任务差异大时能更快减少排名误差。

随着智能体具备越来越广泛的任务能力,对其进行全面评估的复杂性和成本显著上升。传统评估方法面临任务相关性与随机性问题,需大量样本才能准确比较,导致成本增加。本文提出智能体跨任务主动评估的形式化定义与概念框架,将排名算法性能作为评估样本数量的函数进行衡量。不同于预处理阶段的数据筛选或压缩,我们采用在线决策机制:每轮由排名算法自主选择任务与智能体进行评分采样,评估算法在每轮输出智能体排名,并基于真实排名动态评估其表现。我们在合成数据及模拟的Atari游戏智能体真实评估数据上对比多种基线方法。结果表明,尽管埃洛系统理论上有已知缺陷,但实际中仍是高效降低排名误差的可靠选择;而近期提出的软康多塞优化方法在合成数据上表现接近埃洛,在真实Atari评估中显著优于埃洛。当任务与真实排序偏差较大时,基于比例代表性的任务选择策略可带来更高的排名误差下降速率。

原文摘要 · Abstract (English)

As intelligent agents become more generally-capable, i.e. able to master a wide variety of tasks, the complexity and cost of properly evaluating them rises significantly. Tasks that assess specific capabilities of the agents can be correlated and stochastic, requiring many samples for accurate comparisons, leading to added costs. In this paper, we propose a formal definition and a conceptual framework for active evaluation of agents across multiple tasks, which assesses the performance of ranking algorithms as a function of number of evaluation data samples. Rather than curating, filtering, or compressing existing data sets as a preprocessing step, we propose an online framing: on every iteration, the ranking algorithm chooses the task and agents to sample scores from. Then, evaluation algorithms report a ranking of agents on each iteration and their performance is assessed with respect to the ground truth ranking over time. Several baselines are compared under different experimental contexts, with synthetic generated data and simulated online access to real evaluation data from Atari game-playing agents. We find that the classical Elo rating system -- while it suffers from well-known failure modes, in theory -- is a consistently reliable choice for efficient reduction of ranking error in practice. A recently-proposed method, Soft Condorcet Optimization, shows comparable performance to Elo on synthetic data and significantly outperforms Elo on real Atari agent evaluation. When task variation from the ground truth is high, selecting tasks based on proportional representation leads to higher rate of ranking error reduction.

智能体评估主动学习排名优化Atari

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。