用中等难度任务集可高效准确评估AI智能体排名。
Efficient Benchmarking of AI Agents
- 仅评估通过率30%-70%的任务,避免性能偏差。
- 任务数减少44%-70%,排名准确性仍高于随机采样。
- 适合需要快速比较智能体性能的研究者使用。
在综合基准上评估AI智能体成本高昂,因每次评估需与工具交互并进行多步推理。我们研究小任务子集是否能在更低成本下保持智能体排名。不同于静态语言模型基准,智能体评估受框架(scaffold)驱动的分布偏移影响,性能依赖于底层模型的封装方式。在8个基准、33种智能体框架和70多个模型配置下,我们发现绝对分数预测在此偏移下表现下降,而排名预测保持稳定。利用这一差异,提出无需优化的简单协议:仅在历史通过率介于30%-70%的任务上评估新智能体。该中等难度过滤策略基于项目反应理论,使评估任务数减少44%-70%,同时在框架和时间偏移下仍保持高排名保真度。其结果优于随机采样(种子间方差大)和贪婪任务选择(在分布偏移下表现差)。结果表明,可靠排行榜无需完整基准评估。
原文摘要 · Abstract (English)
Evaluating AI agents on comprehensive benchmarks is expensive because each evaluation requires interactive rollouts with tool use and multi-step reasoning. We study whether small task subsets can preserve agent rankings at substantially lower cost. Unlike static language model benchmarks, agent evaluation is subject to scaffold-driven distribution shift, since performance depends on the framework wrapping the underlying model. Across eight benchmarks, 33 agent scaffolds, and 70+ model configurations, we find that absolute score prediction degrades under this shift, while rank-order prediction remains stable. Exploiting this asymmetry, we propose a simple optimization-free protocol: evaluate new agents only on tasks with intermediate historical pass rates (30-70%). This mid-range difficulty filter, motivated by Item Response Theory, reduces the number of evaluation tasks by 44-70% while maintaining high rank fidelity under scaffold and temporal shifts. It provides more reliable rankings than random sampling, which exhibits high variance across seeds, and outperforms greedy task selection under distribution shift. These results suggest that reliable leaderboard ranking does not require full-benchmark evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。