arXiv:2609.07785cs.AI2026-09

排行榜分数不能直接说明哪个智能体更好,要看评估条件和不确定性。

What Does an LLM-Agent Leaderboard Rank Actually Compare?

论文配图:What Does an LLM-Agent Leaderboard Rank Actually Compare?
图 1 · 摘自论文原文
  • 用明确的目标和测量来源做对比,避免误判
  • 相近排名常无法确定优劣,需考虑不确定性
  • 适合评估系统设计者和研究者参考

大型语言模型智能体排行榜常被理解为排名靠前的智能体更优,但公开评估记录可能因任务组合、标签来源、发布细节或成本规则不同而无法支持此结论。本文研究排行榜分数实际衡量什么,并界定何时可做出两两优劣判断。提出一种以估计量为中心的成对比较方法,明确定义比较目标与数据来源,检查共同支持范围,并基于指定不确定性规则和实际差异阈值评估差异显著性。受控实验在小样本条件下验证决策标签,揭示目标权重重分配时不确定性的重要性。在SWE-bench、AgentRewardBench和tau2-bench上,排名接近的系统常无法区分;代理标签和效用规则也可能改变优选结果。DataAgentBench和Open Agent展示粗粒度公开记录仍可估计的部分内容。排行榜分数仅反映已发布的评估结果,而精细的优劣主张还依赖于估计量及不确定性规则的解释。

原文摘要 · Abstract (English)

An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent. Public evaluation logs may not support that conclusion when systems differ in task mixture, label source, release detail, or cost rule. We study what leaderboard scores estimate and when they justify pairwise superiority conclusions. Our estimand-aware pairwise procedure states the comparison target and measurement source, checks common support, and evaluates the supported difference using a stated uncertainty rule and practical margin. Controlled checks evaluate the decision labels under known finite-sample conditions and show why uncertainty must be included when judging sensitivity to target reweighting. Across SWE-bench, AgentRewardBench, and tau2-bench, close rank differences are often unresolved; proxy labels and utility rules can also change which system is selected. DataAgentBench and Open Agent show what remains estimable from coarser public records. A leaderboard score summarizes a released evaluation, whereas a fine-grained superiority claim additionally depends on the estimand and uncertainty rule used to interpret the difference.

智能体评估排行榜分析不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。