arXiv:2607.12338cs.AI2026-07

研究少样本评估能否得出与完整测试一致的智能体结论。

How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks

论文配图:How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks
图 1 · 摘自论文原文
  • 通过重放公开数据集任务记录,检验部分评估的可靠性
  • 不同基准需不同任务比例:AppWorld仅15%即可,SWE-bench需90%
  • 提出评估报告应明确性能差距、任务选择等关键规则

智能体评测常在所有任务完成后比较两个智能体,但评估成本高导致部分运行颇具吸引力。仅看任务占比无法判断部分评估是否支持完整评测的结论。本文通过重放SWE-bench、AppWorld和tau-bench的公开任务级记录,研究此问题。一个部分预算被认为足够,当且仅当它支持完整评测的决策、覆盖必要任务组,且未解决的比较不超过目标比例。所需任务比例差异显著:在5%预算网格的0%严格阈值下,AppWorld在15%时达标,tau-bench在25%时达标,SWE-bench Verified在90%时达标;而SWE-bench Lite在主要覆盖率规则下,95%仍未达标。部分评估报告应说明一智能体需领先多少、任务如何选取、所需覆盖率规则、决策机制及可保留未决比较数。

原文摘要 · Abstract (English)

Agent benchmarks often compare two agents after all tasks have run, but costly evaluations make partial runs tempting. A task fraction alone does not show whether a partial run supports the same pairwise conclusion as the completed benchmark. We study this question by replaying completed public task-level records from SWE-bench, AppWorld, and tau-bench. A partial budget counts as enough only when it supports the completed benchmark's decision, covers required task groups, and leaves no more than a target fraction of comparisons unresolved. The required task fraction varies sharply. At the strict 0 percentage point threshold on a 5 percentage point budget grid, AppWorld first meets all targets at 15 percent, tau-bench at 25 percent, and SWE-bench Verified at 90 percent; SWE-bench Lite does not meet all targets by 95 percent under the primary coverage rule. Partial-evaluation reports should state how much one agent must outperform another, how tasks are selected, what coverage rule is required, what decision rule is used, and how many comparisons may remain unresolved.

智能体评测基准测试部分评估任务覆盖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。