arXiv:2602.03238cs.AI2026-02被引 1

LLM代理性能不是单一指标,不同评估目标需区分对待。

"LLM Agent Performance" Is Not a Single Evaluation Target

  • 区分模型对比与系统评估,避免混淆评价标准。
  • 同一配置下比较模型差异,结果更可信。
  • 适合关注评估公平性与系统鲁棒性的研究者。

LLM代理基准测试得分不仅受模型影响,还受代理框架、环境、评估器和推理预算等因素影响。统一执行环境通过在相同配置下评估候选模型,使观察到的差异更归因于模型本身。然而,模型比较只是代理基准的一种用途。其他评估则用于比较完整代理系统,或检验固定模型/系统在预设运行条件变化下的稳定性。这些结果均被统称为「LLM代理性能」。本文认为,『LLM代理性能』不代表单一评价目标。在参考堆栈下进行模型比较与完整系统比较回答的是不同问题,而鲁棒性测试则关注任一结论是否在预设条件下仍成立。因此,得分所支持的结论取决于声明的候选边界和条件策略。我们推导了对排行榜、结果报告和基准版本管理的启示,表明区分这些类别可在保持公平比较的同时,容纳系统创新与鲁棒性分析。

原文摘要 · Abstract (English)

LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget. Unified execution controls these non-model factors by evaluating candidate models under the same configuration, making observed differences more attributable to the models themselves. However, model comparison is only one use of agent benchmarks. Other evaluations compare complete agent systems or test whether a fixed model or system remains stable across predeclared changes in its operating conditions. These results can all be reported under the common label of "LLM agent performance." Our position is that "LLM agent performance" does not denote a single evaluation target. Model comparisons under a reference stack and comparisons of complete agent systems answer different questions, while robustness asks whether either conclusion persists across predeclared conditions. The claim supported by a score therefore depends on the declared candidate boundary and condition policy. We derive implications for leaderboards, result reporting, and benchmark versioning, showing how distinguishing these classes preserves fair comparison while accommodating system innovation and robustness analysis.

LLM代理评估基准系统鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。