arXiv:2608.26623cs.AI2026-08中稿 · EMNLP

首个系统评估LLM在复杂工具调用任务中判别能力的基准,揭示其固有局限。

AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling

论文配图:AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
图 1 · 摘自论文原文
  • 构建涵盖六种拓扑、三难度等级的3808个实例,测试LLM判官在有无真值条件下的表现。
  • 难度越高,判官一致性越差;无真值时所有判官收敛至77-82%区间,体现任务结构性天花板。
  • 真值并非总有益,部分模型反而因过度依赖真值导致评分偏差,结构化评分表效果有限。

LLM判官广泛用于评估智能体工具调用系统,但其在结构化、依赖驱动工作流中的可靠性尚未得到充分检验。我们提出AgentJudgeBench,首个系统研究LLM作为判官在工作流有向无环图(DAG)上对智能体工具调用的可靠性,区别于开放文本或偏好评估的通用判官任务。该基准包含3,808个实例,覆盖六种DAG拓扑和三种难度层级,由五种生成器(3B–70B开源模型及GPT-5.4)与六种判官(20B至前沿规模)在有/无真值条件下进行配对评估。判官一致性随任务难度单调下降,无真值时下降速度加快1.5倍;在高难度任务下无真值时,所有六名判官均收敛至77–82%的狭窄区间,表明任务难度构成主要结构性上限,尽管弱生成器的上限高度部分受提示影响,模型规模无法突破。真值暴露并非普遍有益:对GPT-5.4(下降1.5个百分点)和Gemini-2.5-Pro(下降3.9个百分点)反而降低一致性,符合过锚定现象。在缓解策略中,思维链推理与判官温度调节影响微乎其微,而结构化评分标准可提升一致性达6.5个百分点,但对不同判官-生成器组合泛化性不一。有真值时,QwQ-32B最接近程序参考;人类验证研究显示,GPT-OSS-120B最具人类对齐性;无真值时,前沿判官仅在共享天花板内边际领先。这些结果揭示当前LLM判官的根本局限,并为智能体系统的可靠评估提供实践指导。

原文摘要 · Abstract (English)

LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systematically study LLM-as-a-judge reliability for agentic tool-calling over workflow DAGs, as distinct from the broader LLM-as-a-judge task of open-ended text or preference evaluation. The benchmark comprises 3,808 instances spanning six DAG topologies and three difficulty tiers, evaluated with five generators (3B-70B open-weight models and GPT-5.4) and six judges (20B to frontier scale) under paired with- and without-ground-truth conditions. Judge alignment degrades monotonically with task difficulty, 1.5x faster without ground truth, and on hard queries without ground truth all six judges converge to a narrow 77-82% band regardless of scale, revealing a structural ceiling driven primarily by task difficulty, though its height is partly prompt-dependent for weaker generators, that model capacity alone cannot overcome. Ground-truth exposure is not uniformly beneficial: it reduces alignment for GPT-5.4 (1.5 pp) and Gemini-2.5-Pro (3.9 pp), consistent with over-anchoring. Among mitigation strategies, chain-of-thought reasoning and judge temperature both have negligible effect, while structured evaluation rubrics improve alignment by up to 6.5 pp but do not generalize uniformly across judge-generator pairs. With ground truth, QwQ-32B best matches the programmatic reference, while a human validation study identifies GPT-OSS-120B as the most human-aligned judge; without it, frontier judges lead only marginally within the shared ceiling. These results expose fundamental limitations of current LLM judges and yield practical guidelines for reliable evaluation in agentic systems.

大模型评估智能体工具调用判官可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。