构建首个代理任务元评估数据集,用于衡量大模型评价的可靠性。
Counsel: A Meta-Evaluation Dataset for Agentic Tasks
- 用人类标注对比大模型对代理任务轨迹的批评,检验其准确性。
- 发现更强模型和更多推理可提升评价与人工一致率,最高达88%。
- 适合研究评测模型对齐、改进自动评估系统的人使用。
随着代理系统处理越来越复杂的多步骤任务,对其行为轨迹的评估成为主要瓶颈——在主流代理基准上人工标注单条轨迹需数小时,难以规模化用于性能测量或训练数据构建。这导致广泛依赖大模型作为评判者(LLM-as-a-judge, LLMJ)进行大规模过程与结果级评估,但其评判质量常未被验证。本文提出Counsel,首个公开的代理任务元评估数据集。该数据集包含开放权重大模型对tau-bench(客户支持代理)和DA-Code(编码代理)两个基准的流程级批评,以及人类对这些批评的元评估。人类标注者将每项标记为“准确”、“位置正确但推理差”或“不应标记”,达成可靠的标注一致性(Krippendorff's alpha = 0.78)。数据集按错误定位与推理质量分层,可用于校准、改进或训练大模型评判者。对比不同开放权重评判模型,我们发现更强大模型和更多推理努力均能提升与人工的一致性,最强模型在定位上达到约88%,在推理上约65%。Counsel基于开放权重模型生成,采用宽松许可,旨在推动社区对大模型评价器的严谨研究与对齐优化。
原文摘要 · Abstract (English)
As agentic systems tackle increasingly complex multi-step tasks, evaluating their trajectories presents a major bottleneck - human annotation of a single trajectory on popular agentic benchmarks can take hours, making it difficult to scale evaluations for measuring performance or curating training data. This has driven widespread reliance on automated approaches such as LLM-as-a-judge (LLMJ) to critique agents at the process and outcome-levels at scale, however, the soundness of LLMJ critiques often goes unmeasured. Here, we introduce Counsel, the first public dataset of meta-evaluations for agentic tasks. Counsel consists of process-level critiques from open-weight LLMJs on two agent benchmarks: tau-bench (customer support agents) and DA-Code (coding agents), and human meta-evaluations of these critiques. Human annotators label critiques on each flagged error as "spot on", "correct location but poor reasoning", or "should not have flagged", achieving reliable inter-annotator agreement (Krippendorff's alpha of 0.78). The resulting dataset stratifies LLMJ critiques by human alignment across both error location within a trajectory and reasoning quality, serving as valuable data to calibrate, improve, or train LLMJs for agents. Comparing open-weight judges, we find that more capable judge models and more reasoning effort both enabled improved human agreement, with the strongest judge reaching ~88% agreement on location and ~65% on reasoning. Counsel is generated using open-weight models and is permissively licensed for broad community use, which we hope will enable rigorous study and improved alignment of LLM-based evaluators for agentic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。