arXiv:2508.05508cs.AI2025-08被引 7

提出通用智能体评估框架,更贴近人类判断任务完成度。

Auto-Eval Judge: Towards a General Agentic Framework for Task Completion Evaluation

  • 将任务拆解为步骤,逐项验证推理过程与输出
  • 在两个基准上比现有方法提升10%以上评估一致性
  • 适合需要细致过程评估的复杂任务研究者

基础模型作为智能体在多领域广泛应用,亟需可靠的评估框架。当前方法如大模型作裁判仅关注最终输出,忽略智能体决策中的逐步推理;而现有智能体作裁判系统多局限于特定领域。为此,我们提出一种可泛化、模块化的智能体任务完成评估框架,模拟人类评价方式:将任务分解为子任务,利用智能体输出和推理信息逐项验证,各模块贡献评估不同方面,结果聚合生成最终结论。我们在GAIA和BigCodeBench两个基准上评估Magentic-One Actor Agent,结果显示,该裁判智能体与人工评价的一致性更高,分别较GPT-4o基线提升4.76%和10.52%的对齐准确率,证明了所提通用评估框架的有效性。

原文摘要 · Abstract (English)

The increasing adoption of foundation models as agents across diverse domains necessitates a robust evaluation framework. Current methods, such as LLM-as-a-Judge, focus only on final outputs, overlooking the step-by-step reasoning that drives agentic decision-making. Meanwhile, existing Agent-as-a-Judge systems, where one agent evaluates another's task completion, are typically designed for narrow, domain-specific settings. To address this gap, we propose a generalizable, modular framework for evaluating agent task completion independent of the task domain. The framework emulates human-like evaluation by decomposing tasks into sub-tasks and validating each step using available information, such as the agent's output and reasoning. Each module contributes to a specific aspect of the evaluation process, and their outputs are aggregated to produce a final verdict on task completion. We validate our framework by evaluating the Magentic-One Actor Agent on two benchmarks, GAIA and BigCodeBench. Our Judge Agent predicts task success with closer agreement to human evaluations, achieving 4.76% and 10.52% higher alignment accuracy, respectively, compared to the GPT-4o based LLM-as-a-Judge baseline. This demonstrates the potential of our proposed general-purpose evaluation framework.

智能体评估任务完成度大模型评测自动化判断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。