arXiv:2604.06111cs.AIcs.CL2026-04被引 3

新基准可精准调控任务难度和时长,让智能体评估更高效可靠。

AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments

  • 用网格规划任务实现可调时长与难度控制
  • 13个模型测试显示性能差异显著,评估结果可解释
  • 轻量环境设计使评估快且可复现,适合训练验证

现有智能体评估基准存在两大缺陷:环境交互开销高(占总评估时间41%),任务时长与难度分布不均导致整体评分不可靠。为此,我们提出AgentCE-Bench,基于统一的网格规划任务,要求智能体在部分完成的计划中填充隐藏槽位,需满足局部与全局约束。该基准通过两个正交维度实现细粒度控制:可扩展时长(由隐藏槽数H决定)与可控难度(由误导性干扰项数量B决定)。关键创新在于所有工具调用均通过静态JSON文件处理,采用轻量环境设计,消除部署开销,支持快速、可复现的评估,适用于训练阶段验证。我们验证了H与B能有效控制任务时长与难度,且该基准具备强领域一致性和模型区分能力。在6个领域对13个不同规模与架构的模型进行综合实验,揭示显著的跨模型性能差异,证实AgentCE-Bench能提供可解释、可控的智能体推理评估。

原文摘要 · Abstract (English)

Existing Agent benchmarks suffer from two critical limitations: high environment interaction overhead (up to 41\% of total evaluation time) and imbalanced task horizon and difficulty distributions that make aggregate scores unreliable. To address these issues, we propose AgentCE-Bench built around a unified grid-based planning task, where agents must fill hidden slots in a partially completed schedule subject to both local slot constraints and global constraints. Our benchmark offers fine-grained control through two orthogonal axes: \textbf{Scalable Horizons}, controlled by the number of hidden slots $H$, and \textbf{Controllable Difficulty}, governed by a decoy budget $B$ that determines the number of globally misleading decoy candidates. Crucially, all tool calls are resolved via static JSON files under a \textbf{Lightweight Environment} design, eliminating setup overhead and enabling fast, reproducible evaluation suitable for training-time validation. We first validate that $H$ and $B$ provide reliable control over task horizon and difficulty, and that AgentCE-Bench exhibits strong domain consistency and model discriminability. We then conduct comprehensive experiments across 13 models of diverse sizes and families over 6 domains, revealing significant cross-model performance variation and confirming that AgentCE-Bench provides interpretable and controllable evaluation of agent reasoning.

智能体评估轻量环境可调难度基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。