arXiv:2605.19099cs.AIcs.CL2026-05被引 2

构建评估长时程智能体任务委派的基准,揭示当前系统委派能力仍有巨大提升空间。

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows

论文配图:DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
图 1 · 摘自论文原文
  • 设计统一基准框架,支持多种模型与委派策略对比评测。
  • 实测委派准确率仅7.5%至29.5%,远低于理论上限15-31个百分点。
  • 适合研究智能体协作、任务调度与自动化流程优化的研究者使用。

我们提出DecisionBench,一个用于长时程智能体工作流中涌现式委派的基准框架。该框架固定了任务集合(GAIA、tau-bench、BFCL多轮)、模型池(11个模型,7个供应商系列)、委派接口(call_model + 可选read_profile通道)、确定性技能标注层及多维度评估指标,涵盖质量、成本、延迟、委派率、路由保真度-at-k、供应商自偏好和反事实委派上限。该框架对同伴信息生成方式无偏,支持学习型路由器、增强记忆、自适应配置构建及多步委派等方法的评估。通过在全模型池上进行五条件基准测试(n=23,375任务实例),发现:(i) 四种意识条件下任务平均质量无统计差异(|beta| ≤ 0.010,p ≥ 0.21),仅以质量评估会忽略编排信号;(ii) 路由保真度-at-1在不同条件下为7.5%至29.5%,且交付通道(按需工具调用 vs. 预加载描述)显著影响表现;(iii) 反事实上限表明,理想委派性能比当前高出15-31个百分点,揭示未来编排方法的巨大潜力。我们开源完整框架、标注层、参考干预集、分析管道及220个条件运行存档。

原文摘要 · Abstract (English)

We introduce DecisionBench, a benchmark substrate for emergent delegation in long-horizon agentic workflows. The substrate fixes a task suite (GAIA, tau-bench, BFCL multi-turn), a peer-model pool (11 models, 7 vendor families), a delegation interface (call_model plus an optional read_profile channel), a deterministic skill-annotation layer, and a multi-axis metric suite covering quality, cost, latency, delegation rate, routing fidelity-at-k, vendor self-preference, and a counterfactual-delegation ceiling. The substrate is agnostic to how peer information is generated or delivered, so learned routers, richer peer memories, adaptive profile construction, and multi-step delegation can all be evaluated against it. We characterize the substrate with a five-condition reference sweep on the full pool (n=23,375 task instances). Three benchmark-level findings emerge: (i) mean end-task quality is statistically indistinguishable across the four awareness conditions (|beta| <= 0.010, p >= 0.21), so quality-only evaluation would miss the orchestration signal; (ii) routing fidelity-at-1 ranges from 7.5% to 29.5% across conditions at near-equal mean quality, with delivery channel (on-demand tool vs. preloaded description) dominating description content; (iii) a counterfactual ceiling places perfect delegation 15-31 percentage points above measured performance on every suite, locating large unrealized headroom for future orchestration methods. We release the substrate, annotation layer, reference intervention suite, analysis pipeline, and 220 per-condition run archives.

智能体任务委派评估基准长时程任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。