arXiv:2609.06059cs.AI2026-09

新基准DAREBench评估智能体在真实部署中的可靠性与效率表现。

DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

论文配图:DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents
图 1 · 摘自论文原文
  • 基于统一执行环境和契约协议,整合233项任务构建多维度评估矩阵。
  • 测试23个商用模型与12个开源模型,发现无模型在所有场景领先。
  • 强调根据任务类型和成本权衡选型,避免单一分数误导决策。

随着大语言模型从问答系统演变为通用智能体,评估需超越静态答案正确性,涵盖多模态感知、多步执行、工具使用及成果交付能力。然而现有基准常绑定特定任务类型、执行环境或评分规则,影响可比性、可解释性与部署可靠性。本文提出DAREBench(Deployment-Aware and Reliable Evaluation of Models as Agents),一个面向部署的可靠智能体评估基准。基于共享OpenClaw执行环境,将22个源基准中的233项任务重构为由输入模态和执行形式构成的$2\times3$工作负载矩阵,并采用统一契约协议与证据驱动评分审计机制进行评估。对23个商用API模型与12个本地部署开源模型开展7,587次模型-任务运行,报告准确率、令牌消耗及商用模型参考成本。结果表明:无单一模型在所有工作负载组中占优;文本与多模态任务呈现显著的准确率-成本权衡;本地开源模型在部分组别具竞争力,但整体仍落后于前沿商用模型。研究提示:智能体部署与模型选择应综合考虑工作负载特征、部署模式及准确率-成本平衡,而非依赖单一总分。

原文摘要 · Abstract (English)

As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception, multi-step execution, tool use, and artifact delivery. However, existing benchmarks are often tied to specific task types, execution environments, or scoring protocols, limiting their comparability, interpretability, and reliability for deployment decisions. We introduce DAREBench (Deployment-Aware and Reliable Evaluation of Models as Agents), a benchmark designed to capture workload variation and support reliable agent evaluation. Built on a shared OpenClaw execution environment, DAREBench organizes 233 tasks selected and adapted from 22 source benchmarks into a $2\times3$ workload matrix defined by input modality and execution form, and evaluates them under a unified contract-based protocol with evidence-based score auditing. We evaluate 23 commercial API models and 12 locally deployed open-weight models over 7,587 model--task runs, reporting accuracy and token consumption alongside reference costs for API models. Results show that no single model dominates all workload groups, text and multimodal tasks exhibit distinct accuracy--cost trade-offs, and local open-weight models are competitive in several groups but still trail frontier commercial models overall. These findings suggest that agent deployment and model selection should consider workload profiles, deployment mode, and accuracy--cost trade-offs rather than rely on a single aggregate score.

智能体评估模型部署基准测试成本权衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。