arXiv:2606.23654cs.CLcs.SE2026-06被引 1

基于真实工作场景构建企业智能体评测基准,揭示评估复杂性。

EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions

论文配图:EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions
图 1 · 摘自论文原文
  • 从真实企业会话中提取852个可复现任务,还原完整上下文与规则
  • 顶尖模型表现仅0.663分,凸显企业任务评估需多维指标
  • 强调评估应披露工具链、成本、生成质量等细节,不可单一打分

企业智能体正深入工作空间:读取异构文件、调用工具并交付业务成果。我们提出EnterpriseClawBench,一个基于私有真实企业会话构建的智能体评测基准。从大规模工作会话存档出发,该基准生成852个可复现任务,每项任务均配有恢复的环境配置、重写后的提示词、角色类别、技能子类、硬性规则及语义评分标准。由于会话包含内部企业内容,数据不对外公开;我们的核心贡献是构建与评估协议的可复用性。在EnterpriseClawBench上,最佳配置(Codex + GPT-5.5)得分仅为0.663。结果表明,企业智能体评估必须报告工具链-模型组合、成果交付、视觉质量、成本、运行时间及技能迁移行为,而非简化为单一分数。代码已开源。

原文摘要 · Abstract (English)

Enterprise agents increasingly operate inside workspaces: they read heterogeneous files, invoke tools, and deliver business artifacts. We introduce EnterpriseClawBench, an enterprise agent benchmark constructed from proprietary, real-world agent sessions. Starting from a large archive of workplace sessions, the EnterpriseClawBench produces 852 reproducible tasks, each paired with recovered fixtures, rewritten prompts, role classes, skill subclasses, hard rules, and semantic rubrics. Because the sessions contain internal enterprise content, we do not release the benchmark data; instead, our reusable contribution is the construction and evaluation protocol. On EnterpriseClawBench, the best configuration reaches only 0.663 (Codex with GPT-5.5). These results show that enterprise agent evaluation must report harness--model combinations, artifact delivery, visual quality, cost, runtime, and skill-transfer behavior, rather than collapsing performance into a single score. Code: https://github.com/FrontisAI/EnterpriseClawBench

智能体评测企业AI基准测试真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。