实时动态评估大模型在真实工作流中的执行能力。
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows

- 构建可更新的实时工作流基准,分离需求信号与固定任务快照。
- 13个前沿模型仅66.7%任务通过,多系统业务流程仍是难点。
- 适合关注真实场景自动化、评估模型执行可信度的研究者。
大语言模型代理需跨软件工具、商业服务和本地工作区完成端到端任务,但现有基准常冻结任务集并仅评价最终输出,难以评估代理对动态工作流需求的响应能力。本文提出Claw-Eval-Live,一个实时工作流代理基准,将可刷新的需求信号层与可复现的时间戳发布快照分离。每轮发布基于公开的工作流需求信号,结合ClawHub Top-500技能,生成具有固定环境、服务、工作区和评分器的受控任务。评分时记录执行轨迹、审计日志、服务状态和运行后工作区产物,证据充分时采用确定性检查,语义维度则使用结构化LLM判断。当前版本含105项任务,涵盖受控商业服务与本地工作区修复,评估13个前沿模型。实验显示,可靠工作流自动化仍远未解决:领先模型仅通过66.7%任务,无一模型达70%。失败集中在人力资源、管理及多系统业务流程,本地修复虽较易但未饱和。排行榜排名不足以为据,相似通过率下完成度差异显著,任务级区分集中于中等难度任务。结果表明,工作流代理评估应同时基于最新外部需求与可验证的代理行为。
原文摘要 · Abstract (English)
LLM agents are expected to complete end-to-end units of work across software tools, business services, and local workspaces. Yet many agent benchmarks freeze a curated task set at release time and grade mainly the final response, making it difficult to evaluate agents against evolving workflow demand or verify whether a task was executed. We introduce Claw-Eval-Live, a live benchmark for workflow agents that separates a refreshable signal layer, updated across releases from public workflow-demand signals, from a reproducible, time-stamped release snapshot. Each release is constructed from public workflow-demand signals, with ClawHub Top-500 skills used in the current release, and materialized as controlled tasks with fixed fixtures, services, workspaces, and graders. For grading, Claw-Eval-Live records execution traces, audit logs, service state, and post-run workspace artifacts, using deterministic checks when evidence is sufficient and structured LLM judging only for semantic dimensions. The release contains 105 tasks spanning controlled business services and local workspace repair, and evaluates 13 frontier models under a shared public pass rule. Experiments reveal that reliable workflow automation remains far from solved: the leading model passes only 66.7% of tasks and no model reaches 70%. Failures are structured by task family and execution surface, with HR, management, and multi-system business workflows as persistent bottlenecks and local workspace repair comparatively easier but unsaturated. Leaderboard rank alone is insufficient because models with similar pass rates can diverge in overall completion, and task-level discrimination concentrates in a middle band of tasks. Claw-Eval-Live suggests that workflow-agent evaluation should be grounded twice, in fresh external demand and in verifiable agent action.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。