用自动化框架评测真实场景下的智能代理表现
STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios

- 基于任务提示自动生成真实环境与验证程序
- 40个任务评估11个模型,以系统状态正确性为准
- 适合研究智能代理真实表现与可靠性的人
大语言模型正被用于构建日常应用的个人代理,但其评估仍面临挑战。现有基准依赖沙箱环境、静态任务设计和粗粒度评分,限制了可扩展性并阻碍可靠评估进展。本文提出STAGE-Claw,一个自动化框架,可在基于状态的个人计算环境中构建并评估真实个人代理场景。给定任务提示后,该框架自动创建并验证包含环境、任务提示、真值及验证程序的基准任务。代理在真实操作系统环境中运行,性能通过最终系统状态的正确性衡量,而非仅文本响应。本文利用STAGE-Claw构建了40个具有挑战性的真实场景任务,评估了11个前沿模型,分析其任务得分、成本、工具调用可靠性及常见失败模式。整体上,STAGE-Claw为真实用户场景中的代理评估提供了可扩展、基于状态的解决方案。
原文摘要 · Abstract (English)
Large language models are increasingly used to power personal agents for everyday applications, but evaluating these agents remains a challenge. Existing benchmarks still rely on sandboxed artifacts, static task design, and coarse scoring, which hinder scalability and limit progress toward reliable personal-agent evaluation. This paper introduces STAGE-Claw, an automated framework for building and evaluating realistic personal-agent scenarios in state-based personal-computing environments. Given a task hint, STAGE-Claw automatically creates and validates a realistic benchmark task with its environment, task prompts, ground truth, and related verification programs. Agents are then evaluated in realistic operating environments, where performance is measured by the correctness of the final system state rather than only the textual response. Using STAGE-Claw, this paper creates a benchmark with 40 challenging real scenario agent tasks, evaluates 11 frontier models, and analyzes their task scores, costs, tool-call reliability, and common failure patterns. Overall, STAGE-Claw offers a scalable, state-based way to evaluate agents in realistic user scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。