用真实开发者与代理的对话构建可复现的评测基准,更贴近实际使用场景。
RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions

- 从真实OpenClaw会话中提取任务,重建执行环境并设计可验证评分器。
- 包含281个可执行任务,最佳模型仅解决65.8%的任务,仍有巨大提升空间。
- 适合评估代理在真实开发场景中的实际能力,尤其关注环境依赖和意图模糊问题。
代理评测应反映用户实际向部署的代理提出的需求,但现有评测常缺乏真实开发者-代理会话的关键现实特征。我们提出RealClawBench,一个基于真实OpenClaw会话的实时评测框架,以捕捉部署代理使用的分布、多样性与真实难度。真实用户请求难以评测,因其常依赖本地执行环境、包含隐含或未明确的意图,且需非平凡验证。RealClawBench通过两个核心机制应对:重构的执行环境与确定性可验证评分器,将真实会话转化为可复现、自动评分的任务。最终发布包含281个可执行任务,源自更大规模的真实会话池,同时保留原始分布,最大最终-源分布的Jensen-Shannon散度为0.0448。对14个当代模型的评估显示,最佳系统仅解决65.8%的任务,揭示了在真实开发者-代理工作负载上的显著提升空间。通过将真实部署会话转化为受控评测实例,RealClawBench为构建更能反映实际使用能力的评测提供了可行路径。代码已公开于:https://anonymous.4open.science/r/real-claw-bench-582B。
原文摘要 · Abstract (English)
Agent benchmarks should reflect what users actually ask deployed agents to do, yet existing benchmarks often miss key realism properties of real developer-agent sessions. We introduce RealClawBench, a live benchmark framework built from real OpenClaw sessions to capture the distribution, diversity, and real-world difficulty of deployed agent use. Real user requests are challenging to benchmark because they often depend on local execution environments, involve implicit or underspecified intent, and require nontrivial verification. RealClawBench addresses these challenges with two core mechanisms: reconstructed execution environments and deterministic verifiable scorers, which together convert real sessions into reproducible, automatically scored tasks. The resulting release contains 281 executable tasks sampled from a much larger real-session pool while preserving the source distribution, with maximum final-vs-source Jensen-Shannon divergence of 0.0448. Evaluating 14 contemporary models shows that the best system solves only 65.8% of tasks, revealing substantial headroom on realistic developer-agent workloads. By turning real deployed sessions into controlled evaluation instances, RealClawBench provides a practical path toward benchmarks that better measure agent capability in actual use. Code is available at:https://anonymous.4open.science/r/real-claw-bench-582B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。