比较两个智能体系统,发现资源消耗差异大,评估自主AI需结合任务完成与资源使用。
Resource Constraints and Performance in Agentic AI Systems

- 对比OpenClaw与NanoBot在任务执行中的表现和资源消耗
- 两系统全任务完成率分别为31%与25%,但资源开销差距显著
- 强调应记录每项结果的执行过程以准确评估智能体性能
向更自主的AI发展越来越依赖于将语言模型与工具、记忆、状态管理及多步执行相结合的智能体系统。这些机制既影响任务能力也带来操作负担。我们通过配对主基准和更详细的仪器化提示子集,比较了OpenClaw与NanoBot这两个完整智能体系统。在主基准中,OpenClaw的全任务完成率为31%,NanoBot为25%,六个百分点差异在95%任务自助区间(-3至15个百分点)内,未显示任一系统具有统计显著的全完成优势。在仪器化层,两系统均实现26%全完成率,但NanoBot在43%提示中至少部分完成,而OpenClaw仅26%。OpenClaw在83%提示中耗时更长,且每个提示的峰值内存更高,几何平均比值分别为2.98(运行时间)和19.44(峰值内存)。在十组详细提示中,至少一个系统达成部分或全完成的场景下,NanoBot在八组中占优;但在全部23组提示中,其十八个占优案例中有十个是因联合失败而廉价胜出。不同证据层级的结果标签不一致,揭示了为何智能体系统评估必须将能力与资源测量关联到尝试级执行与评分溯源。研究结果表明,评估自主AI进展应基于验证的任务完成、观测的资源使用以及将每项结果与其生成执行相联系的记录。
原文摘要 · Abstract (English)
Progress toward more autonomous AI increasingly depends on agentic systems that combine a language model with tools, memory, state management, and multi-step execution. These mechanisms shape both task capability and operational burden. We compare OpenClaw and NanoBot as complete agentic systems using a paired primary benchmark and a more detailed instrumented subset of paired prompts. In the primary benchmark, the rate of full task completion was 31% for OpenClaw and 25% for NanoBot, a six-percentage-point difference with a 95% task-bootstrap interval from -3 to 15 percentage points, providing no statistically established full-completion advantage for either system. In the instrumented layer, both systems achieved 26% full completion, while NanoBot reached at least partial completion on 43% of prompts compared with 26% for OpenClaw. OpenClaw took longer on 83% of prompts and had a higher recorded peak-memory value on every prompt, with geometric mean ratios of 2.98 for wall time and 19.44 for peak memory. Among the ten detailed-layer prompts on which at least one system achieved partial or full completion, NanoBot weakly dominated on eight; across all 23 prompts, however, ten of its eighteen dominance cases were cheaper joint failures. Outcome labels differ across the two evidence layers, showing why agent-system evaluation should connect capability and resource measurements to attempt-level execution and scoring provenance. These findings show that progress toward more autonomous AI should be evaluated through verified task completion, observed resource use and records linking each result to the execution that produced it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。