自动构建可更新的具身空间智能评测基准,减少人工成本。
Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction

- 用三个智能体协作五阶段流程自动生成评测基准
- 支持多种机器人载体和空间推理任务,可持续更新
- 适合研究者快速搭建可靠、可复现的评测体系
评测基准对评估具身空间智能至关重要,但其构建过程耗时费力、难以复用和维护。现有具身基准多为静态,模型进步后易饱和,无法区分新能力。我们提出Embodied-BenchClaw,一个用于构建具身空间智能评测基准的自主代理系统。给定用户指定的评估意图,该系统通过五阶段流程——意图蓝图设计、数据采集、结构化与清洗、基准合成、评估报告生成——自动生成完整且可持续更新的基准包。流程由规划、构建、评估三个智能体协同完成。为提升复用性与可靠性,系统引入可扩展的技能库和过程质量控制机制,使基准构建具备组合性、可验证性和可修复性。我们实例化了涵盖室内空间推理、室外空间推理、机器人操作、四足机器人导航、无人机/航拍理解及静态基准增强的多个基准。这些基准覆盖多样具身载体、数据源和空间能力。人类评估、裁判评分、一致性检查、成本分析及消融实验表明,Embodied-BenchClaw能以更少人工投入,构建可验证、可执行、可维护且具有诊断价值的具身空间智能基准。
原文摘要 · Abstract (English)
Benchmarks are essential for evaluating embodied spatial intelligence, yet their construction is labor-intensive, hard to reuse, and difficult to maintain. Existing embodied benchmarks are often static and may quickly become saturated as models improve, limiting their ability to distinguish new capabilities. We propose Embodied-BenchClaw, an autonomous agentic system for constructing embodied spatial intelligence benchmarks. Given a user-specified evaluation intent, Embodied-BenchClaw automatically produces a complete and continually updatable benchmark package through a five-stage pipeline: intent blueprinting, data collection, structuring and cleaning, benchmark synthesis, and evaluation reporting. The pipeline is coordinated by three agents for planning, construction, and evaluation. To improve reusability and reliability, Embodied-BenchClaw introduces an extensible Skill Library and process quality control, enabling benchmark construction to be composable, verifiable, and repairable. We instantiate multiple benchmarks covering indoor spatial reasoning, outdoor spatial reasoning, robotic manipulation, quadruped robot navigation, UAV/aerial-view understanding, and static benchmark enhancement. These benchmarks span diverse embodied carriers, data sources, and spatial capabilities. Experiments with human evaluation, judge-based assessment, consistency checks, cost analysis, and ablations show that Embodied-BenchClaw can construct verifiable, executable, maintainable, and diagnostically useful embodied spatial benchmarks with reduced manual effort.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。