构建真实机器人评估框架,验证通用机械臂的推理与操作能力。
ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation
- 设计标准化真实场景实验框架,包含20个任务和1.08万条专家轨迹。
- 实测发现模型表现受训练数据、微调方式等影响显著,非仅由架构决定。
- 适合研究机器人泛化能力、失败原因分析的科研人员使用。
视觉-语言-动作(VLA)模型与世界动作模型已成为通用机器人智能的核心范式,但其实际进展受限于缺乏兼具物理真实性和诊断可控性的评估协议。仿真基准虽具规模与可复现性,却无法反映感知噪声、接触动力学、延迟、校准误差及硬件限制带来的现实差距。而真实机器人评估往往分散在不同平台、场景、物体和评分规则之间,难以公平比较与归因失败。我们提出ManipArena,一个标准化的真实机器人评估框架,用于在一致物理条件下研究操作泛化。该框架包含20个任务、10,812条专家轨迹、1350万帧图像,以及约188小时机器人运行时间,覆盖桌面与移动操作。结合任务模板变化、分层内域、视觉偏移与语义分布外测试、子任务部分计分、三级语言标注、低层级运动信号,以及从物理场景重建的成对真实-仿真环境。基于ManipArena,我们评估了七种桌面配置下的VLA与世界动作模型策略。结果表明,真实机器人结论不仅依赖模型架构,还受模型来源、微调策略、数据采样方式与标注粒度影响。ManipArena为诊断具身泛化的能力边界与失效模式提供了可复现且可解释的基础。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models and world-action models have emerged as central paradigms for general-purpose robotic intelligence, yet their empirical progress remains constrained by the absence of evaluation protocols that are both physically realistic and diagnostically controlled. Simulator-centric benchmarks provide scale and reproducibility, but cannot fully capture the reality gap induced by perception noise, contact dynamics, latency, calibration error, and hardware constraints. Conversely, real-robot evaluations are often fragmented across platforms, scenes, objects, and scoring rules, making fair comparison and failure attribution difficult. We introduce ManipArena, a standardized real-robot evaluation framework for studying manipulation generalization under matched physical conditions. ManipArena comprises 20 tasks, 10,812 expert trajectories, 13.5M frames, and approximately 188 robot hours across tabletop and mobile manipulation. The framework combines schema-defined task variation, stratified in-domain, visualshift, and semantic-OOD trials, subtask-level partial-credit scoring, three-level language annotations, low-level motor signals, and paired real-to-sim environments reconstructed from physical scenes. Using ManipArena, we evaluate seven tabletop configurations spanning VLA and world-action-model policies. The results show that real-robot conclusions depend not only on architecture, but also on model provenance, fine-tuning regime, data sampling, and annotation granularity. ManipArena thus provides a reproducible and interpretable foundation for diagnosing capability boundaries and failure modes in embodied generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。