构建移动操作基准,加速机器人模型验证与优化
MobileManiBench: Simplifying Model Verification for Mobile Manipulation
- 模拟先行生成多样化操作轨迹,支持可控验证
- 涵盖300K条轨迹,覆盖100个真实场景和630类物体
- 适合研究数据效率、泛化能力及多模态感知的学者
视觉-语言-动作模型虽推动了机器人操作发展,但仍受限于依赖大量由远程操控收集的数据集,且主要集中于静态桌面场景。我们提出一种以模拟为核心的验证框架,并引入MobileManiBench——一个面向移动机器人的大规模操作基准。基于NVIDIA Isaac Sim并结合强化学习,该流程可自动生成包含丰富标注(语言指令、多视角RGB-D分割图像、同步的物体/机器人状态与动作)的多样化操作轨迹。MobileManiBench包含2种移动平台(平行夹爪与灵巧手机器人)、2个同步摄像头(头部与右腕)、630个物体(20类)、5项基础技能(开、关、拉、推、拾取),在100个真实场景中完成超100项任务,生成300K条轨迹。该设计支持对机器人本体、感知模态与策略架构的可控、可扩展研究,加速数据效率与泛化性相关研究。我们对代表性VLA模型进行了基准测试,揭示了复杂模拟环境中感知、推理与控制的表现规律,所有代码、数据集与模型均已开源。
原文摘要 · Abstract (English)
Vision-language-action models have advanced robotic manipulation but remain constrained by reliance on the large, teleoperation-collected datasets dominated by the static, tabletop scenes. We propose a simulation-first framework to verify VLA architectures before real-world deployment and introduce MobileManiBench, a large-scale benchmark for mobile-based robotic manipulation. Built on NVIDIA Isaac Sim and powered by reinforcement learning, our pipeline autonomously generates diverse manipulation trajectories with rich annotations (language instructions, multi-view RGB-depth-segmentation images, synchronized object/robot states and actions). MobileManiBench features 2 mobile platforms (parallel-gripper and dexterous-hand robots), 2 synchronized cameras (head and right wrist), 630 objects in 20 categories, 5 skills (open, close, pull, push, pick) with over 100 tasks performed in 100 realistic scenes, yielding 300K trajectories. This design enables controlled, scalable studies of robot embodiments, sensing modalities, and policy architectures, accelerating research on data efficiency and generalization. We benchmark representative VLA models and report insights into perception, reasoning, and control in complex simulated environments, with all code, datasets, and models publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。