构建23万+虚拟场景,支撑机器人导航与操作的大规模评测。
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
- 用23万+多样室内环境和13万+带标注物体构建开放生态
- 实测显示仿真到现实的迁移相关性高达0.98
- 适合做机器人泛化能力研究与零样本策略评估
大规模部署机器人需应对日常场景中无穷无尽的变化。现有机器人基准难以覆盖真实环境中的场景布局、物体几何与任务要求的多样性。衡量这种泛化能力需要超越物理测试规模与多样性。我们提出MolmoSpaces,一个完全开源的生态系统,支持机器人策略的大规模基准评测。该系统包含超过23万种多样的室内环境,涵盖手工设计的家居场景与程序生成的多房间住宅,配有13万多个丰富标注的物体资产,包括48,000个可操作物体及4200万次稳定抓取。关键的是,这些环境与模拟器无关,兼容MuJoCo、Isaac和ManiSkill等主流平台。生态系统支持完整体化任务:静态与移动操作、导航,以及跨整个室内环境的长时程多任务协调感知、规划与交互。我们还设计了MolmoSpaces-Bench,包含8项任务的基准套件,机器人在多样化场景与丰富标注物体中进行交互。实验表明,该基准具有强仿真到现实的相关性(R = 0.96,ρ = 0.98),验证了新式更强的零样本策略优于旧版本,并揭示了提示词、初始关节位置与相机遮挡的关键敏感性。通过MolmoSpaces及其开源资产与工具链,我们为机器人学习研究提供可扩展的数据生成、策略训练与基准创建基础。
原文摘要 · Abstract (English)
Deploying robots at scale demands robustness to the long tail of everyday situations. The countless variations in scene layout, object geometry, and task specifications that characterize real environments are vast and underrepresented in existing robot benchmarks. Measuring this level of generalization requires infrastructure at a scale and diversity that physical evaluation alone cannot provide. We introduce MolmoSpaces, a fully open ecosystem to support large-scale benchmarking of robot policies. MolmoSpaces consists of over 230k diverse indoor environments, ranging from handcrafted household scenes to procedurally generated multiroom houses, populated with 130k richly annotated object assets, including 48k manipulable objects with 42M stable grasps. Crucially, these environments are simulator-agnostic, supporting popular options such as MuJoCo, Isaac, and ManiSkill. The ecosystem supports the full spectrum of embodied tasks: static and mobile manipulation, navigation, and multiroom long-horizon tasks requiring coordinated perception, planning, and interaction across entire indoor environments. We also design MolmoSpaces-Bench, a benchmark suite of 8 tasks in which robots interact with our diverse scenes and richly annotated objects. Our experiments show MolmoSpaces-Bench exhibits strong sim-to-real correlation (R = 0.96, \r{ho} = 0.98), confirm newer and stronger zero-shot policies outperform earlier versions in our benchmarks, and identify key sensitivities to prompt phrasing, initial joint positions, and camera occlusion. Through MolmoSpaces and its open-source assets and tooling, we provide a foundation for scalable data generation, policy training, and benchmark creation for robot learning research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。