arXiv:2412.05789cs.RO2024-12被引 23

构建统一可扩展的视觉语言机器人仿真框架,解决资产分散问题

InfiniteWorld: A Unified Scalable Simulation Framework for General Visual-Language Robot Interaction

  • 基于Nvidia Isaac Sim打造统一仿真平台,整合3D资产生成与标注流程
  • 推出四个新基准,涵盖场景图协作探索与开放世界社交移动操作任务
  • 适合机器人交互、具身智能与多模态学习研究者使用

实现具身AI中的缩放定律已成为研究重点。然而,以往工作分散在多个仿真平台中,资产与模型缺乏统一接口,导致研究效率低下。为此,我们提出InfiniteWorld,一个基于Nvidia Isaac Sim的统一且可扩展的通用视觉-语言机器人交互仿真框架。该框架集成了一系列生成驱动的3D资产构建方法、Real2Sim、自动化标注框架及统一3D资产处理流程,提供一体化的机器人交互与学习平台。此外,为模拟真实机器人交互,我们构建了四个新基准,包括场景图协作探索和开放世界社交移动操作。前者关注机器人探索环境并构建场景知识这一常被忽视的任务,后者在此基础上模拟与不同知识水平智能体的交互。这些基准能更全面评估具身智能体在环境理解、任务规划与执行及智能交互方面的能力。我们希望本工作能为社区提供系统化的资产接口,缓解高质量资产匮乏问题,并提供更全面的机器人交互评估体系。

原文摘要 · Abstract (English)

Realizing scaling laws in embodied AI has become a focus. However, previous work has been scattered across diverse simulation platforms, with assets and models lacking unified interfaces, which has led to inefficiencies in research. To address this, we introduce InfiniteWorld, a unified and scalable simulator for general vision-language robot interaction built on Nvidia Isaac Sim. InfiniteWorld encompasses a comprehensive set of physics asset construction methods and generalized free robot interaction benchmarks. Specifically, we first built a unified and scalable simulation framework for embodied learning that integrates a series of improvements in generation-driven 3D asset construction, Real2Sim, automated annotation framework, and unified 3D asset processing. This framework provides a unified and scalable platform for robot interaction and learning. In addition, to simulate realistic robot interaction, we build four new general benchmarks, including scene graph collaborative exploration and open-world social mobile manipulation. The former is often overlooked as an important task for robots to explore the environment and build scene knowledge, while the latter simulates robot interaction tasks with different levels of knowledge agents based on the former. They can more comprehensively evaluate the embodied agent's capabilities in environmental understanding, task planning and execution, and intelligent interaction. We hope that this work can provide the community with a systematic asset interface, alleviate the dilemma of the lack of high-quality assets, and provide a more comprehensive evaluation of robot interactions.

机器人仿真视觉语言具身智能基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。