arXiv:2503.06814cs.ROcs.AI2025-03

通过模块化与大规模学习结合,打造能零样本完成复杂操作的通用机器人。

Unlocking Generalization for Robotics via Modularity and Scale

  • 用规划强制模块化结构,避免端到端学习复杂性。
  • 利用经典规划器在仿真中生成大规模监督数据,提升泛化能力。
  • 整合规划、控制、场景生成,实现从仿真到现实的零样本迁移。

如何构建通用机器人系统?仅靠规模不足,因机器人任务具有显著多模态性、数据获取困难及物理部署挑战。然而,当前多数部署系统本质模块化,各模块可独立泛化。因此,本论文提出将模块化与大规模学习结合,构建通用机器人智能体。首先,如何引入模块化与层次结构?关键洞察是:不依赖代理端到端学习层次与底层控制,而是通过规划强制模块化,提升学习效率与能力。其次,规模在构建通用机器人系统中的作用?神经网络需海量多样化数据、表达力强的架构及监督信号。我们利用经典规划作为强大监督源,虽其泛化能力强但运行成本高且需特权信息。我们在仿真中用规划器监督大规模策略学习,训练出通用智能体。最后,如何统一模块化与大规模策略学习以构建可执行零样本操作的真实机器人系统?通过紧密整合模块化高层/中层规划、学习的局部控制、过程式场景生成及大规模策略学习,实现从仿真到现实的迁移。实验证明,该方法可生成单一通用智能体,在真实世界解决复杂长时程操纵任务。

原文摘要 · Abstract (English)

How can we build generalist robot systems? Scale may not be enough due to the significant multimodality of robotics tasks, lack of easily accessible data and the challenges of deploying on physical hardware. Meanwhile, most deployed robotic systems today are inherently modular and can leverage the independent generalization capabilities of each module to perform well. Therefore, this thesis seeks to tackle the task of building generalist robot agents by integrating these components into one: combining modularity with large-scale learning for general purpose robot control. The first question we consider is: how can we build modularity and hierarchy into learning systems? Our key insight is that rather than having the agent learn hierarchy and low-level control end-to-end, we can enforce modularity via planning to enable more efficient and capable robot learners. Next, we come to the role of scale in building generalist robot systems. To scale, neural networks require vast amounts of diverse data, expressive architectures to fit the data and a source of supervision to generate the data. We leverage a powerful supervision source: classical planning, which can generalize, but is expensive to run and requires access to privileged information to perform well in practice. We use these planners to supervise large-scale policy learning in simulation to produce generalist agents. Finally, we consider how to unify modularity with large-scale policy learning to build real-world robot systems capable of performing zero-shot manipulation. We do so by tightly integrating key ingredients of modular high and mid-level planning, learned local control, procedural scene generation and large-scale policy learning for sim2real transfer. We demonstrate that this recipe can produce a single, generalist agent that can solve challenging long-horizon manipulation tasks in the real world.

机器人模块化泛化仿真到现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。