用扩散模型统一控制机器人,让语言指令直接生成动作。
Pixel Motion Diffusion is What We Need for Robot Control
- 用扩散模型建模高阶意图与低阶动作,中间通过像素运动表示衔接。
- 在CALVIN和MetaWorld上表现顶尖,真实世界仅需微调即可迁移。
- 适合想用扩散模型做机器人控制的研究者和工程师。
我们提出DAWN(Diffusion is All We Need for robot control),一个基于扩散模型的统一框架,用于语言引导的机器人操作任务。该框架通过结构化的像素运动表示,连接高层运动意图与底层机器人动作。在DAWN中,高层与低层控制器均以扩散过程建模,实现全可训练、端到端的系统,并具备可解释的中间运动抽象。DAWN在具有挑战性的CALVIN基准上达到当前最优性能,展现出强大的多任务能力,并在MetaWorld上进一步验证有效性。尽管存在仿真与现实之间的显著领域差异且真实数据有限,我们仍证明其在真实世界中仅需极少微调即可可靠迁移,展示了基于扩散模型的动作抽象在机器人控制中的实际可行性。结果表明,将扩散建模与以运动为中心的表示结合,是实现可扩展、鲁棒机器人学习的强大基线。
原文摘要 · Abstract (English)
We present DAWN (Diffusion is All We Need for robot control), a unified diffusion-based framework for language-conditioned robotic manipulation that bridges high-level motion intent and low-level robot action via structured pixel motion representation. In DAWN, both the high-level and low-level controllers are modeled as diffusion processes, yielding a fully trainable, end-to-end system with interpretable intermediate motion abstractions. DAWN achieves state-of-the-art results on the challenging CALVIN benchmark, demonstrating strong multi-task performance, and further validates its effectiveness on MetaWorld. Despite the substantial domain gap between simulation and reality and limited real-world data, we demonstrate reliable real-world transfer with only minimal finetuning, illustrating the practical viability of diffusion-based motion abstractions for robotic control. Our results show the effectiveness of combining diffusion modeling with motion-centric representations as a strong baseline for scalable and robust robot learning. Project page: https://eronguyen.github.io/DAWN/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。