arXiv:2410.10088cs.ROcs.AI2024-10ICRA被引 86

提出新架构让机器人用扩散模型高效完成复杂操作,免调参。

The Ingredients for Robotic Diffusion Transformers

  • 设计可扩展的扩散变换器架构,适配多任务机器人
  • 在双臂ALOHA上实现1500+步长任务,性能超越现有方法
  • 10小时多模态语言标注数据训练,提升模型泛化能力

近年来,机器人学家借助高容量Transformer架构和生成式扩散模型,在灵巧机器人硬件上实现了越来越通用的任务求解。然而,将这两项独立进展结合却出人意料地困难,因为缺乏清晰、明确的设计准则。本文系统研究并优化了高容量扩散变换器策略的关键架构决策。所提出的模型可在多种机器人平台上高效解决多样化任务,且无需针对每套设备反复调参。通过整合研究发现与改进组件,我们提出新型架构 extit{DIT-Policy},在双臂ALOHA机器人上显著优于现有最先进方法,成功解决长达1500+时间步的灵巧操作任务。此外,使用10小时高度多模态、带语言标注的ALOHA示范数据训练时,策略展现出更优的可扩展性。我们希望该工作能推动未来机器人学习技术的发展,融合生成式扩散建模的效率与大规模Transformer架构的可扩展性。代码、机器人数据集与视频见:https://dit-policy.github.io

原文摘要 · Abstract (English)

In recent years roboticists have achieved remarkable progress in solving increasingly general tasks on dexterous robotic hardware by leveraging high capacity Transformer network architectures and generative diffusion models. Unfortunately, combining these two orthogonal improvements has proven surprisingly difficult, since there is no clear and well-understood process for making important design choices. In this paper, we identify, study and improve key architectural design decisions for high-capacity diffusion transformer policies. The resulting models can efficiently solve diverse tasks on multiple robot embodiments, without the excruciating pain of per-setup hyper-parameter tuning. By combining the results of our investigation with our improved model components, we are able to present a novel architecture, named \method, that significantly outperforms the state of the art in solving long-horizon ($1500+$ time-steps) dexterous tasks on a bi-manual ALOHA robot. In addition, we find that our policies show improved scaling performance when trained on 10 hours of highly multi-modal, language annotated ALOHA demonstration data. We hope this work will open the door for future robot learning techniques that leverage the efficiency of generative diffusion modeling with the scalability of large scale transformer architectures. Code, robot dataset, and videos are available at: https://dit-policy.github.io

机器人扩散模型Transformer生成式策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。