解决3D策略学习的不稳定性问题,提升泛化与迁移能力
R3D: Revisiting 3D Policy Learning

- 设计基于Transformer的3D编码器与扩散解码器,增强训练稳定性
- 在复杂操作基准上显著超越现有3D方法,实现更鲁棒的模仿学习
- 适合关注3D机器人控制与大规模预训练的科研人员
3D策略学习有望实现更强的泛化能力和跨体态迁移,但训练不稳定性与严重过拟合阻碍了其发展,限制了对强大3D感知模型的应用。本文系统诊断失败原因,发现缺失3D数据增强和批归一化负面影响是主因。提出一种新架构,结合可扩展的Transformer-based 3D编码器与扩散解码器,专为大规模稳定性设计,并支持大规模预训练。该方法在具有挑战性的操控基准上显著优于当前最优3D基线,建立了可扩展3D模仿学习的新基础。
原文摘要 · Abstract (English)
3D policy learning promises superior generalization and cross-embodiment transfer, but progress has been hindered by training instabilities and severe overfitting, precluding the adoption of powerful 3D perception models. In this work, we systematically diagnose these failures, identifying the omission of 3D data augmentation and the adverse effects of Batch Normalization as primary causes. We propose a new architecture coupling a scalable transformer-based 3D encoder with a diffusion decoder, engineered specifically for stability at scale and designed to leverage large-scale pre-training. Our approach significantly outperforms state-of-the-art 3D baselines on challenging manipulation benchmarks, establishing a new and robust foundation for scalable 3D imitation learning. Project Page: https://r3d-policy.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。