用隐式场景监督提升机器人视觉模仿学习的泛化能力
ISS Policy : Scalable Diffusion Policy with Implicit Scene Supervision
- 基于点云输入的扩散策略,通过隐式场景监督优化动作预测
- 在MetaWorld和Adroit数据集上达到当前最佳性能,真实场景表现稳健
- 适用于需要强泛化能力的复杂操作任务,如机械臂与灵巧手操控
基于视觉的模仿学习已实现令人瞩目的机器人操作技能,但其依赖物体外观而忽略底层三维场景结构,导致训练效率低且泛化能力差。为此,我们提出基于3D视觉-运动DiT的扩散策略——隐式场景监督(ISS)策略,从点云观测中预测连续动作序列。通过引入新颖的隐式场景监督模块,使模型输出与场景几何演化保持一致,从而提升策略性能与鲁棒性。ISS策略在单臂操作任务(MetaWorld)和灵巧手操作任务(Adroit)上均达到当前最优水平。真实世界实验表明其具备强泛化能力与鲁棒性。消融研究显示,该方法在数据量与参数规模上均表现出良好可扩展性。代码与视频将公开。
原文摘要 · Abstract (English)
Vision-based imitation learning has enabled impressive robotic manipulation skills, but its reliance on object appearance while ignoring the underlying 3D scene structure leads to low training efficiency and poor generalization. To address these challenges, we introduce \emph{Implicit Scene Supervision (ISS) Policy}, a 3D visuomotor DiT-based diffusion policy that predicts sequences of continuous actions from point cloud observations. We extend DiT with a novel implicit scene supervision module that encourages the model to produce outputs consistent with the scene's geometric evolution, thereby improving the performance and robustness of the policy. Notably, ISS Policy achieves state-of-the-art performance on both single-arm manipulation tasks (MetaWorld) and dexterous hand manipulation (Adroit). In real-world experiments, it also demonstrates strong generalization and robustness. Additional ablation studies show that our method scales effectively with both data and parameters. Code and videos will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。