用噪声依赖策略让机器人从低质数据中高效学习,提升数据利用率。
Ambient Diffusion Policy: Imitation Learning from Suboptimal Data in Robotics

- 通过噪声依赖机制只在特定扩散阶段使用低质数据
- 在六个任务上提升33%性能,优于现有方法
- 适合处理真实场景中混杂的低质量机器人数据
我们提出环境扩散策略(Ambient Diffusion Policy),一种从机器人低质数据中进行模仿学习的简单且原理清晰的方法。高质量、任务特定的机器人数据采集成本高,而低质量或分布外的演示数据却大量存在。现有方法在同时利用两类数据时,难以区分有效与有害特征。我们的方法通过引入噪声依赖的数据使用机制,仅在高噪声和低噪声扩散阶段启用低质数据。我们首先观察到机器人动作数据呈现谱功率律,由此推导出最优扩散策略具有全局-局部层级结构和局部性,理论验证了该设计。实验在四类低质动作数据(含噪声轨迹、仿真到现实差距、任务不匹配、大规模数据混合)上验证,覆盖六个任务。结果表明,该方法能有效从任意来源的低质数据中学习。尤其在大型异构数据集Open X-Embodiment上,相比现有联合训练基线最高提升33%。整体上,该方法显著提升了低质演示的可用性,拓展了机器人可利用的数据源。
原文摘要 · Abstract (English)
We propose Ambient Diffusion Policy, a simple and principled method for imitation learning from suboptimal data in robotics. High-quality, task-specific robot data is expensive and time-consuming to collect, while suboptimal datasets with lower-quality or out-of-distribution demonstrations are abundant. Existing methods that co-train on both data sources in robotics often fail to separate the meaningful and the harmful features in the suboptimal samples. In contrast, our method extracts only the useful features by introducing a new axis to co-training in robotics: noise-dependent data usage. Ambient Diffusion Policy restricts the contribution of suboptimal data during training to only the high and low diffusion times. To rigorously justify our approach, we first observe that robot action data exhibits a spectral power law. This induces two important properties on the optimal Diffusion Policy that we exploit: a global-to-local hierarchy and locality. We theoretically formalize this discussion using a simplified model. Our experiments validate Ambient Diffusion Policy on four types of suboptimal action data (noisy trajectories, sim-to-real gap, task mismatch, and large-scale data mixtures) across six tasks. The results show that it effectively learns from arbitrary sources of suboptimal data. Notably, it outperforms existing co-training baselines by up to 33% when scaled to Open X-Embodiment - a large dataset with heterogeneous data quality and unstructured distribution shifts. Overall, Ambient Diffusion Policy increases the utility of suboptimal demonstrations and expands the set of usable data sources in robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。