无需训练即可快速模仿,直接用数据集生成动作
Training-Free Imitation Learning with Closed-Form Diffusion Policies

- 从示范数据推导出闭式梯度,实现零训练策略生成
- 在移动端CPU上毫秒级响应,推理速度超越神经网络扩散模型
- 可对预训练模型进行实时编辑,支持新演示增强与策略引导
尽管基于扩散的策略表现优异且表达能力强,但其漫长的离线训练过程制约了数据采集与策略部署效率。我们提出闭式扩散策略(Closed-Form Diffusion Policies, CFDP),一种无需训练的模仿学习方法,通过从示范数据集中推导闭式得分函数实现策略生成。在硬件实验中,CFDP可在移动CPU上实现毫秒级实时推理,性能优于依赖数小时训练的神经扩散策略。在多个模仿学习基准测试中,CFDP在训练时间与性能之间取得良好平衡。此外,我们展示了闭式扩散策略作为可组合原语,能够实现在推理阶段对预训练神经扩散策略进行数据驱动的编辑,包括策略引导和新增示范增强。
原文摘要 · Abstract (English)
While diffusion-based policies have impressive performance and expressivity, their long offline training slows down the data collection and policy deployment loop. We introduce Closed-Form Diffusion Policies, a class of training-free diffusion-based policies for imitation learning using the closed-form score derived from the demonstration dataset. We deploy CFDP with real-time inference with a mobile CPU in hardware experiments, showing it can successfully perform imitation directly from the dataset in milliseconds and with faster inference than neural diffusion policies. In experiments on imitation learning benchmarks, we show that CFDP is competitive against neural baselines that require hours of training, providing a favorable tradeoff between training time and performance. Finally, we show how closed-form diffusion policies act as a composable primitive that enables data-driven inference-time editing of pre-trained neural diffusion policies, including policy guidance and novel demonstration augmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。