用概率框架提升机器人策略在少样本下的泛化能力
PAC-DP: PAC-Bayesian Diffusion Policy Learning

- 将扩散策略建模为贝叶斯神经网络,引入先验后验约束
- 在低数据场景下成功率显著提升,负对数似然降低
- 适合数据稀缺的复杂机器人任务,理论与实践兼备
扩散策略(DPs)可执行复杂操作任务,但通常通过最小化去噪目标训练,在机器人领域常见的有限数据条件下泛化控制能力较弱。本文提出PAC-DP,通过将DP视为贝叶斯神经网络,并定义PAC-Bayes泛化界,推导出一种新训练目标:在标准去噪损失基础上增加后验与先验参数分布间的KL散度正则项。理论上,该方法为DP训练提供了有理论依据的正则化机制,且不显著增加训练时间。实验表明,该方法在多个机器人操作基准上提升了去噪性能,降低了变分负对数似然,提高了成功率。尤其在低数据训练和复杂任务中效果最为显著,验证了PAC-DP作为机器人策略学习的理论可靠框架。
原文摘要 · Abstract (English)
Diffusion Policies (DPs) are able to perform complex manipulation tasks. However, DPs are typically trained by minimizing a denoising objective, which provides limited control over generalization in the finite-data regimes common in robotics. In this letter, we propose PAC-DP, an approach that increases the performance of DPs in robotic manipulation tasks. By modeling the DP as a Bayesian neural network, and defining a PAC-Bayes generalization bound, we derive a novel training objective that augments the standard denoising loss with a Kullback-Leibler divergence regularizer between the posterior and prior parameter distributions. From the theoretical perspective, our approach provides a principled approach to regularize the training of DPs without significantly increasing the training time. From the practical point of view, experimental results demonstrate improved denoising performance, lower variational negative log-likelihood, and higher success rates across multiple robotic manipulation benchmarks. Crucially, the largest improvements are observed in low-data training regimes and complex tasks, establishing PAC-DP as a theoretically grounded framework for robot policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。