arXiv:2601.01003cs.LGcs.RO2026-01被引 1

通过收缩采样提升扩散策略鲁棒性,减少动作误差。

Contractive Diffusion Policies: Robust Action Diffusion via Contractive Score-Based Sampling with Differential Equations

  • 引入收缩性采样机制,让相近轨迹相互靠近
  • 在数据稀缺下仍保持优异性能,动作方差显著降低
  • 可无缝集成到现有模型,计算开销极小

扩散策略作为离线策略学习的强大生成模型,其采样过程可通过引导随机微分方程(SDE)的得分函数严格刻画。然而,这种基于得分的SDE建模虽然赋予模型学习多样化行为的灵活性,也带来了求解器误差、得分匹配误差、高数据需求以及动作生成不一致等问题。这些问题在图像生成中影响较小,但在连续控制任务中会累积并导致失败。本文提出收缩性扩散策略(CDPs),通过在扩散采样动态中引入收缩行为,使相近轨迹更紧密,从而增强对求解器误差和得分匹配误差的鲁棒性,并减少不必要的动作方差。我们提供了深入的理论分析及实用实现方案,可仅以极小修改与计算成本将CDPs融入现有扩散策略架构。在仿真与真实世界设置中进行广泛实验验证,结果显示,在多个基准测试中CDPs普遍优于基线策略,尤其在数据稀缺条件下优势明显。

原文摘要 · Abstract (English)

Diffusion policies have emerged as powerful generative models for offline policy learning, whose sampling process can be rigorously characterized by a score function guiding a stochastic differential equation (SDE). However, the same score-based SDE modeling that grants diffusion policies the flexibility to learn diverse behavior also incurs solver and score-matching errors, large data requirements, and inconsistencies in action generation. While less critical in image generation, these inaccuracies compound and lead to failure in continuous control settings. We introduce contractive diffusion policies (CDPs) to induce contractive behavior in the diffusion sampling dynamics. Contraction pulls nearby flows closer to enhance robustness against solver and score-matching errors while reducing unwanted action variance. We develop an in-depth theoretical analysis along with a practical implementation recipe to incorporate CDPs into existing diffusion policy architectures with minimal modification and computational cost. We evaluate CDPs for offline learning by conducting extensive experiments in simulation and real-world settings. Across benchmarks, CDPs often outperform baseline policies, with pronounced benefits under data scarcity.

扩散模型强化学习策略学习鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。