揭示扩散策略的表达能力与统计代价的权衡,指导实际训练中参数选择。
Expressivity and Statistical Trade-offs in Diffusion Policy Learning

- 以漂移项的Lipschitz常数K为关键变量,量化策略表达能力
- 样本量n下性能误差达$ ilde{O}(n^{-2/(m+6)})$,理论可优化
- 提出按样本数选K值的实用原则,适合强化学习研究者
基于扩散的策略近期成为强化学习中强大的策略参数化方法,将状态条件下的动作分布表示为带有参数化漂移项的扩散过程的终端分布。该表示在实践中展现出显著的表达灵活性,能建模复杂、多模态且高度非高斯的动作分布;然而,其数学本质驱动机制以及在有限数据下如何充分利用仍不明确。本文识别出漂移项的Lipschitz预算$K$是控制扩散策略表达力和统计行为的核心量。通过逼近理论,我们证明:具有$K$-Lipschitz漂移的扩散策略可集中在最优确定性策略附近,价值逼近误差为$1/K$量级;同时在非退化扩散噪声下建立了匹配的下界。更高的表达力伴随统计代价:当漂移由神经网络参数化时,增大$K$虽提升逼近能力,但增加统计复杂度。二者平衡后,在通用神经网络漂移下,有限样本性能差距为$ ilde{O}(n^{-2/(m+6)})$,而在单侧耗散漂移类中可达更优的$ ilde{O}(n^{-2/(m+4)})$,其中$n$为样本量,$m$为状态空间维数。数值实验验证了$K$随样本变化的权衡效应,支持理论预测。框架还提出实践原则:根据可用样本量选择扩散预算$K$,再选用对应固定Lipschitz系数的神经网络结构。
原文摘要 · Abstract (English)
Diffusion-based policies have recently emerged as powerful policy parameterizations for reinforcement learning, representing state-conditioned action distributions as terminal laws of diffusion processes with parameterized drifts. This terminal-law representation has shown substantial expressive flexibility in practice, enabling diffusion policies to model complex, multimodal, and highly non-Gaussian action distributions; however, it remains unclear what mathematically drives this expressivity and how to fully exploit it when the policy is learned from finite data. In this paper, we identify the drift Lipschitz budget $K$ as a central quantity governing the expressivity and statistical behavior of diffusion policies. We quantify expressivity through approximation: diffusion policies with $K$-Lipschitz drifts can concentrate near optimal deterministic policies and achieve value approximation error of order $1/K$; moreover, we prove a matching lower bound under nondegenerate diffusion noise. This increased expressivity comes with a statistical cost. When the drift is parameterized by neural networks, increasing $K$ improves approximation but increases statistical complexity. Balancing these two terms yields a finite-sample performance gap of order $\tilde{O}(n^{-2/(m+6)})$ for generic neural-network drifts, and a sharper rate $\tilde{O}(n^{-2/(m+4)})$ for one-sided dissipative drift classes, where $n$ is the sample size and $m$ is the dimension of the state space. Numerical experiments provide empirical evidence for the sample-dependent trade-off in $K$, supporting both theoretical regimes. Our framework also suggests a practical implementation principle: choose the diffusion budget $K$ according to the available sample size, and then select a neural-network architecture with the corresponding fixed Lipschitz coefficient.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。