arXiv:2605.27877cs.LGcs.AI2026-05

通过局部残差修正提升离线策略优化,避免动作偏离数据分布。

SPAR: Support-Preserving Action Rectification

论文配图:SPAR: Support-Preserving Action Rectification
图 1 · 摘自论文原文
  • 以冻结行为克隆策略为锚点,局部修正残差空间中的策略
  • 在D4RL上显著提升低质量基线性能,达当前最优水平
  • 适合离线强化学习中需安全改进策略的研究者

离线策略改进面临最大化价值与拟合数据分布之间的固有矛盾。基于样本内加权回归的方法虽稳定,却因过度保守而抑制分布尾部的高价值动作;而基于梯度的方法常出现拟合-优化冲突,导致策略偏离数据流形。为此,我们提出支持保持型动作修正(SPAR),将全局学习重构为以冻结纯行为克隆策略为锚点的局部残差修正。该框架在残差空间中进行细粒度拟合与局部策略改进,从而缩小搜索空间。进一步引入潜在自模仿机制,采用潜在采样加权回归解决残差空间中的拟合-改进梯度冲突。理论上证明该机制消除了标准价值梯度的流形法向漂移;大量D4RL实验表明,SPAR能从次优基线中提取显著增益,实现当前最优性能。

原文摘要 · Abstract (English)

Offline policy improvement faces an inherent conflict between maximizing value and fitting the data distribution. While in-sample weighted regression is stable, it suffers from over-conservatism that suppresses high-value actions in the distribution tail; conversely, gradient-based approaches often exhibit a fitting-optimization conflict of gradients, which drives the policy off the data manifold. To address this, we propose Support-Preserving Action Rectification (SPAR), which reframes global learning as a local residual rectification anchored to a frozen pure behavior cloning policy. This framework performs fine-grained fitting and local policy improvement in the residual space, thereby contracting the search space. We further introduce Latent Self-Imitation, utilizing a latent-sampling weighted-regression mechanism to address fitting-improvement gradient conflict in the residual space. Theoretically, we prove this mechanism eliminates the manifold-normal drift of standard value gradients, while extensive D4RL experiments show SPAR extracts significant gains from suboptimal baselines to achieve state-of-the-art performance.

离线强化学习策略改进残差修正D4RL

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。