提出统一约束优化框架,自动调节约束类型与强度。
Automatic Constraint Policy Optimization based on Continuous Constraint Interpolation Framework for Offline Reinforcement Learning
- 通过插值参数统一三类约束,实现平滑切换。
- 在D4RL和NeoRL2上达当前最优,提升稳定性和泛化性。
- 适合需要高效策略优化的离线强化学习场景。
离线强化学习依赖策略约束以缓解外推误差,而约束形式与强度对性能影响重大。然而现有方法多固定使用单一约束族:加权行为克隆、密度正则化或支持约束,缺乏统一解释其关联与权衡的理论框架。本文提出连续约束插值(CCI)框架,使这三类约束成为同一约束谱上的特例。该框架引入单一插值参数,实现约束类型间的平滑过渡与合理组合。基于CCI,我们设计自动约束策略优化(ACPO),一种基于拉格朗日对偶更新的实用原始-对偶算法,可自适应调整插值参数。此外,我们建立了最大熵性能差异引理,并推导出闭式最优策略及其参数投影的性能下界。在D4RL和NeoRL2上的实验表明,该方法在多种任务中均取得稳健提升,整体达到当前最优性能。
原文摘要 · Abstract (English)
Offline Reinforcement Learning (RL) relies on policy constraints to mitigate extrapolation error, where both the constraint form and constraint strength critically shape performance. However, most existing methods commit to a single constraint family: weighted behavior cloning, density regularization, or support constraints, without a unified principle that explains their connections or trade-offs. In this work, we propose Continuous Constraint Interpolation (CCI), a unified optimization framework in which these three constraint families arise as special cases along a common constraint spectrum. The CCI framework introduces a single interpolation parameter that enables smooth transitions and principled combinations across constraint types. Building on CCI, we develop Automatic Constraint Policy Optimization (ACPO), a practical primal--dual algorithm that adapts the interpolation parameter via a Lagrangian dual update. Moreover, we establish a maximum-entropy performance difference lemma and derive performance lower bounds for both the closed-form optimal policy and its parametric projection. Experiments on D4RL and NeoRL2 demonstrate robust gains across diverse domains, achieving state-of-the-art performance overall.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。