arXiv:2502.00850cs.LGcs.AI2025-02

解决离线强化学习中模型与策略不一致的问题,提升真实环境下的稳定性。

Dual Alignment Maximin Optimization for Offline Model-based RL

  • 提出双对齐最大化优化框架,统一处理模型与策略一致性。
  • 在多个基准任务上表现优异,有效避免分布外状态和动作。
  • 适合关注离线强化学习稳定性和泛化能力的研究者。

离线强化学习代理因合成数据与真实环境间的分布差异面临部署挑战。现有研究多聚焦于提升合成采样的保真度及引入无偏策略机制,但直接集成的范式常无法保证在存在偏差的模型和环境动态下策略行为的一致性,这源于行为策略与学习策略之间的固有差异。本文将关注点从模型可靠性转向策略差异,在优化期望回报的同时,自洽地整合合成数据,提出一种新的演员-评论家框架:双对齐最大化优化(DAMO)。该框架统一确保模型-环境策略一致性与合成数据与离线数据的兼容性。内层最小化执行双重保守价值估计,对齐策略与轨迹,避免分布外状态和动作;外层最大化则确保策略改进与内层价值估计保持一致。实验表明,DAMO能有效实现模型与策略对齐,在多种基准任务中取得具有竞争力的性能。

原文摘要 · Abstract (English)

Offline reinforcement learning agents face significant deployment challenges due to the synthetic-to-real distribution mismatch. While most prior research has focused on improving the fidelity of synthetic sampling and incorporating off-policy mechanisms, the directly integrated paradigm often fails to ensure consistent policy behavior in biased models and underlying environmental dynamics, which inherently arise from discrepancies between behavior and learning policies. In this paper, we first shift the focus from model reliability to policy discrepancies while optimizing for expected returns, and then self-consistently incorporate synthetic data, deriving a novel actor-critic paradigm, Dual Alignment Maximin Optimization (DAMO). It is a unified framework to ensure both model-environment policy consistency and synthetic and offline data compatibility. The inner minimization performs dual conservative value estimation, aligning policies and trajectories to avoid out-of-distribution states and actions, while the outer maximization ensures that policy improvements remain consistent with inner value estimates. Empirical evaluations demonstrate that DAMO effectively ensures model and policy alignments, achieving competitive performance across diverse benchmark tasks.

强化学习离线学习策略对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。