通过解耦干预与自然演化,提升模型预测的样本效率。
IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning

- 将动作影响与环境演变分离建模,避免策略偏差
- 在五个MuJoCo任务中实现更高采样效率与竞争力回报
- 适合追求高效强化学习的算法研究者
基于模型的强化学习(MBRL)通过学习环境动态生成合成经验,是实现高效决策的有前景方法。现有方法多通过不确定性估计、模型正则化和保守价值学习改进动态预测与策略优化,但通常将转移模型与评价器视为整体预测器,忽略了策略引起的数据偏差。这导致动作与环境演化纠缠,不均衡的动作覆盖会扭曲用于策略改进的反事实价值估计。为此,我们提出IADD-TR框架,融合干预感知动态解耦(IADD)与目标正则化(TR)。IADD将转移过程分解为动作干预阶段与无动作自然演化阶段,利用零动作锚点解决两阶段分解的非唯一性问题,其隐变量与状态对齐分量分别在可逆块内变换和逐点意义下可识别。对于策略学习,我们从重放状态策略梯度函数的高效影响函数推导出TR,通过动作密度加权残差修正增强评价器,并优化目标损失,在评价器或重放动作密度一致时实现双重稳健的策略梯度估计。在五个MuJoCo任务上的大量实验表明,IADD-TR在保持竞争力回报的同时显著提升了样本效率。
原文摘要 · Abstract (English)
Model-based reinforcement learning (MBRL), which learns environment dynamics to generate synthetic experience, is a promising approach to sample-efficient decision making. Numerous methods have been developed to improve dynamics prediction and policy optimization for MBRL through uncertainty estimation, model regularization, and conservative value learning. However, these methods typically treat the transition model and critic as monolithic predictors, overlooking the policy-induced data bias. Consequently, action can become entangled with environmental evolution, while uneven action coverage may distort the counterfactual value estimates used for policy improvement. To address this, we propose IADD-TR, a unified framework combining Intervention-Aware Dynamics Decoupling (IADD) and Targeted Regularization (TR). IADD factorizes transitions into an action-intervention stage and an action-free natural evolution stage, using a zero-action anchor to resolve the non-uniqueness of this two-stage factorization for robust generalization. Its latent and state-aligned components are identifiable up to an invertible within-block transformation and pointwise, respectively. For policy learning, we derive TR from the efficient influence function of a replay-state policy-gradient functional. TR augments the critic with an action-density-scaled residual correction and optimizes a targeted loss, yielding doubly robust policy-gradient estimation when either the critic or the replay action density is consistently specified. Extensive experiments on five MuJoCo tasks show that IADD-TR achieves competitive returns with improved sample efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。