通过密度感知提升离线强化学习的安全性,避免探索无效状态
Variational OOD State Correction for Offline Reinforcement Learning
- 在变分框架中同时优化动作结果与数据密度,引导智能体留在安全区域
- 在MuJoCo和AntMaze上验证,显著提升离线强化学习的性能与稳定性
- 适合研究离线强化学习安全性、分布外状态处理的开发者参考
离线强化学习的性能受状态分布偏移问题严重影响,而分布外(OOD)状态修正是一种常见应对方法。本文提出一种名为密度感知安全感知(DASP)的新方法,通过鼓励智能体优先选择能带来高数据密度结果的动作,从而促进其在分布内或返回分布内(安全)区域运行。该方法在变分框架下优化目标,同时考虑决策可能的结果及其密度,为安全决策提供关键上下文信息。我们在离线MuJoCo和AntMaze基准上进行了大量实验验证,结果表明所提方法在有效性与可行性方面均表现优异。
原文摘要 · Abstract (English)
The performance of Offline reinforcement learning is significantly impacted by the issue of state distributional shift, and out-of-distribution (OOD) state correction is a popular approach to address this problem. In this paper, we propose a novel method named Density-Aware Safety Perception (DASP) for OOD state correction. Specifically, our method encourages the agent to prioritize actions that lead to outcomes with higher data density, thereby promoting its operation within or the return to in-distribution (safe) regions. To achieve this, we optimize the objective within a variational framework that concurrently considers both the potential outcomes of decision-making and their density, thus providing crucial contextual information for safe decision-making. Finally, we validate the effectiveness and feasibility of our proposed method through extensive experimental evaluations on the offline MuJoCo and AntMaze suites.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。