arXiv:2508.02159cs.LG2025-08ICML被引 9

用特权信息提升部分可观测强化学习的安全与性能

PIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement Learning

  • 设计不对称约束的马尔可夫决策过程,理论分析特权信息优势
  • 通过特权表征对齐和异构演员-评论家结构,显著提升安全性和效率
  • 适合研究安全强化学习、世界模型与特权信息融合的学者

部分可观测性给安全强化学习带来重大挑战,因难以识别潜在风险与奖励。在训练中利用特定类型的特权信息以缓解部分可观测性问题,已取得显著实证成功。本文提出非对称约束的部分可观测马尔可夫决策过程(ACPOMDPs),从理论上分析特权信息在安全强化学习中的优势。基于ACPOMDPs,我们提出特权信息引导的Dreamer(PIGDreamer),一种基于模型的强化学习方法,通过特权表征对齐和异构演员-评论家结构,提升智能体的安全性与性能。实验表明,PIGDreamer显著优于现有安全强化学习方法,并在性能、鲁棒性与效率上超越其他特权强化学习方案。代码已开源:https://github.com/hggforget/PIGDreamer。

原文摘要 · Abstract (English)

Partial observability presents a significant challenge for Safe Reinforcement Learning (Safe RL), as it impedes the identification of potential risks and rewards. Leveraging specific types of privileged information during training to mitigate the effects of partial observability has yielded notable empirical successes. In this paper, we propose Asymmetric Constrained Partially Observable Markov Decision Processes (ACPOMDPs) to theoretically examine the advantages of incorporating privileged information in Safe RL. Building upon ACPOMDPs, we propose the Privileged Information Guided Dreamer (PIGDreamer), a model-based RL approach that leverages privileged information to enhance the agent's safety and performance through privileged representation alignment and an asymmetric actor-critic structure. Our empirical results demonstrate that PIGDreamer significantly outperforms existing Safe RL methods. Furthermore, compared to alternative privileged RL methods, our approach exhibits enhanced performance, robustness, and efficiency. Codes are available at: https://github.com/hggforget/PIGDreamer.

强化学习安全学习世界模型特权信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。