arXiv:2607.26040cs.LGstat.ML2026-07

用潜在引导提升世界模型,让智能体学得更快更准。

Reinformed Dreamer: An Asymmetric World Model Efficiently Trained through Latent Guidance

  • 用潜在空间中的额外信息指导训练,优化状态表示。
  • 在多个基准上表现优于传统方法,提升更稳定。
  • 适合想提升模型训练效率的强化学习研究者。

与人类学习时需要指导类似,强化学习算法在训练中也可受益于奖励之外的额外监督。通过引入额外信息来学习更好的表征和行为,是不对称强化学习的核心思想。该范式在部分可观测环境下(有额外状态信息)已证明有效,即便在完全可观测条件下,当存在更精细的状态信息时也适用。本文聚焦基于模型的强化学习,研究不对称学习对观察表征和特权信息表征的影响。首先,我们发现已有算法Informed Dreamer在特权信息表征上存在局限;随后提出一种新的基于潜在引导的不对称表征学习目标,构建出新算法Reinformed Dreamer。在多个基准测试中,其性能一致性超越先前的不对称方法,显示出更强的泛化能力与稳定性。

原文摘要 · Abstract (English)

Much like humans benefit from guidance while learning, reinforcement learning algorithms may benefit from additional supervision beyond rewards. Leveraging additional information during training to learn better representations and behaviors has been the focus of asymmetric reinforcement learning. This learning paradigm has proven effective under partial observability when additional state information is available, but also under full observability when more refined state information is available. Focusing on model-based reinforcement learning, we study the effect of asymmetric learning on observation representations and on privileged information representations. First, we identify a limitation in the privileged information representations learned by an asymmetric model-based algorithm known as the Informed Dreamer. Then, we propose a novel asymmetric representation learning objective using latent guidance, resulting in a new algorithm called the Reinformed Dreamer. Experiments across several benchmarks show a more consistent improvement over Dreamer than previous asymmetric approaches.

强化学习世界模型表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。