arXiv:2412.05766cs.LGcs.AI2024-12NeurIPS被引 6

让世界模型专注关键信息,避免被无关细节干扰。

Policy-shaped prediction: avoiding distractions in model-based reinforcement learning

  • 用分割模型+任务感知重建损失+对抗学习,引导模型聚焦有用信息。
  • 在复杂背景环境中,性能显著优于现有方法,超越DreamerV3等主流模型。
  • 适合追求高效、鲁棒的MBRL研究者和工业应用开发者。

基于模型的强化学习(MBRL)是实现样本高效策略优化的有前景方向。然而,基于重构的MBRL存在一个已知弱点:当世界中某些细节高度可预测但与策略无关时,模型会消耗大量容量学习这些无用内容,从而忽略重要环境动态。我们通过设计一个新环境验证了该问题在主流方法(如DreamerV3和DreamerPro)中的持续影响,其中背景干扰复杂、可预测且对决策无用。为此,我们提出一种融合预训练分割模型、任务感知重建损失与对抗学习的方法,有效引导世界模型容量聚焦关键信息。实验表明,该方法在多种去噪策略中表现最优,推动了更稳健的模型基强化学习发展。

原文摘要 · Abstract (English)

Model-based reinforcement learning (MBRL) is a promising route to sample-efficient policy optimization. However, a known vulnerability of reconstruction-based MBRL consists of scenarios in which detailed aspects of the world are highly predictable, but irrelevant to learning a good policy. Such scenarios can lead the model to exhaust its capacity on meaningless content, at the cost of neglecting important environment dynamics. While existing approaches attempt to solve this problem, we highlight its continuing impact on leading MBRL methods -- including DreamerV3 and DreamerPro -- with a novel environment where background distractions are intricate, predictable, and useless for planning future actions. To address this challenge we develop a method for focusing the capacity of the world model through synergy of a pretrained segmentation model, a task-aware reconstruction loss, and adversarial learning. Our method outperforms a variety of other approaches designed to reduce the impact of distractors, and is an advance towards robust model-based reinforcement learning.

MBRL强化学习注意力机制模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。