让智能体在变化环境中自主学习行为,提升适应能力。
Adaptive World Models: Learning Behaviors by Latent Imagination Under Non-Stationarity
- 提出隐参数-部分可观测马尔可夫决策过程新框架
- 在多种非平稳强化学习任务中实现鲁棒行为学习
- 无监督生成结构化、任务感知的潜在空间
构建基础世界模型是具身智能的关键研究方向,而适应非平稳环境的能力是重要评判标准。本文提出一种新形式——隐参数-POMDP,用于基于自适应世界模型的控制。实验表明,该方法可在多种非平稳强化学习基准上学习出鲁棒行为。此外,该框架能无监督地学习任务抽象,生成结构化的、任务感知的潜在空间。
原文摘要 · Abstract (English)
Developing foundational world models is a key research direction for embodied intelligence, with the ability to adapt to non-stationary environments being a crucial criterion. In this work, we introduce a new formalism, Hidden Parameter-POMDP, designed for control with adaptive world models. We demonstrate that this approach enables learning robust behaviors across a variety of non-stationary RL benchmarks. Additionally, this formalism effectively learns task abstractions in an unsupervised manner, resulting in structured, task-aware latent spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。