让机器人提前‘想象’隐藏参数,快速适应真实环境。
PrivilegedDreamer: Explicit Imagination of Privileged Information for Rapid Adaptation of Learned Policies
- 通过双循环结构显式估计隐藏参数并用于策略优化
- 在5个任务中超越现有最先进方法,适应速度更快
- 适合需模拟到现实迁移的机器人控制场景
许多真实世界的控制问题涉及受不可观测隐藏参数影响的动力学和目标,如自动驾驶和机器人操作,导致仿真到现实的迁移性能下降。为此,我们采用隐参数马尔可夫决策过程(HIP-MDPs)建模此类问题,其中隐藏变量参数化转移和奖励函数。现有方法如领域随机化、领域自适应和元学习仅将隐藏参数的影响视为额外方差,难以有效处理参数化奖励的HIP-MDP问题。本文提出PrivilegedDreamer,一种基于模型的强化学习框架,通过引入显式参数估计模块扩展现有方法。其新颖的双循环架构能从有限历史数据中显式估计隐藏参数,并将这些估计值用于条件化模型、智能体和价值网络。在五个多样化的HIP-MDP任务上的实证分析表明,PrivilegedDreamer优于现有的基于模型、无模型及领域自适应学习算法。此外,我们通过消融实验验证了各组件的有效性。
原文摘要 · Abstract (English)
Numerous real-world control problems involve dynamics and objectives affected by unobservable hidden parameters, ranging from autonomous driving to robotic manipulation, which cause performance degradation during sim-to-real transfer. To represent these kinds of domains, we adopt hidden-parameter Markov decision processes (HIP-MDPs), which model sequential decision problems where hidden variables parameterize transition and reward functions. Existing approaches, such as domain randomization, domain adaptation, and meta-learning, simply treat the effect of hidden parameters as additional variance and often struggle to effectively handle HIP-MDP problems, especially when the rewards are parameterized by hidden variables. We introduce Privileged-Dreamer, a model-based reinforcement learning framework that extends the existing model-based approach by incorporating an explicit parameter estimation module. PrivilegedDreamer features its novel dual recurrent architecture that explicitly estimates hidden parameters from limited historical data and enables us to condition the model, actor, and critic networks on these estimated parameters. Our empirical analysis on five diverse HIP-MDP tasks demonstrates that PrivilegedDreamer outperforms state-of-the-art model-based, model-free, and domain adaptation learning algorithms. Additionally, we conduct ablation studies to justify the inclusion of each component in the proposed architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。