arXiv:2506.19785cs.AI2025-06ICLR被引 1

通过隐动态建模任务信念相似性,提升元强化学习在稀疏奖励下的适应效率。

Learning Task Belief Similarity with Latent Dynamics for Meta-Reinforcement Learning

  • 基于贝叶斯自适应马尔可夫决策过程,用隐任务信念度量相似性
  • 在稀疏奖励环境下,比现有方法更快识别任务并完成探索
  • 适合需要快速适应新环境的机器人控制场景

元强化学习依赖于探索中获取的任务分布先验信息,以快速适应未知任务。代理探索效率取决于对当前任务的准确识别。近期基于贝叶斯自适应的深度强化学习方法通常依赖重构环境奖励信号,但在稀疏奖励场景下难以实现,导致次优利用。受双模拟度量启发,我们提出SimBelief——一种通过测量贝叶斯自适应马尔可夫决策过程(BAMDP)中任务信念相似性的新型元强化学习框架。该方法有效提取相似任务分布的共性特征,实现稀疏奖励环境中高效的任务识别与探索。引入隐任务信念度量,学习相似任务的共同结构,并融入具体任务信念。通过学习跨任务分布的隐动态,将共享的隐任务信念特征与特定任务特征关联,促进快速任务识别与适应。在稀疏奖励的MuJoCo和panda-gym任务上,本方法优于现有最优基线。

原文摘要 · Abstract (English)

Meta-reinforcement learning requires utilizing prior task distribution information obtained during exploration to rapidly adapt to unknown tasks. The efficiency of an agent's exploration hinges on accurately identifying the current task. Recent Bayes-Adaptive Deep RL approaches often rely on reconstructing the environment's reward signal, which is challenging in sparse reward settings, leading to suboptimal exploitation. Inspired by bisimulation metrics, which robustly extracts behavioral similarity in continuous MDPs, we propose SimBelief-a novel meta-RL framework via measuring similarity of task belief in Bayes-Adaptive MDP (BAMDP). SimBelief effectively extracts common features of similar task distributions, enabling efficient task identification and exploration in sparse reward environments. We introduce latent task belief metric to learn the common structure of similar tasks and incorporate it into the specific task belief. By learning the latent dynamics across task distributions, we connect shared latent task belief features with specific task features, facilitating rapid task identification and adaptation. Our method outperforms state-of-the-art baselines on sparse reward MuJoCo and panda-gym tasks.

元强化学习稀疏奖励任务识别贝叶斯推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。