arXiv:2409.14216cs.ROcs.AI2024-09ICRA被引 13

用主动推理提升机器人在稀疏奖励下的视觉控制能力

R-AIF: Solving Sparse-Reward Robotic Tasks from Pixels with Active Inference and World Models

论文配图:R-AIF: Solving Sparse-Reward Robotic Tasks from Pixels with Active Inference and World Models
图 1 · 摘自论文原文
  • 结合世界模型与主动推理,从像素中推断隐藏状态
  • 在连续动作空间下实现更高成功率和更稳定表现
  • 适合研究稀疏奖励下强化学习与认知建模的学者

尽管主动推理(AIF)在马尔可夫决策过程(MDP)中展现出良好效果,但在部分可观测马尔可夫决策过程(POMDP)场景下的研究仍较少。在POMDP中,智能体需从原始感官输入(如图像像素)推断未观测环境状态。尤其在连续动作空间、稀疏奖励信号的复杂控制任务中,现有方法表现有限。本文提出新型先验偏好学习机制与自我修正调度策略,显著提升智能体在稀疏奖励、连续动作、目标导向的机器人控制任务中的性能。实验表明,所提方法在累积奖励、相对稳定性及成功率达上优于当前最优模型。

原文摘要 · Abstract (English)

Although research has produced promising results demonstrating the utility of active inference (AIF) in Markov decision processes (MDPs), there is relatively less work that builds AIF models in the context of environments and problems that take the form of partially observable Markov decision processes (POMDPs). In POMDP scenarios, the agent must infer the unobserved environmental state from raw sensory observations, e.g., pixels in an image. Additionally, less work exists in examining the most difficult form of POMDP-centered control: continuous action space POMDPs under sparse reward signals. In this work, we address issues facing the AIF modeling paradigm by introducing novel prior preference learning techniques and self-revision schedules to help the agent excel in sparse-reward, continuous action, goal-based robotic control POMDP environments. Empirically, we show that our agents offer improved performance over state-of-the-art models in terms of cumulative rewards, relative stability, and success rate. The code in support of this work can be found at https://github.com/NACLab/robust-active-inference.

主动推理稀疏奖励世界模型机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。