arXiv:2410.08822cs.LGcs.AI2024-10被引 24

用对象中心的潜空间建模环境动态,提升机器人操作的样本效率。

SOLD: Slot Object-Centric Latent Dynamics Models for Relational Manipulation Learning from Pixels

  • 基于槽注意力机制,从像素中无监督学习对象级动态模型。
  • 在多个需关系推理的机器人任务上超越DreamerV3和TD-MPC2。
  • 潜空间结构清晰,适合行为模型进行可解释性决策。

学习潜在动力学模型能提供一种与任务无关的环境理解表征。利用该知识进行模型基于强化学习(RL),可通过想象的轨迹回放提高样本效率,优于无模型方法。此外,由于潜在空间作为行为模型的输入,世界模型所学的丰富表征有助于高效学习所需技能。现有方法多依赖环境状态的整体表征,而人类则通过对象及其交互来推理,预测动作对周围特定部分的影响。受此启发,我们提出槽注意力对象中心潜动力学模型(SOLD),一种新颖的模型基于强化学习算法,可从像素输入中无监督地学习对象中心的动力学模型。实验表明,该结构化的潜空间不仅提升了模型可解释性,还为行为模型提供了有价值的推理输入。SOLD在多个需关系推理与操作能力的基准机器人环境中表现优于当前最优的模型基于方法DreamerV3与TD-MPC2。视频演示见https://slot-latent-dynamics.github.io/。

原文摘要 · Abstract (English)

Learning a latent dynamics model provides a task-agnostic representation of an agent's understanding of its environment. Leveraging this knowledge for model-based reinforcement learning (RL) holds the potential to improve sample efficiency over model-free methods by learning from imagined rollouts. Furthermore, because the latent space serves as input to behavior models, the informative representations learned by the world model facilitate efficient learning of desired skills. Most existing methods rely on holistic representations of the environment's state. In contrast, humans reason about objects and their interactions, predicting how actions will affect specific parts of their surroundings. Inspired by this, we propose Slot-Attention for Object-centric Latent Dynamics (SOLD), a novel model-based RL algorithm that learns object-centric dynamics models in an unsupervised manner from pixel inputs. We demonstrate that the structured latent space not only improves model interpretability but also provides a valuable input space for behavior models to reason over. Our results show that SOLD outperforms DreamerV3 and TD-MPC2 - state-of-the-art model-based RL algorithms - across a range of benchmark robotic environments that require relational reasoning and manipulation capabilities. Videos are available at https://slot-latent-dynamics.github.io/.

强化学习对象中心潜动力学机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。