让机器人生成符合物理规律的目标,提升自主学习效率。
Physically-Grounded Goal Imagination: Physics-Informed Variational Autoencoder for Self-Supervised Reinforcement Learning
- 将物理规律嵌入变分自编码器,分离动态与外观变量
- 生成的目标符合碰撞、物体恒常性等物理约束
- 在抓取、推移等任务中显著提升探索效率和技能掌握
自监督目标条件强化学习使机器人无需人工标注即可自主习得多样化技能。但核心挑战在于目标设定:机器人需提出当前环境中可实现且多样化的可行目标。现有方法如RIG(基于想象目标的视觉强化学习)使用变分自编码器(VAE)在学习的隐空间生成目标,但存在生成物理上不合理的目标,阻碍学习效率的问题。本文提出物理信息增强的RIG(PI-RIG),通过一种新型增强型物理信息变分自编码器(Enhanced p3-VAE),将物理约束直接融入VAE训练过程,实现物理一致且可达成的目标生成。其关键创新在于显式分离隐空间中的物理变量(控制物体动力学)与环境变量(捕捉视觉外观),并通过微分方程约束与守恒定律强制物理一致性。这使得生成的目标尊重基本物理原理,如物体恒常性、碰撞约束与动态可行性。大量实验表明,该物理信息目标生成显著提升了目标质量,促进更有效的探索,在包含抓取、推动和拾放等视觉机器人操作任务中实现了更好的技能习得。
原文摘要 · Abstract (English)
Self-supervised goal-conditioned reinforcement learning enables robots to autonomously acquire diverse skills without human supervision. However, a central challenge is the goal setting problem: robots must propose feasible and diverse goals that are achievable in their current environment. Existing methods like RIG (Visual Reinforcement Learning with Imagined Goals) use variational autoencoder (VAE) to generate goals in a learned latent space but have the limitation of producing physically implausible goals that hinder learning efficiency. We propose Physics-Informed RIG (PI-RIG), which integrates physical constraints directly into the VAE training process through a novel Enhanced Physics-Informed Variational Autoencoder (Enhanced p3-VAE), enabling the generation of physically consistent and achievable goals. Our key innovation is the explicit separation of the latent space into physics variables governing object dynamics and environmental factors capturing visual appearance, while enforcing physical consistency through differential equation constraints and conservation laws. This enables the generation of physically consistent and achievable goals that respect fundamental physical principles such as object permanence, collision constraints, and dynamic feasibility. Through extensive experiments, we demonstrate that this physics-informed goal generation significantly improves the quality of proposed goals, leading to more effective exploration and better skill acquisition in visual robotic manipulation tasks including reaching, pushing, and pick-and-place scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。