通过预测物体部件运动,实现跨机器人的动作规划。
Embodiment-Agnostic Action Planning via Object-Part Scene Flow
- 基于物体部件的3D场景流预测动作轨迹
- 在虚拟环境上性能提升27.7%和26.2%
- 仅需人类示范即可部署到多种机器人
我们发现机器人动作规划的关键在于理解末端执行器操作目标物体部件时的运动。为此,提出生成3D物体部件场景流并提取其变换,以求解多样机器人的动作轨迹。该方法从物体运动预测显式推导机器人动作,提升策略鲁棒性。不同于依赖特定机器人的训练数据,本方法具备机器人无关性,可跨多种机器人泛化,并能从人类示范中学习。方法包含三部分:物体部件定位器、RGBD视频生成器和轨迹规划器。即使在无轨迹标注的视频数据上训练,仍显著优于现有方法,在MetaWorld和Franka-Kitchen虚拟环境中分别提升27.7%和26.2%。真实世界实验表明,仅用人类示范训练的策略可成功部署于多种机器人平台。
原文摘要 · Abstract (English)
Observing that the key for robotic action planning is to understand the target-object motion when its associated part is manipulated by the end effector, we propose to generate the 3D object-part scene flow and extract its transformations to solve the action trajectories for diverse embodiments. The advantage of our approach is that it derives the robot action explicitly from object motion prediction, yielding a more robust policy by understanding the object motions. Also, beyond policies trained on embodiment-centric data, our method is embodiment-agnostic, generalizable across diverse embodiments, and being able to learn from human demonstrations. Our method comprises three components: an object-part predictor to locate the part for the end effector to manipulate, an RGBD video generator to predict future RGBD videos, and a trajectory planner to extract embodiment-agnostic transformation sequences and solve the trajectory for diverse embodiments. Trained on videos even without trajectory data, our method still outperforms existing works significantly by 27.7% and 26.2% on the prevailing virtual environments MetaWorld and Franka-Kitchen, respectively. Furthermore, we conducted real-world experiments, showing that our policy, trained only with human demonstration, can be deployed to various embodiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。