用时序差分与度量学习提升机器人路径规划的准确性
Physics-informed Temporal Difference Metric Learning for Robot Motion Planning
- 结合时序差分与度量学习,更准确求解运动规划中的Eikonal方程
- 在2到12自由度系统中,复杂与未见环境下的成功率显著提升
- 适合需要高效自监督路径规划的机器人系统研发者
运动规划旨在为机器人从起始配置到目标配置找到无碰撞路径。近年来,自监督学习方法无需昂贵专家示范即可解决该问题,通过求解Eikonal方程训练神经网络,实现高效求解。然而,这些方法在复杂环境中表现不佳,因未能保持Eikonal方程的关键性质,如最优值函数和测地距离。为此,我们提出一种新型自监督时序差分度量学习方法,更精确求解Eikonal方程,提升复杂及未见环境中的规划性能。该方法在有限区域内施加贝尔曼最优性原理,利用时序差分学习避免虚假局部极小值,同时结合度量学习保留Eikonal方程的核心测地性质。实验表明,该方法在2至12自由度的机器人系统中,显著优于现有自监督学习方法,在复杂与未见环境中均表现更优。
原文摘要 · Abstract (English)
The motion planning problem involves finding a collision-free path from a robot's starting to its target configuration. Recently, self-supervised learning methods have emerged to tackle motion planning problems without requiring expensive expert demonstrations. They solve the Eikonal equation for training neural networks and lead to efficient solutions. However, these methods struggle in complex environments because they fail to maintain key properties of the Eikonal equation, such as optimal value functions and geodesic distances. To overcome these limitations, we propose a novel self-supervised temporal difference metric learning approach that solves the Eikonal equation more accurately and enhances performance in solving complex and unseen planning tasks. Our method enforces Bellman's principle of optimality over finite regions, using temporal difference learning to avoid spurious local minima while incorporating metric learning to preserve the Eikonal equation's essential geodesic properties. We demonstrate that our approach significantly outperforms existing self-supervised learning methods in handling complex environments and generalizing to unseen environments, with robot configurations ranging from 2 to 12 degrees of freedom (DOF).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。