将机器人示范视为几何曲线,用距离场生成可控运动指令。
Geometry-aware Policy Imitation
- 把示范数据看作几何曲线,构建距离场生成控制流。
- 仿真与实机测试中成功率更高,速度是扩散模型的20倍。
- 适合需要快速响应和可解释性的机器人控制任务。
我们提出几何感知策略模仿(GPI)方法,将示范视为几何曲线而非离散的状态-动作样本集合。基于这些曲线,GPI 构建距离场,生成两种互补的控制原语:沿专家轨迹推进的前进流和纠正偏差的吸引流。二者结合形成可调控的非参数向量场,直接指导机器人行为。该框架将度量学习与策略合成解耦,支持在低维状态空间与高维感知输入间模块化适配。GPI 通过保留不同示范为独立模型自然支持多模态,并可通过简单叠加距离场高效组合新示范。我们在多种任务的仿真与真实机器人上进行了评估。结果表明,GPI 在成功率上优于扩散模型,运行速度快20倍,内存占用更低,且对扰动具有鲁棒性。这些结果确立了GPI作为生成式方法在机器人模仿学习中的高效、可解释且可扩展的替代方案。
原文摘要 · Abstract (English)
We propose a Geometry-aware Policy Imitation (GPI) approach that rethinks imitation learning by treating demonstrations as geometric curves rather than collections of state-action samples. From these curves, GPI derives distance fields that give rise to two complementary control primitives: a progression flow that advances along expert trajectories and an attraction flow that corrects deviations. Their combination defines a controllable, non-parametric vector field that directly guides robot behavior. This formulation decouples metric learning from policy synthesis, enabling modular adaptation across low-dimensional robot states and high-dimensional perceptual inputs. GPI naturally supports multimodality by preserving distinct demonstrations as separate models and allows efficient composition of new demonstrations through simple additions to the distance field. We evaluate GPI in simulation and on real robots across diverse tasks. Experiments show that GPI achieves higher success rates than diffusion-based policies while running 20 times faster, requiring less memory, and remaining robust to perturbations. These results establish GPI as an efficient, interpretable, and scalable alternative to generative approaches for robotic imitation learning. Project website: https://yimingli1998.github.io/projects/GPI/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。