arXiv:2605.17601cs.RO2026-05

用环境约束代替轨迹模仿,实现一次示范的高效泛化。

From a Single Demonstration to a General Policy for Contact-Rich Manipulation

论文配图:From a Single Demonstration to a General Policy for Contact-Rich Manipulation
图 1 · 摘自论文原文
  • 将示范分解为利用环境约束的行为单元,分离通用结构与具体细节
  • 在7个真实任务中实现超90%成功率,跨物体姿态和接触动态泛化
  • 适合需要快速适应新场景的机器人操作任务

我们提出一种学习从示范(LfD)框架,实现在多阶段、高接触性操作任务中的一次性泛化。核心思路是将环境约束作为归纳偏置。通过将示范表示为利用环境约束的一系列行为,机器人将任务通用结构(约束类型及其转换关系)与实例特定细节(如精确轨迹、姿态、局部几何)分离。我们的四阶段流程在此表示基础上构建完整策略:首先将单次示范抽象为环境约束基元,接着通过自引导探索消除歧义,然后融入针对性的人类修正以应对分布外变化,最后在线通过柔顺交互恢复被抽象掉的细节。由于策略遵循约束而非模仿轨迹,因此能跨物体姿态、局部几何及未建模接触动力学实现泛化。我们在七个真实世界多阶段高接触性操作任务上验证了该方法,成功率超过90%。大量实验结果确立了环境约束作为学习从示范中高效泛化的基础构件。

原文摘要 · Abstract (English)

We present a Learning from Demonstration (LfD) framework that achieves one-shot generalization in multi-stage, contact-rich manipulation tasks. Central to our approach is the utilization of environmental constraints as the inductive bias. By representing a demonstration as a sequence of behaviors that exploit environmental constraints, the robot separates task-general structure -- the constraint types and their transitions -- from instance-specific details such as exact demonstration trajectories, poses, and local geometries. Our four-stage pipeline builds a complete policy on this representation: the robot first abstracts a single demonstration into environmental-constraint primitives, then disambiguates them through self-guided exploration, next assimilates targeted human corrections that handle out-of-distribution variations, and finally recovers the abstracted-away details online through compliant interaction. Because the resulting policy follows constraints rather than mimics trajectories, it generalizes across object poses, local geometries, and unmodeled contact dynamics. We validate our approach on seven real-world multi-stage contact-rich manipulation tasks and achieve over 90% success. These extensive experimental results establish environmental constraints as fundamental building blocks for efficient generalization in learning from demonstration.

模仿学习泛化能力机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。