用3D关键点定义任务,让机器人灵活模仿复杂动作。
Correspondence-Oriented Imitation Learning: Flexible Visuomotor Control with 3D Conditioning
- 以场景中可变数量的关键点运动定义任务,支持灵活的时空控制。
- 在真实机械臂任务上,对稀疏和密集指令均表现优于现有方法。
- 自监督训练生成对应标签,无需人工标注,适合复杂操作场景。
我们提出对应关系导向的模仿学习(COIL),一种面向视觉-运动控制的条件策略学习框架,支持三维空间中的灵活任务表达。每个任务由场景中选定物体上的关键点运动目标定义,不假设固定关键点数量或均匀时间间隔,可适应不同用户意图与任务需求。为将这种对应关系导向的任务表示稳健映射到动作,我们设计了一种基于时空注意力机制的条件策略,有效融合多模态输入信息。策略通过可扩展的自监督流程训练,利用仿真中收集的演示数据,并在事后自动生成对应标签。COIL在跨任务、跨物体及运动模式下具有良好泛化能力,在真实世界操作任务中,无论是稀疏还是密集指令,均显著优于先前方法。
原文摘要 · Abstract (English)
We introduce Correspondence-Oriented Imitation Learning (COIL), a conditional policy learning framework for visuomotor control with a flexible task representation in 3D. At the core of our approach, each task is defined by the intended motion of keypoints selected on objects in the scene. Instead of assuming a fixed number of keypoints or uniformly spaced time intervals, COIL supports task specifications with variable spatial and temporal granularity, adapting to different user intents and task requirements. To robustly ground this correspondence-oriented task representation into actions, we design a conditional policy with a spatio-temporal attention mechanism that effectively fuses information across multiple input modalities. The policy is trained via a scalable self-supervised pipeline using demonstrations collected in simulation, with correspondence labels automatically generated in hindsight. COIL generalizes across tasks, objects, and motion patterns, achieving superior performance compared to prior methods on real-world manipulation tasks under both sparse and dense specifications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。