自动标注操作关键点,实现跨机器人零样本迁移。
Keypose Exploration: Efficient Automatic Trajectory Labelling and Cross-Embodiment Policy Transfer

- 用视觉语言模型+轨迹分析自动标记关键动作点
- 生成的标签让扩散策略性能达基线水平
- 通过可达性筛选关键点,提升跨机器人迁移效果
基于关键点的操作将长时序任务分解为关键路径点,以简化策略学习。现有方法依赖任务特异性启发式或人工标注来提取演示中的关键点。本文提出一种针对抓取类任务的自动轨迹标注流程:结合视觉语言模型进行语义事件检测,与经典轨迹分析实现精准时间对齐,仅需对重复演示中的一个样本进行VLM推理。利用标注数据训练关键点引导的扩散策略(DP),通过关键点条件控制示范分布。探索了该特性在跨机器人迁移中的应用:候选关键点通过可达性图采样并筛选,引导策略向目标机器人可实现的关键点收敛。初步实验证明,标注数据生成的策略性能达到标准扩散策略基线水平;在多模态插入任务中,当存在可行候选点时,可达性过滤后的关键点条件可促进零样本迁移。
原文摘要 · Abstract (English)
Keypose-based manipulation decomposes tasks into critical waypoints to simplify policy learning for long-horizon tasks, but existing approaches rely on task-specific heuristics or manual annotation to extract keyposes from demonstrations. We present an automatic trajectory labelling pipeline for grasp-related tasks. This pipeline combines vision-language models (VLMs) for semantic event detection with classical trajectory analysis for precise temporal alignment, requiring VLM inference only on one single demo among repeating ones per task. Using the labelled data, we train a keypose-guided Diffusion Policy (DP) that exploits keypose conditioning to intervene demonstration distributions. We explore the possibility to apply this property for cross-embodiment transfer: candidate keyposes are sampled and filtered via a reachability map, steering the policy toward kinematically feasible keyposes for the target robot. As a preliminary feasibility study, experiments on two robomimic tasks show that the labelled data produces policies matching a standard DP baseline, and that reachability-filtered keypose conditioning may benefit zero-shot transfer on the multimodal insertion task when feasible candidates are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。