从单次演示中提取任务相关空间帧,实现机器人技能的快速泛化。
TReF-6: Inferring Task-Relevant Frames from a Single Demonstration for One-Shot Skill Generalization
- 基于轨迹几何识别关键点,构建6自由度任务参考帧。
- 结合视觉语言模型与Grounded-SAM,实现跨场景语义定位。
- 支持真实世界操作任务中的一次性模仿学习。
机器人常因缺乏可迁移、可解释的空间表征,难以从单次示范中泛化技能。本文提出TReF-6方法,从单一轨迹中推断出简化的6自由度任务相关帧(Task-Relevant Frame)。该方法仅通过轨迹几何确定一个影响点作为局部坐标系原点,并以此为基础参数化动态运动基元(DMP),从而超越传统起点-终点模仿的局限。通过视觉语言模型赋予该帧语义意义,并利用Grounded-SAM在新场景中精确定位,实现功能一致的技能泛化。我们在仿真环境中验证了TReF-6对轨迹噪声的鲁棒性,并在真实世界操作任务中部署端到端流程,证明其能有效保持任务意图,在多种物体配置下实现一次示例学习的泛化能力。
原文摘要 · Abstract (English)
Robots often struggle to generalize from a single demonstration due to the lack of a transferable and interpretable spatial representation. In this work, we introduce TReF-6, a method that infers a simplified, abstracted 6DoF Task-Relevant Frame from a single trajectory. Our approach identifies an influence point purely from the trajectory geometry to define the origin for a local frame, which serves as a reference for parameterizing a Dynamic Movement Primitive (DMP). This influence point captures the task's spatial structure, extending the standard DMP formulation beyond start-goal imitation. The inferred frame is semantically grounded via a vision-language model and localized in novel scenes by Grounded-SAM, enabling functionally consistent skill generalization. We validate TReF-6 in simulation and demonstrate robustness to trajectory noise. We further deploy an end-to-end pipeline on real-world manipulation tasks, showing that TReF-6 supports one-shot imitation learning that preserves task intent across diverse object configurations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。