arXiv:2603.03278cs.ROcs.AI2026-03被引 2

用少量示范实现机器人自主交互,自动积累高质量数据。

Tether: Autonomous Functional Play with Correspondence-Driven Trajectory Warping

  • 基于场景语义关键点对齐,将少量示范动作映射到新环境。
  • 实现在真实家庭场景中持续数小时的自主多任务交互,生成超1000条专家级轨迹。
  • 适合研究机器人自主学习与少样本模仿的开发者,无需人工干预。

机器人自主交互与经验学习是解决传统人工演示成本高的关键挑战。本文提出Tether方法,实现结构化、任务导向的自主功能式玩耍。首先,设计一种新型开环策略,通过将目标场景中的语义关键点与不超过10个源示范动作对齐,实现动作的轨迹扭曲,具备极强的数据效率和对空间及语义变化的鲁棒性。其次,结合视觉语言模型的视觉理解能力,构建任务选择、执行、评估与优化的闭环流程,实现真实世界中的持续自主玩耍。在类家庭多物体环境中,该方法首次仅凭少量示范即实现数十小时的自主多任务玩耍,生成多样化高质量数据集。这些数据持续提升闭环模仿策略性能,最终获得超过1000条专家级轨迹,训练出的策略可媲美人类示范数据训练的模型。

原文摘要 · Abstract (English)

The ability to conduct and learn from interaction and experience is a central challenge in robotics, offering a scalable alternative to labor-intensive human demonstrations. However, realizing such "play" requires (1) a policy robust to diverse, potentially out-of-distribution environment states, and (2) a procedure that continuously produces useful robot experience. To address these challenges, we introduce Tether, a method for autonomous functional play involving structured, task-directed interactions. First, we design a novel open-loop policy that warps actions from a small set of source demonstrations (<=10) by anchoring them to semantic keypoint correspondences in the target scene. We show that this design is extremely data-efficient and robust even under significant spatial and semantic variations. Second, we deploy this policy for autonomous functional play in the real world via a continuous cycle of task selection, execution, evaluation, and improvement, guided by the visual understanding capabilities of vision-language models. This procedure generates diverse, high-quality datasets with minimal human intervention. In a household-like multi-object setup, our method is the first to perform many hours of autonomous multi-task play in the real world starting from only a handful of demonstrations. This produces a stream of data that consistently improves the performance of closed-loop imitation policies over time, ultimately yielding over 1000 expert-level trajectories and training policies competitive with those learned from human-collected demonstrations.

机器人自主学习少样本模仿视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。