arXiv:2505.11680cs.RO2025-05被引 2

用视觉基础模型实现零样本技能迁移,让机器人跨物体通用操作。

Grounded Task Axes: Zero-Shot Semantic Skill Generalization via Task-Axis Controllers and Visual Foundation Models

  • 将技能拆解为基于物体关键点的可适配控制轴,实现语义级控制
  • 通过SD-DINO等模型在新物体上找到相似语义特征,完成零样本迁移
  • 实测验证了拧螺丝、倒液体、刮铲等任务的鲁棒性与泛化能力

在开放世界机器人操作中,跨物体技能迁移仍是核心挑战。泛化需兼顾不同物体的高层结构差异,同时保持底层交互控制的一致性。本文提出一种基于示例的零样本技能迁移方法,不将技能视为原子单元,而是将其分解为一系列具有优先级的接地任务轴(GTA)控制器。每个GTA控制器定义在某一轴上的可适应控制方式,如位置或力控,其接地依据是物体的关键点与轴线(例如螺钉头部相对位置或螺杆轴线)。零样本迁移通过在新目标物体上寻找语义相似的接地特征实现。我们利用视觉基础模型(如SD-DINO)检测物体间语义相似的关键点,完成示例式接地。在真实机器人上评估了拧螺丝、倒液体和刮铲等任务,验证了框架在各类任务中均具备强鲁棒性和高度泛化能力。

原文摘要 · Abstract (English)

Transferring skills between different objects remains one of the core challenges of open-world robot manipulation. Generalization needs to take into account the high-level structural differences between distinct objects while still maintaining similar low-level interaction control. In this paper, we propose an example-based zero-shot approach to skill transfer. Rather than treating skills as atomic, we decompose skills into a prioritized list of grounded task-axis (GTA) controllers. Each GTAC defines an adaptable controller, such as a position or force controller, along an axis. Importantly, the GTACs are grounded in object key points and axes, e.g., the relative position of a screw head or the axis of its shaft. Zero-shot transfer is thus achieved by finding semantically-similar grounding features on novel target objects. We achieve this example-based grounding of the skills through the use of foundation models, such as SD-DINO, that can detect semantically similar keypoints of objects. We evaluate our framework on real-robot experiments, including screwing, pouring, and spatula scraping tasks, and demonstrate robust and versatile controller transfer for each.

机器人操作零样本迁移视觉基础模型技能泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。