无需标注数据,让机器人学会在真实环境中泛化操作物体。
UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
- 用大模型自动标注海量指令-视觉操作关系数据
- 仅用10次演示就能泛化到新物体和任务
- 适合做开放任务的机器人操作研究
理解物体的细粒度操作属性对机器人在非结构化环境中执行开放式任务指令至关重要。现有视觉操作预测方法通常依赖人工标注数据或仅适用于预设任务集。我们提出UAD(无监督操作属性蒸馏),从基础模型中蒸馏操作知识到一个任务条件化的操作模型,无需任何人工标注。通过结合大型视觉模型与视觉语言模型的优势,UAD自动构建大规模的《指令,视觉操作》配对数据集。仅在冻结特征基础上训练轻量级任务条件解码器,UAD在真实场景和多种人类活动下表现出显著泛化能力,尽管训练仅使用仿真中的渲染物体。以UAD提供的操作信息作为观测空间,我们展示的模仿学习策略在仅训练10次演示后,即可泛化至未见过的物体实例、物体类别以及任务指令的变化。
原文摘要 · Abstract (English)
Understanding fine-grained object affordances is imperative for robots to manipulate objects in unstructured environments given open-ended task instructions. However, existing methods of visual affordance predictions often rely on manually annotated data or conditions only on a predefined set of tasks. We introduce UAD (Unsupervised Affordance Distillation), a method for distilling affordance knowledge from foundation models into a task-conditioned affordance model without any manual annotations. By leveraging the complementary strengths of large vision models and vision-language models, UAD automatically annotates a large-scale dataset with detailed $<$instruction, visual affordance$>$ pairs. Training only a lightweight task-conditioned decoder atop frozen features, UAD exhibits notable generalization to in-the-wild robotic scenes and to various human activities, despite only being trained on rendered objects in simulation. Using affordance provided by UAD as the observation space, we show an imitation learning policy that demonstrates promising generalization to unseen object instances, object categories, and even variations in task instructions after training on as few as 10 demonstrations. Project website: https://unsup-affordance.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。