arXiv:2503.09423cs.ROcs.CV2025-03中稿 · CoRL被引 5

用动作先验对齐视觉语言模型,提升复杂场景下语言控制抓取的效率与泛化能力。

Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

  • 通过一个注意力层对齐无条件动作先验与3D视觉语言先验。
  • 在仿真和真实世界中实现更高成功率,且步骤更少,支持未见物体与指令。
  • 共享抓取与放置策略,适合需要多模态动作适应的机器人任务。

我们研究在杂乱环境中基于语言指令的抓取与放置任务,即机器人需从开放杂乱中抓取目标物体并移动至指定位置。现有方法或依赖大规模数据端到端训练,或在零样本设置下组合基础模型,但存在级联误差问题;且多集中于视觉与语言基础模型,忽视动作先验的作用。本文提出A²方法,通过学习一个注意力层,将无条件动作先验与3D视觉-语言先验对齐,构建高效策略。该方法使策略在更少数据下训练的同时保留零样本泛化能力。我们发现共享抓取与放置策略能提升各自性能,并引入策略适配机制以应对动作的多模态特性。大量仿真与真实世界实验表明,本方法在杂乱场景中对抓取与放置任务均取得更高成功率,且步数更少,有效泛化至未见物体与语言指令。视频与代码见https://xukechun.github.io/papers/A2。

原文摘要 · Abstract (English)

We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation models in a zero-shot setting, suffering from cascading errors. In addition, they primarily leverage vision and language foundation models, focusing less on action priors. In this paper, we aim to develop an effective policy by integrating foundation priors from vision, language, and action. We propose A$^2$, an action prior alignment method that aligns unconditioned action priors with 3D vision-language priors by learning one attention layer. The alignment formulation enables our policy to train with less data and preserve zero-shot generalization capabilities. We show that a shared policy for both pick and place actions enhances the performance for each task, and introduce a policy adaptation scheme to accommodate the multi-modal nature of actions. Extensive experiments in simulation and the real-world show that our policy achieves higher task success rates with fewer steps for both pick and place tasks in clutter, effectively generalizing to unseen objects and language instructions. Videos and codes are available at https://xukechun.github.io/papers/A2.

机器人动作先验语言控制零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。