让机器人更聪明地抓东西,不同手型都能用。
Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

- 把抓取动作和任务理解分开,先生成可行抓法再结合上下文调整。
- 在多种机械手上都表现良好,真实世界实验验证了跨手型泛化能力。
- 支持空间、认知、时间三类先验,适合复杂动态环境中的智能抓取。
本文提出 AdaRoboVLG,一种任务自适应的视觉-语言-抓取框架,支持不同机械手间的通用抓取合成。与现有将基础模型与端到端抓取策略紧密耦合的方法不同,该框架学习一个高效且可泛化的基础抓取策略,通过显式的运动学映射和基于力闭合的稳定性评估生成并筛选物理上可行的抓取候选,同时将任务依赖的理解交由专用的基础模型模块完成。这些模块提供可组合的先验信息,融入抓取合成过程,实现无需重训练基础抓取策略的上下文自适应抓取。大量仿真与真实世界实验证明:(i) 基础策略具备高效学习能力和强跨手型泛化性能;(ii) 框架能有效融合空间、认知与时间先验,解决三类典型抓取挑战,性能不弱于当前最优方法;(iii) 这些先验可协同工作,在杂乱与动态环境中实现功能性抓取。结果表明,将物理抓取合成与任务理解解耦,为机器人抓取提供可扩展范式,使未来基础模型进展可直接转化为抓取能力提升,而无需重设计或重训练底层策略。补充视频见 https://adarobovlg.github.io/
原文摘要 · Abstract (English)
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient generalizable base policy that generates and evaluates physically feasible grasp candidates through explicit kinematic mapping and force-closure-based stability estimation, while offloading task-dependent understanding to specialized foundation-model modules. These modules provide composable priors that are integrated into the grasp synthesis process, enabling contextually adaptive grasp synthesis without retraining the underlying grasp policy. Through extensive simulation and real-world experiments, we demonstrate that (i) the base policy exhibits efficient learning and strong cross-hand generalization, (ii) the framework effectively incorporates spatial, cognitive, and temporal priors to address three representative grasping challenges without compromising grasp synthesis performance compared to state-of-the-art methods, and (iii) these priors can operate jointly to enable functional grasping in cluttered and dynamic environments. These results indicate that decoupling physical grasp synthesis from task-dependent understanding provides a scalable paradigm for robotic grasping, allowing future advances in foundation models to be directly translated into improved grasp capabilities without redesigning or retraining the underlying grasp policy. Supplementary videos are available at https://adarobovlg.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。