用视觉语言模型+扩散动作模型,让机器人在复杂环境里巧用墙壁等外力抓取大而平的物体。
DexDiff: Towards Extrinsic Dexterity Manipulation of Ungraspable Objects in Unrestricted Environments
- 结合视觉语言模型与条件扩散动作模型,实现长时序任务规划与执行
- 仿真中成功率比基线高47%,能有效抓取未见过的物体
- 适合需要灵活利用环境外力的现实机器人操作场景
抓取大型平坦物体(如书本或煎锅)常被视为不可抓取任务,因其抓取姿态难以到达。以往方法虽利用墙壁或桌面边缘等外部辅助实现外在灵巧性,但受限于特定任务策略且缺乏任务规划来发现预抓取条件,难以适应多变环境和外力约束。为此,我们提出DexDiff,一种面向外在灵巧性的鲁棒机器人操作方法,支持长时序规划。具体而言,首先使用视觉语言模型(VLM)感知环境状态并生成高层任务计划,随后由目标条件动作扩散(GCAD)模型预测低层动作序列。该模型通过离线数据学习,以高层规划生成的累积奖励作为目标条件,从而提升动作预测能力。实验表明,该方法不仅能有效完成不可抓取任务,还具备对未见物体的泛化能力。在仿真中成功率较基线高出47%,并在真实场景中实现高效部署与操作。
原文摘要 · Abstract (English)
Grasping large and flat objects (e.g. a book or a pan) is often regarded as an ungraspable task, which poses significant challenges due to the unreachable grasping poses. Previous works leverage Extrinsic Dexterity like walls or table edges to grasp such objects. However, they are limited to task-specific policies and lack task planning to find pre-grasp conditions. This makes it difficult to adapt to various environments and extrinsic dexterity constraints. Therefore, we present DexDiff, a robust robotic manipulation method for long-horizon planning with extrinsic dexterity. Specifically, we utilize a vision-language model (VLM) to perceive the environmental state and generate high-level task plans, followed by a goal-conditioned action diffusion (GCAD) model to predict the sequence of low-level actions. This model learns the low-level policy from offline data with the cumulative reward guided by high-level planning as the goal condition, which allows for improved prediction of robot actions. Experimental results demonstrate that our method not only effectively performs ungraspable tasks but also generalizes to previously unseen objects. It outperforms baselines by a 47% higher success rate in simulation and facilitates efficient deployment and manipulation in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。