用吸盘+夹爪一体设计,让机器人能擦玻璃开抽屉
VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
- 夹爪与真空吸盘集成一体,可自由切换或协同工作
- 在DexVLA和Pi0框架下成功完成多种复杂操作任务
- 低成本硬件方案开源,适合需要多功能末端执行器的研究者
视觉语言动作模型通过利用大规模预训练的视觉和语言表示,显著提升了通用机器人操作能力。现有多数VLA系统默认采用并行双指夹爪作为末端执行器,但在处理擦拭玻璃、无手柄抽屉等任务时受限于接触面积不足或缺乏附着力。为此,本文提出一种低成本、一体化硬件设计,将机械双指夹爪与真空吸盘结合,实现单一末端执行器下的双模态操作。该系统支持两种模式的灵活切换或协同使用,显著拓展了可行任务范围。我们在两个先进的VLA框架(DexVLA和Pi0)中验证了该设计的有效性与实用性。实验表明,搭载该混合末端执行器的机器人可成功完成多个传统夹爪无法实现的复杂任务。所有硬件设计与控制方案将公开发布。
原文摘要 · Abstract (English)
Vision Language Action models have significantly advanced general purpose robotic manipulation by harnessing large scale pretrained vision and language representations. Among existing approaches, a majority of current VLA systems employ parallel two finger grippers as their default end effectors. However, such grippers face inherent limitations in handling certain real world tasks such as wiping glass surfaces or opening drawers without handles due to insufficient contact area or lack of adhesion. To overcome these challenges, we present a low cost, integrated hardware design that combines a mechanical two finger gripper with a vacuum suction unit, enabling dual mode manipulation within a single end effector. Our system supports flexible switching or synergistic use of both modalities, expanding the range of feasible tasks. We validate the efficiency and practicality of our design within two state of the art VLA frameworks: DexVLA and Pi0. Experimental results demonstrate that with the proposed hybrid end effector, robots can successfully perform multiple complex tasks that are infeasible for conventional two finger grippers alone. All hardware designs and controlling systems will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。