arXiv:2605.02037cs.ROcs.AI2026-05

低成本机器人平台实现视觉语言动作模型端到端部署,可安全抓取易碎物品。

VILAS: A VLA-Integrated Low-cost Architecture with Soft Grasping for Robotic Manipulation

论文配图:VILAS: A VLA-Integrated Low-cost Architecture with Soft Grasping for Robotic Manipulation
图 1 · 摘自论文原文
  • 用模块化低成本硬件集成视觉语言动作模型,支持从数据采集到部署全流程
  • 在葡萄抓取任务中成功部署3个主流VLA模型,验证系统有效性
  • 设计折纸结构软夹爪,无需力传感器即可实现温和、重复的接触

我们提出VILAS,一个完全低成本、模块化的机器人操作平台,旨在支持在廉价硬件上进行端到端视觉-语言-动作(VLA)策略学习与部署。系统集成Fairino FR5协作臂、Jodell RG52-50电动夹爪和双相机感知模块,通过基于ZMQ的通信架构统一协调遥操作、数据采集和策略部署。为在不依赖显式力传感的情况下安全操作易碎物体,我们设计了一种基于折纸结构的软性顺应夹爪,可在压缩载荷下产生可预测变形,实现对脆弱目标的柔和且重复的接触。我们在VILAS平台上部署并评估了三个最先进的VLA模型:pi_0、pi_0.5和GR00T N1.6。所有模型均使用相同的示范数据集,从公开发布的预训练检查点微调而来,该数据集通过我们的遥操作流程收集。在葡萄抓取任务上的实验验证了所提系统的有效性,证实了高性能操作策略可在低成本模块化硬件上成功训练与部署。研究结果进一步提供了当前VLA模型在真实场景中部署特性的实用见解。

原文摘要 · Abstract (English)

We present VILAS, a fully low-cost, modular robotic manipulation platform designed to support end-to-end vision-language-action (VLA) policy learning and deployment on accessible hardware. The system integrates a Fairino FR5 collaborative arm, a Jodell RG52-50 electric gripper, and a dual-camera perception module, unified through a ZMQ-based communication architecture that seamlessly coordinates teleoperation, data collection, and policy deployment within a single framework. To enable safe manipulation of fragile objects without relying on explicit force sensing, we design a kirigami-based soft compliant gripper extension that induces predictable deformation under compressive loading, providing gentle and repeatable contact with delicate targets. We deploy and evaluate three state-of-the-art VLA models on the VILAS platform: pi_0, pi_0.5, and GR00T N1.6. All models are fine-tuned from publicly released pretrained checkpoints using an identical demonstration dataset collected via our teleoperation pipeline. Experiments on a grape grasping task validate the effectiveness of the proposed system, confirming that capable manipulation policies can be successfully trained and deployed on low-cost modular hardware. Our results further provide practical insights into the deployment characteristics of current VLA models in real-world settings.

机器人操作视觉语言动作低成本系统软夹爪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。