arXiv:2510.17369cs.ROcs.AI2025-10中稿 · NeurIPS被引 1

将视觉语言动作模型部署到软体机械臂,实现安全人机交互。

Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots

  • 针对软体机械臂特性,对VLA模型进行针对性微调。
  • 微调后软臂性能达到刚性机械臂水平,突破部署瓶颈。
  • 适合关注柔性机器人与通用控制融合的研究者。

机器人系统正面临在以人为中心的非结构化环境中运行的需求,安全性、适应性和泛化能力至关重要。视觉语言动作(VLA)模型被视为一种由语言引导的通用控制框架,适用于真实机器人。然而,其部署仍局限于传统串联机械臂。由于刚性结构和基于学习的控制不可预测性,与环境的安全交互能力缺失且关键。本文首次将VLA模型部署于软体连续体机械臂,实现自主安全的人机交互。我们构建了系统化的微调与部署流程,评估了两种先进VLA模型(OpenVLA-OFT 和 $π_0$)在典型操作任务中的表现。结果显示,未经微调的原生策略因本体差异失效,但通过针对性微调,软体机械臂性能可媲美刚性对手。研究证明,微调是弥合本体差距的关键,且将VLA模型与软体机器人结合,可在人共享环境中实现安全灵活的具身智能。

原文摘要 · Abstract (English)

Robotic systems are increasingly expected to operate in human-centered, unstructured environments where safety, adaptability, and generalization are essential. Vision-Language-Action (VLA) models have been proposed as a language guided generalized control framework for real robots. However, their deployment has been limited to conventional serial link manipulators. Coupled by their rigidity and unpredictability of learning based control, the ability to safely interact with the environment is missing yet critical. In this work, we present the deployment of a VLA model on a soft continuum manipulator to demonstrate autonomous safe human-robot interaction. We present a structured finetuning and deployment pipeline evaluating two state-of-the-art VLA models (OpenVLA-OFT and $π_0$) across representative manipulation tasks, and show while out-of-the-box policies fail due to embodiment mismatch, through targeted finetuning the soft robot performs equally to the rigid counterpart. Our findings highlight the necessity of finetuning for bridging embodiment gaps, and demonstrate that coupling VLA models with soft robots enables safe and flexible embodied AI in human-shared environments.

软体机器人VLA模型具身智能人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。