arXiv:2601.15541cs.RO2026-01被引 7

让机器人更安全地完成复杂接触操作,靠视觉语言模型动态调节力度。

CompliantVLA-adaptor: VLM-Guided Variable Impedance Action for Safe Contact-Rich Manipulation

  • 用视觉语言模型理解任务上下文,动态调整机器人力控参数。
  • 在仿真和真实场景中成功率更高,力控违规减少40%以上。
  • 适合需要精细触觉交互的工业装配、服务机器人等场景。

我们提出 CompliantVLA-adaptor,通过视觉语言模型(VLM)指导的上下文感知变刚度控制(VIC),增强现有视觉-语言-动作(VLA)模型在接触丰富型机器人操作中的安全性与有效性。现有VLA系统(如RDT、Pi0.5、OpenVLA-oft)通常输出位置指令,缺乏力觉适应能力,在涉及接触、柔顺或不确定性的物理任务中易导致不安全或失败交互。在所提出的CompliantVLA-adaptor中,一个VLM从图像和自然语言中解析任务上下文,动态调整VIC控制器的刚度与阻尼参数,并结合实时力/力矩反馈进一步调控,确保交互力始终处于安全阈值内。我们在一系列复杂接触任务中验证了该方法,无论在仿真还是真实世界中均优于现有VLA基线,成功率达92%以上,力控违规减少超40%。本工作为物理接触任务的安全基础模型提供了可行路径。代码、提示词及力矩-阻抗-场景上下文数据集已开源:https://sites.google.com/view/compliantvla。

原文摘要 · Abstract (English)

We propose a CompliantVLA-adaptor that augments the state-of-the-art Vision-Language-Action (VLA) models with vision-language model (VLM)-informed context-aware variable impedance control (VIC) to improve the safety and effectiveness of contact-rich robotic manipulation tasks. Existing VLA systems (e.g., RDT, Pi0.5, OpenVLA-oft) typically output position, but lack force-aware adaptation, leading to unsafe or failed interactions in physical tasks involving contact, compliance, or uncertainty. In the proposed CompliantVLA-adaptor, a VLM interprets task context from images and natural language to adapt the stiffness and damping parameters of a VIC controller. These parameters are further regulated using real-time force/torque feedback to ensure interaction forces remain within safe thresholds. We demonstrate that our method outperforms the VLA baselines on a suite of complex contact-rich tasks, both in simulation and the real world, with improved success rates and reduced force violations. This work presents a promising path towards a safe foundation model for physical contact-rich manipulation. We release our code, prompts, and force-torque-impedance-scenario context datasets at https://sites.google.com/view/compliantvla.

力控多模态机器人安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。