arXiv:2606.00515cs.ROcs.AI2026-06被引 1

让视觉语言模型安全操控高接触任务,防止错误指令导致物理失控。

PaCo-VLA: Passivity-Shielded Compliance Prior for Contact-Rich Vision-Language-Action Manipulation

论文配图:PaCo-VLA: Passivity-Shielded Compliance Prior for Contact-Rich Vision-Language-Action Manipulation
图 1 · 摘自论文原文
  • 用合规性先验替代直接控制,将模型输出转为任务级动作建议。
  • 实测在复杂插接任务中精度显著提升,对抗性扰动下零违规。
  • 适合需高安全性、强物理交互的机器人操作场景。

高接触任务需要高层次语义推理与高频接触动力学的安全调控。尽管视觉-语言-动作(VLA)模型具备前所未有的语义泛化能力,但其低频输出难以满足力敏感任务的可靠性要求。为此,我们提出PaCo-VLA,一种被动性防护的合规性先验,重构了VLA接口。不直接赋予模型电机控制权,而是将其输出视为任务级合规性提议:语义绑定、任务阶段与导纳调度。一个高频、独立于提议的被动性防护机制通过能量罐会计与边界检查,阻止无效、过时或未经验证的预测绕过底层接触物理。该解耦架构支持因果评估,可分离语义贡献与几何捷径。大量仿真与真实世界插接实验表明,PaCo-VLA在精度上优于无防护的VLA基线,在对抗性合规扰动下仍保持零被动性违规。该框架在导纳端口建立可证明的采样被动性运行契约,并为在高接触领域部署基础模型提供运行时接口。

原文摘要 · Abstract (English)

Contact-rich manipulation demands both high-level semantic reasoning and the safe regulation of high-frequency contact dynamics. While Vision-Language-Action (VLA) models provide unprecedented semantic generalization, their low-rate outputs lack the reliability required for direct plant authority in force-sensitive tasks. To bridge this semantic-to-control gap, we introduce PaCo-VLA, a passivity-shielded compliance prior that recasts the VLA interface. Rather than trusting VLAs with direct motor commands, PaCo-VLA treats network outputs as task-level compliance proposals: semantic bindings, task stages, and admittance schedules. A high-frequency, proposal-independent passivity shield governs these proposals through energy-tank accounting and boundary checks, preventing invalid, stale, or unverified model predictions from bypassing low-level contact physics. This decoupled architecture also enables causal evaluation, isolating semantic contributions from geometric shortcuts. Extensive simulated and real-world connector-insertion experiments demonstrate that PaCo-VLA achieves superior precision over unshielded VLA baselines, sustaining zero passivity violations even under adversarial compliance shifts. This framework establishes a provably sampled-passive runtime contract at the admittance port and provides a runtime interface for deploying foundation models in contact-rich domains.

机器人操控合规控制多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。