给视觉语言动作模型加安全层,让机器人避障更稳、任务成功率更高。
VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer
- 用控制屏障函数设计可插拔的安全约束层,无缝接入现有模型。
- 在复杂场景中避障率提升超50%,任务成功率提高近10%。
- 适合需要高安全性的机器人操作应用,如家庭服务或工业协作。
视觉-语言-动作(VLA)模型在泛化多样机器人操作任务方面表现出色。然而,在非结构化环境中部署这些模型仍具挑战性,尤其需同时保证任务执行与安全性,防止物理交互中的潜在碰撞。本文提出名为AEGIS的视觉-语言-安全动作(VLSA)架构,包含基于控制屏障函数构建的可插拔安全约束(SC)层。AEGIS可直接集成至现有VLA模型,提升安全性并提供理论保障,同时保持原始指令遵循性能。为评估该架构有效性,我们构建了涵盖不同空间复杂度与障碍物干扰场景的安全关键基准SafeLIBERO。大量实验表明,本方法显著优于现有基线:在避障率上提升超过50%,任务成功率提升近10%。所有基准数据集、代码及补充材料已公开于https://vlsa-aegis.github.io/。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in generalizing across diverse robotic manipulation tasks. However, deploying these models in unstructured environments remains challenging due to the critical need for simultaneous task compliance and safety assurance, particularly in preventing potential collisions during physical interactions. In this work, we introduce a Vision-Language-Safe Action (VLSA) architecture, named AEGIS, which contains a plug-and-play safety constraint (SC) layer formulated via control barrier functions. AEGIS integrates directly with existing VLA models to improve safety with theoretical guarantees, while maintaining their original instruction-following performance. To evaluate the efficacy of our architecture, we construct a comprehensive safety-critical benchmark SafeLIBERO, spanning distinct manipulation scenarios characterized by varying degrees of spatial complexity and obstacle intervention. Extensive experiments demonstrate the superiority of our method over state-of-the-art baselines. Notably, AEGIS achieves over 50% improvement in obstacle avoidance rate while substantially increasing the task success rate by nearly 10%. All benchmark datasets, code, and supplementary materials are publicly available at https://vlsa-aegis.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。