构建首个视觉-语言-动作模型的安全评估基准,验证其在复杂场景下的安全性。
LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models

- 通过关键帧驱动生成大量安全示范数据,提升数据效率。
- 发现高多样性训练虽提升安全性,但任务成功率仍受限于轨迹生成与语义对齐。
- 适合研究机器人安全、具身智能和视觉-语言-动作模型的开发者使用。
尽管视觉-语言-动作(VLA)模型具备出色的操控能力,但在严格约束下的操作安全性仍缺乏验证。为此,我们提出一个参数化安全基准,可程序化生成具有全面随机性的安全关键场景。为克服人工遥操作的可扩展性瓶颈,我们开发了一种新型关键帧驱动的数据生成流程。基于此框架,我们构建了包含19,664条严格无碰撞示范的大规模数据集,并进行了广泛域随机化。随后,我们对八种VLA模型和两种具身基础模型进行了系统性跨范式评估。分析揭示出关键的泛化-安全权衡:虽然高多样性训练能促进更安全的轨迹,但任务成功率仍受制于次优的轨迹合成与语义错位。通过提供可扩展的数据生成管道、可靠的数据集及深层失败模式洞察,LIBERO-Safety为开发安全可靠的VLA模型奠定了重要基础。
原文摘要 · Abstract (English)
Despite the impressive manipulation capabilities of Vision-Language-Action (VLA) models, their operational safety under strict constraints remains largely unverified. To address this, we introduce a parametric safety benchmark to procedurally generate safety-critical scenarios with comprehensive stochasticity. To overcome the scalability bottlenecks of human teleoperation, we develop a novel keypose-driven data generation pipeline. Leveraging this infrastructure, we curate a large-scale dataset of 19,664 strictly collision-free demonstrations with extensive domain randomization. We then conduct a systematic cross-paradigm evaluation of eight VLA and two embodied foundation models. Our analysis reveals a critical generalization-safety tension: although high-diversity training fosters safer trajectories, task success remains fundamentally bottlenecked by sub-optimal trajectory synthesis and semantic misalignment. By providing a scalable pipeline, a robust dataset, and profound failure-mode insights, LIBERO-Safety establishes a crucial foundation for developing safe and reliable VLA models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。