arXiv:2604.12447cs.RO2026-04被引 3

提出HazardArena基准,评估视觉语言动作模型在语义风险下的安全性

HazardArena: Evaluating Semantic Safety in Vision-Language-Action Models

论文配图:HazardArena: Evaluating Semantic Safety in Vision-Language-Action Models
图 1 · 摘自论文原文
  • 构建安全/不安全孪生场景,仅语义不同但动作要求一致
  • 2000+资产、40项任务,7类真实风险,发现模型在安全训练下仍会出错
  • 无需训练的Safety Option Layer可有效降低不安全行为

视觉-语言-动作(VLA)模型从视觉-语言主干中继承丰富世界知识,并通过动作示范获得可执行技能。然而现有评估主要关注动作执行成功率,导致动作策略与视觉-语言语义松散耦合,暴露出系统性漏洞:即使动作执行正确,也可能在语义风险下引发不安全结果。为此,我们提出HazardArena,一个在受控但含风险情境下评估VLA语义安全性的基准。该基准由共享物体、布局和动作要求的成对安全/不安全场景构成,仅语义上下文不同。实验发现,仅在安全场景训练的VLA模型在对应不安全场景中常表现不安全。HazardArena包含超过2000个资产和40项风险敏感任务,覆盖7类基于机器人安全标准的真实风险。为缓解此漏洞,我们提出无需训练的Safety Option Layer,利用语义属性或视觉-语言判别器约束动作执行,在几乎不影响任务性能的前提下显著减少不安全行为。我们希望HazardArena能推动对大规模部署中语义安全评估与强化的重新思考。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models inherit rich world knowledge from vision-language backbones and acquire executable skills via action demonstrations. However, existing evaluations largely focus on action execution success, leaving action policies loosely coupled with visual-linguistic semantics. This decoupling exposes a systematic vulnerability whereby correct action execution may induce unsafe outcomes under semantic risk. To expose this vulnerability, we introduce HazardArena, a benchmark designed to evaluate semantic safety in VLAs under controlled yet risk-bearing contexts. HazardArena is constructed from safe/unsafe twin scenarios that share matched objects, layouts, and action requirements, differing only in the semantic context that determines whether an action is unsafe. We find that VLA models trained exclusively on safe scenarios often fail to behave safely when evaluated in their corresponding unsafe counterparts. HazardArena includes over 2,000 assets and 40 risk-sensitive tasks spanning 7 real-world risk categories grounded in established robotic safety standards. To mitigate this vulnerability, we propose a training-free Safety Option Layer that constrains action execution using semantic attributes or a vision-language judge, substantially reducing unsafe behaviors with minimal impact on task performance. We hope that HazardArena highlights the need to rethink how semantic safety is evaluated and enforced in VLAs as they scale toward real-world deployment.

语义安全视觉语言动作基准测试机器人安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。