arXiv:2606.27079cs.RO2026-06

为视觉语言动作模型设计安全诊断基准,揭示感知与控制对安全的决定性影响。

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models

论文配图:ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 构建13类安全分类体系,从物理交互、语言指令和视觉感知三方面评估风险。
  • 提出累积安全成本和风险暴露时间指标,量化任务过程中的安全风险。
  • 发现视觉和结构变化比语言变化更严重损害安全性,适合机器人安全研究者使用。

在具身智能中,安全是机器人可靠部署于真实世界的基本前提。当前视觉-语言-动作(VLA)模型虽持续提升通用任务能力,但其具身安全边界仍不清晰。为此,我们提出ForesightSafety-VLA,一个以安全为核心评价目标的诊断基准。定义涵盖物理交互(Safe-Core)、语言指令(Safe-Lang)和感知(Safe-Vis)的13类安全分类体系,通过场景结构、语言指令和视觉观测三个可控维度评估策略,使故障原因可诊断而非被单一得分掩盖。除二元任务成功外,还引入累计安全成本(CC)与风险暴露时间(RET),并采用四象限分解实现安全/非安全成功与失败的精细区分。我们在RoboTwin中构建66个安全增强基础场景,覆盖5种具身形态,并报告代表性VLA基线结果。结果显示,即使最强策略仍存在显著安全成本与非安全成功,且结构与视觉变化导致的安全退化远超普通语言变化。表明具身安全紧密依赖感知、具身化与控制能力,无法仅通过事后安全过滤解决。

原文摘要 · Abstract (English)

In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world. Current vision-language-action (VLA) models continue to advance toward general-purpose task capability, yet their embodied safety limits remain poorly understood. To address this gap, we introduce ForesightSafety-VLA, a diagnostic benchmark that makes safety the primary evaluation target for VLA systems. We define a 13-category safety taxonomy covering physical interaction safety (Safe-Core), instruction-side safety (Safe-Lang), and perception-side safety (Safe-Vis), and evaluate policies under three controlled dimensions of variation -- scene structure, language command, and visual observation -- so that failure sources can be diagnosed rather than hidden in a single aggregate score. Beyond binary task success, ForesightSafety-VLA measures process-level risk through cumulative safety cost (CC) and risk exposure time (RET), together with a four-quadrant decomposition of safe/unsafe success and failure. We instantiate 66 safety-augmented base scenarios in RoboTwin across 5 embodiments and report results on representative VLA baselines. Across the evaluated baselines, even the strongest policy incurs non-trivial safety cost and unsafe nominal success, while structure and visual variation induce substantially stronger safety degradation than ordinary language variation. These results suggest that embodied safety is tightly coupled to perception, grounding, and control competence rather than being reducible to post-hoc safety filtering alone.

具身智能安全评估多模态模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。