测试视觉提示的执行能力与安全风险,发现模型可能误听非指令视觉信号。
LIBERO-VIFO: Benchmarking the Capability and Safety of Visual Cue Following in Vision-Language-Action Models

- 构建八类视觉提示基准,分两部分评估模型理解与执行能力。
- 7个VLA模型在无语言指令时仍会响应视觉提示,存在安全隐患。
- 适用于关注机器人视觉安全、人机交互可靠性的研究者。
视觉提示正被广泛用于引导机器人学习,但现有视觉-语言-动作(VLA)模型是否能可靠遵循授权提示而忽略未授权提示尚不明确。现有研究仅覆盖有限提示形式,且仅关注任务成功率,对提示跟随能力评估粗糙。将所有视觉提示视为授权也忽视了未授权跟随的安全风险。为此,我们提出LIBERO-VIFO基准,系统评估VLA模型在视觉提示跟随中的能力与安全性。该基准定义了八类视觉提示,包含两部分共四个协议:第一部分测试提示理解与授权跟随能力,第二部分在语言-提示冲突和无语言条件下评估未授权提示的跟随行为。对七个VLA模型的评估表明,尽管具备提示理解能力,但当前模型在无语言指令时仍可执行由视觉提示指示的任务,暴露了未经授权视觉提示跟随的新兴风险。在场景实例化提示、安全关键设置及真实机器人部署中的扩展实验进一步验证了上述发现。LIBERO-VIFO首次实现对视觉提示跟随能力与安全性的系统化评估,为VLA社区引入以视觉为中心的安全新视角。
原文摘要 · Abstract (English)
Visual cues are increasingly adopted to guide robot learning, but whether Vision-Language-Action (VLA) models can reliably follow authorized cues while disregarding unauthorized ones remains unclear. Existing work covers only a narrow range of cue forms and focuses on final task success, providing only a coarse assessment of cue-following capability. Treating all visual cues as authorized also leaves safety risks of unauthorized following unexplored. To address these gaps, we introduce LIBERO-VIFO, a benchmark to evaluate both the capability and safety of visual cue following in VLA models. LIBERO-VIFO defines eight visual cue families spanning diverse forms. A total of four protocols in two parts are defined: Part I tests cue understanding and authorized following, while Part II evaluates unauthorized visual cue following under language-cue conflict and empty language conditions. Evaluating seven VLA models reveals that although visual cue understanding does not reliably translate into execution, current VLAs are able to execute cue-indicated tasks without language instruction, exposing an emerging risk of unauthorized visual cue following. Extended experiments on scene-instantiated cues, safety-critical settings, and real-robot deployment corroborate these findings. LIBERO-VIFO brings both the capability and safety of visual cue following into systematic evaluation, establishing visual-centric safety as a new perspective for the VLA community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。