测试机器人在相同场景下对指令的敏感度,发现模型常误判危险指令。
GuardianBench: A Same-Scene Instruction-Contrastive Benchmark for Latent Contextual Risk in Embodied AI

- 构建3024组同场景正反指令对,检测隐性安全风险
- 主流模型平均准确率仅24.1%,对指令变化不敏感
- 提出轻量级监督方法,显著提升模型安全判断能力
在具身智能中,安全风险可能是隐性的:一个看似无害的指令与安全场景组合后可能变得危险。现有研究多关注视觉环境变化或执行过程动态,而固定场景仅改变指令这一维度仍待深入。本文提出GuardianBench,一个基于国际安全标准的指令对比基准,通过3,024个同场景安全/非安全对比对,覆盖多种风险类别,系统评估模型对指令-场景组合的敏感性。基准测试显示,主流视觉语言模型(VLMs)存在指令不敏感问题:在相同场景下对安全与危险指令均给出高通过率;主要模型平均配对准确率仅为24.1%。系统性原因分析揭示,模型未能有效关联区分安全与危险组合的关键指令线索。作为后训练案例,轻量级判据级监督方法Verdict Log-Odds Supervision(VLOS)显著提升了开放权重骨干模型性能。本研究通过隐性上下文风险任务设计、标准驱动的对比基准构建、配对级与推理级失败诊断,以及基准支持的判据校准,确立GuardianBench为具身智能中检测并改进指令-场景组合安全推理的可控评估体系。
原文摘要 · Abstract (English)
In embodied AI, safety risk can be latent: a benign instruction and a safe scene become hazardous only when composed. Prior work has advanced embodied safety by varying visual contexts or evaluating execution-time dynamics, but the complementary axis of fixing the scene and varying only the instruction remains underexplored. We introduce GuardianBench, an instruction-contrastive benchmark grounded in international safety standards that isolates this latent contextual risk through 3,024 instruction-scene examples organized as same-scene Safe/Unsafe contrastive pairs across various hazard categories. Benchmarking state-of-the-art vision-language models (VLMs) reveals instruction-insensitive verdicts: models disproportionately approve both instructions under a given scene; across the primary models, average pair accuracy is only 24.1%. Our systematic rationale audit localizes the dominant failure: models fail to bind the instruction-relevant cues that differentiate safe from unsafe compositions. As a post-training case study, Verdict Log-Odds Supervision (VLOS), a lightweight verdict-level objective, substantially improves performance on open-weight backbones. Together, our latent contextual risk task formulation, standards-grounded contrastive benchmark construction, pair-level and rationale-level failure diagnosis, and benchmark-enabled verdict calibration establish GuardianBench as a controlled evaluation suite for exposing and improving safety reasoning over instruction-scene compositions under latent contextual risk.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。