评测视觉语言模型在真实场景中的安全判断能力,发现其对隐蔽危险反应迟钝。
EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards

- 构建1200段第一视角视频,分情境与文字误导双赛道测试
- 模型对显性危险识别率高,但对情境性危险漏判超30%
- 场景文字误导导致模型误判或过度干预,暴露推理缺陷
视觉语言模型(VLMs)被提议作为家庭和工厂中具身智能体的实时安全监护者。可部署的监护者需准确识别真正危险场景,同时避免对日常但表面惊悚的行为误干预,而传统二分类安全基准无法体现这一区别。我们提出EgoSafetyBench,一个包含1,200个机器人视角场景的自视角视频基准,以半秒粒度标注,用于评估VLM作为流式监护者的性能,涵盖两个赛道:情境赛道(800个场景)覆盖四类,从常规安全但可疑的场景到明显且具有上下文关联的危险;视觉通道赛道(400个场景)聚焦场景内文本(如标识、贴纸或标签),这些文本可能歪曲物理状态,每条误导文本均配以真实版本,测试监护者是否能识别文本误导,并判断文本是否扭曲其物理安全判断。两个赛道均采用对比阶梯设计:几乎相同的场景仅在单一可见决定性线索上不同,正确判断必须依赖该线索而非整体场景类型。我们评估了十种开源与闭源VLM。结果表明,尽管模型能可靠识别含危险的视频,却常遗漏具体危险时刻,尤其在情境性危险上表现不佳。此外,误导性场景文本会削弱所有测试模型:脆弱模型错失高达三分之一的危险,而稳健模型则对安全内容过度干预。对照实验显示,看似鲁棒的安全性常源于盲目报警而非真正的物理推理。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are now proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genuinely unsafe situations while avoiding unnecessary intervention on routine but superficially alarming activity, a distinction that binary safety benchmarks obscure. We introduce EgoSafetyBench, an egocentric video benchmark of 1,200 robot-view scenarios annotated at half-second granularity, to evaluate VLMs as streaming guards across two tracks. The situational track (800 scenarios) spans four families, from routine and safe-but-suspicious scenes to obvious and contextual hazards. The visual-channel track (400 scenarios) targets in-scene text-a sign, sticker, or label visible in the scene-that can misrepresent the physical situation, pairing each misleading sign with a truthful version to test both whether a guard flags the text as misleading and whether the text corrupts its physical-safety judgment. Both tracks use contrastive ladders: near-identical scenarios differing only in a single visible deciding cue, so a correct call must hinge on that cue rather than the overall scene type. We evaluate ten open- and closed-source VLMs. We find that while guards reliably recognize videos containing hazards, they often miss specific hazardous moments, particularly contextual hazards. Furthermore, misleading in-scene signs degrade all tested guards: vulnerable models miss up to a third of hazards, while robust models over-intervene on safe content. Matched controls reveal that apparent safety robustness often reflects indiscriminate alarming rather than true physical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。