区分危险与异常,揭示视觉语言模型误把奇怪当危险的缺陷
Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies

- 将危险与异常分离评估,避免模型误把异常当危险
- 多模型测试显示多数模型混淆异常与危险,误判率超60%
- 适合关注AI安全推理、人机交互系统的研究者阅读
现代安全关键系统越来越多依赖人机交互来降低灾难风险并支持应急决策。视觉语言模型(VLMs)因其能解析复杂场景并传递安全相关信息而具有潜力,但其安全推理可靠性仍需严格评估。当前评估常将危险识别简化为二分类(安全/不安全),难以判断模型是识别真实物理危险,还是仅对场景异常敏感。本文提出明确区分危险与异常,并分别识别两类状态。我们在两个数据集上评估多个前沿VLMs,采用多种提示策略,发现模型常将异常性误判为危险性,暴露出对上下文异常的过度依赖。进一步表明,分离评估能更清晰揭示模型安全推理能力,暴露二元判断隐藏的失效模式。数据集已公开于Roboflow:https://app.roboflow.com/vlm-in-context-anomaly-and-hazard-detection/camera-ready-roman-ds。
原文摘要 · Abstract (English)
Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and support decision-making during emergencies. Vision-Language Models (VLMs) are promising for these settings because they can interpret complex scenes and communicate safety-relevant information, but they still require careful evaluation to ensure reliable safety reasoning. In particular, current evaluations often frame danger recognition as a binary decision (Safe/Unsafe), making it unclear whether a model is identifying true physical hazards or merely reacting to unusual scene elements. We address this limitation by introducing an explicit distinction between hazard and anomaly, and by separately recognizing hazardous and anomalous states. We evaluate several state-of-the-art VLMs across two datasets and multiple prompting strategies to test whether this distinction changes model behavior. Our results show that VLMs frequently misinterpret anomalousness as hazardousness, revealing an over-reliance on contextual irregularity as a proxy for danger. We further show that explicitly separating anomaly from hazard provides a more informative evaluation of VLM safety reasoning and exposes failure modes that binary safety judgments can obscure. Our public dataset is available on Roboflow https://app.roboflow.com/vlm-in-context-anomaly-and-hazard-detection/camera-ready-roman-ds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。