发现视觉语言模型在紧急场景识别中过度反应,误判大量安全情况为危险。
Better Safe Than Sorry? Overreaction Problem of Vision Language Models in Visual Emergency Recognition
- 构建对比图像数据集,测试模型对真实与安全场景的区分能力。
- 模型召回率高但精确率低,31%-96%的安全场景被误判为危险。
- 所有模型均错判7个安全场景,反映系统性过激倾向。
视觉语言模型(VLMs)虽能解读视觉内容,但在安全关键场景中的可靠性仍不足。本文提出VERI诊断基准,包含200张合成图像(100组对比对)和50张真实世界图像(25组对比对),每组紧急场景均配有一张视觉相似但安全的对照图像,经人工验证。通过两阶段评估协议(风险识别与应急响应),在医疗紧急、事故及自然灾害场景中测试17个VLMs。分析显示存在“过度反应问题”:模型召回率达70%-100%,但精确率低下,31%-96%的安全情境被错误分类为危险。7个安全场景被所有模型误判。这种“宁可安全也不漏判”的偏差源于上下文过度解读(88%-98%错误由此导致)。合成与真实数据均证实该系统性模式,挑战了VLM在安全关键应用中的可靠性。解决此问题需增强模型在模糊视觉情境下的上下文推理能力。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have shown capabilities in interpreting visual content, but their reliability in safety-critical scenarios remains insufficiently explored. We introduce VERI, a diagnostic benchmark comprising 200 synthetic images (100 contrastive pairs) and an additional 50 real-world images (25 pairs) for validation. Each emergency scene is paired with a visually similar but safe counterpart through human verification. Using a two-stage evaluation protocol (risk identification and emergency response), we assess 17 VLMs across medical emergencies, accidents, and natural disasters. Our analysis reveals an "overreaction problem": models achieve high recall (70-100%) but suffer from low precision, misclassifying 31-96% of safe situations as dangerous. Seven safe scenarios were universally misclassified by all models. This "better-safe-than-sorry" bias stems from contextual overinterpretation (88-98% of errors). Both synthetic and real-world datasets confirm these systematic patterns, challenging VLM reliability in safety-critical applications. Addressing this requires enhanced contextual reasoning in ambiguous visual situations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。