首个真实世界视觉异常检测基准,评估模型理解日常异常能力。
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
- 构建真实场景异常数据集,支持描述、解释、论证三任务
- 发现顶尖视觉语言模型在常识推理上表现不佳
- 适合研究异常检测与人机认知对齐的学者
人类能自然识别、推理并解释环境中的异常。在计算机视觉领域,这一长期挑战仍局限于工业缺陷或人为生成的异常,难以捕捉真实世界中复杂多变的异常现象。本文提出CAVE——首个真实世界视觉异常基准。该数据集支持三个开放式任务:异常描述、解释与论证,并提供细粒度标注,涵盖异常的视觉表现、复杂度、严重性及常见程度。标注基于认知科学中人类识别与解决异常的机制,为视觉-语言模型(VLMs)在异常感知与理解方面的评估提供全面框架。实验表明,即使采用先进提示策略,当前最先进的VLMs在视觉异常感知和常识推理方面仍表现不佳。CAVE作为一个真实且认知基础的基准,将推动异常检测与常识推理研究的发展。
原文摘要 · Abstract (English)
Humans can naturally identify, reason about, and explain anomalies in their environment. In computer vision, this long-standing challenge remains limited to industrial defects or unrealistic, synthetically generated anomalies, failing to capture the richness and unpredictability of real-world anomalies. In this work, we introduce CAVE, the first benchmark of real-world visual anomalies. CAVE supports three open-ended tasks: anomaly description, explanation, and justification; with fine-grained annotations for visual grounding and categorizing anomalies based on their visual manifestations, their complexity, severity, and commonness. These annotations draw inspiration from cognitive science research on how humans identify and resolve anomalies, providing a comprehensive framework for evaluating Vision-Language Models (VLMs) in detecting and understanding anomalies. We show that state-of-the-art VLMs struggle with visual anomaly perception and commonsense reasoning, even with advanced prompting strategies. By offering a realistic and cognitively grounded benchmark, CAVE serves as a valuable resource for advancing research in anomaly detection and commonsense reasoning in VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。