arXiv:2605.23304cs.CV2026-05

用语言规则替代图像样本,实现跨场景的通用危险检测。

General Hazard Detection

论文配图:General Hazard Detection
图 1 · 摘自论文原文
  • 以法规文本定义安全规则,摆脱对标注图像的依赖。
  • 构建3006张多场景图像数据集,支持跨领域合规评估。
  • 结合大模型与人工反馈,提升对抽象安全概念的识别能力。

危险作为抽象概念,通常依赖认知层面的逻辑推理而非具体实例。现有危险检测系统依赖预定义类别和大量标注样本,面临三重挑战:训练数据噪声多且稀疏、安全定义随上下文动态变化、难以泛化到未见场景。为此,我们提出CompliVision数据集——首个面向规则驱动合规评估的通用危险检测数据集,并构建基准评估框架。核心创新在于将危险概念从图像示例中解耦,通过语言规则表达安全要求。基于权威领域法规与ISO标准,定义多领域多样化的危险概念。CompliVision包含3,006张覆盖交通、建筑与仓库环境的图像,每张图像均标注对特定安全规则的合规性,并附自然语言解释说明视觉证据。为实现鲁棒泛化,开发主动学习框架,有效引导并优化视觉-语言模型进行合规评估。尽管先进视觉语言模型表现强劲,仍难以准确处理细粒度、上下文依赖的安全判断。我们提出的通用危险检测框架融合基于LLaVA的视觉推理与人机协同反馈,显著提升评估精度。

原文摘要 · Abstract (English)

Hazard, as an abstract concept, is typically defined through cognitive-level logical reasoning rather than concrete examples. In contrast, existing hazard detection systems rely on predefined hazard categories and require intensive collection of labelled examples within detection or classification architectures. This approach faces three fundamental challenges when addressing abstract safety concepts: (1) noisy and sparse training data, (2) dynamically evolving definitions that change across contexts and time, and (3) limited generalisation to unseen or novel scenarios. To address these limitations, we present the CompliVision dataset, the first general-purpose hazard dataset designed for rule-based compliance assessment, along with a baseline framework for hazard evaluation. Our key innovation is decoupling the hazard concept from image-based examples by expressing safety requirements through language-based rules. We ground our approach in authoritative domain regulations and ISO standards to define diverse hazard concepts across multiple domains. The CompliVision dataset comprises 3,006 images spanning traffic, construction, and warehouse environments, with each image annotated for compliance against specific safety rules, accompanied by natural language explanations highlighting the supporting visual evidence. To achieve robust generalisation, we develop an active learning framework to more effectively guide and refine vision-language models in assessing hazard compliance. While state-of-the-art VLMs demonstrate strong capabilities, they struggle with the fine-grained, context-dependent interpretation required for accurate safety assessment. We proposed a general hazard detection framework to address this limitation which combines LLaVA-based visual reasoning with with human-in-the-loop feedback.

危险检测视觉语言模型主动学习合规评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。