用多模态融合提升监控中可疑行为的可解释性评估
Transformer-Driven Multimodal Fusion for Explainable Suspiciousness Estimation in Visual Surveillance
- 基于改进YOLOv12和双卷积网络,识别武器、表情与姿态异常
- 通过注意力机制融合多源信息,实现高精度实时可疑度评分
- 构建大规模数据集USE50k,适合安全关键场景的智能监控应用
可疑行为评估对复杂环境中主动威胁检测和公共安全至关重要。本文提出一个包含65,500张图像的大规模标注数据集USE50k,涵盖机场、火车站、餐厅、公园等多样非受控环境,涵盖武器、火灾、人群密度、异常面部表情及异常体态等多种线索。基于此,我们设计轻量级模块化系统DeepUSEvision,集成增强版YOLOv12的可疑物体检测器、用于面部表情与体态识别的双深度卷积神经网络(DCNN-I和DCNN-II),以及基于Transformer的判别网络,自适应融合多模态输出,生成可解释的可疑度评分。大量实验表明,该框架在准确性、鲁棒性和可解释性上均优于现有方法。USE50k数据集与DeepUSEvision框架共同为智能监控与实时风险评估提供了强大且可扩展的基础。
原文摘要 · Abstract (English)
Suspiciousness estimation is critical for proactive threat detection and ensuring public safety in complex environments. This work introduces a large-scale annotated dataset, USE50k, along with a computationally efficient vision-based framework for real-time suspiciousness analysis. The USE50k dataset contains 65,500 images captured from diverse and uncontrolled environments, such as airports, railway stations, restaurants, parks, and other public areas, covering a broad spectrum of cues including weapons, fire, crowd density, abnormal facial expressions, and unusual body postures. Building on this dataset, we present DeepUSEvision, a lightweight and modular system integrating three key components, i.e., a Suspicious Object Detector based on an enhanced YOLOv12 architecture, dual Deep Convolutional Neural Networks (DCNN-I and DCNN-II) for facial expression and body-language recognition using image and landmark features, and a transformer-based Discriminator Network that adaptively fuses multimodal outputs to yield an interpretable suspiciousness score. Extensive experiments confirm the superior accuracy, robustness, and interpretability of the proposed framework compared to state-of-the-art approaches. Collectively, the USE50k dataset and the DeepUSEvision framework establish a strong and scalable foundation for intelligent surveillance and real-time risk assessment in safety-critical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。