用视频分析预测危险等级,提升安防系统响应能力
VARS: Vision-based Assessment of Risk in Security Systems
- 基于100段50帧视频,用人评危险分构建评估模型
- Transformer模型在准确率和MAE上表现最优
- 适合需要实时风险预警的安防场景应用
准确预测视频内容中的危险等级对提升安全与安防系统至关重要,尤其在需快速可靠判断的环境中。本研究在自建数据集上开展对比分析,该数据集包含100段视频,每段50帧,经人工标注危险评分(0-10分),并划分为无警报(<7)与高警报(≥7)两类。评估涵盖支持向量机、神经网络及基于Transformer的模型,使用准确率、F1分数和平均绝对误差(MAE)等标准指标进行比较,旨在识别最鲁棒的危险评估方法。本研究推动了视频驱动风险检测框架的准确性与泛化能力发展。
原文摘要 · Abstract (English)
The accurate prediction of danger levels in video content is critical for enhancing safety and security systems, particularly in environments where quick and reliable assessments are essential. In this study, we perform a comparative analysis of various machine learning and deep learning models to predict danger ratings in a custom dataset of 100 videos, each containing 50 frames, annotated with human-rated danger scores ranging from 0 to 10. The danger ratings are further classified into three categories: no alert (less than 7)and high alert (greater than equal to 7). Our evaluation covers classical machine learning models, such as Support Vector Machines, as well as Neural Networks, and transformer-based models. Model performance is assessed using standard metrics such as accuracy, F1-score, and mean absolute error (MAE), and the results are compared to identify the most robust approach. This research contributes to developing a more accurate and generalizable danger assessment framework for video-based risk detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。