构建大规模监控场景检测数据集,推动真实环境目标检测算法发展
OD-VIRAT: A Large-Scale Benchmark for Object Detection in Realistic Surveillance Environments
- 提出两个监控场景检测基准OD-VIRAT Large/Tiny,覆盖10类复杂场景
- 分别包含870万和28万标注实例,涵盖小目标、遮挡等挑战性条件
- 首次系统评估多个前沿检测模型在真实监控图像中的表现
真实人类监控数据集对于在现实条件下训练和评估计算机视觉模型至关重要,有助于提升复杂环境中人及交互物体检测算法的鲁棒性。为此,我们提出了两个视觉目标检测基准:OD-VIRAT Large和OD-VIRAT Tiny,旨在推进监控图像中的视觉理解任务。两个基准的视频序列涵盖了从高处远距离拍摄的10种不同监控场景。数据集提供丰富的边界框与类别标注:OD-VIRAT Large包含599,996张图像中的870万标注实例,OD-VIRAT Tiny包含19,860张图像中的288,901个标注实例。本工作还针对RETMDet、YOLOX、RetinaNet、DETR和Deformable-DETR等先进目标检测架构在该特定版本VIRAT数据集上进行了基准测试。据我们所知,这是首个在复杂背景、遮挡物和小尺度目标等挑战条件下,系统评估这些近期发布的顶尖检测模型性能的工作。提出的基准与实验设置将为理解选定模型的表现提供洞见,并为开发更高效、更鲁棒的目标检测架构奠定基础。
原文摘要 · Abstract (English)
Realistic human surveillance datasets are crucial for training and evaluating computer vision models under real-world conditions, facilitating the development of robust algorithms for human and human-interacting object detection in complex environments. These datasets need to offer diverse and challenging data to enable a comprehensive assessment of model performance and the creation of more reliable surveillance systems for public safety. To this end, we present two visual object detection benchmarks named OD-VIRAT Large and OD-VIRAT Tiny, aiming at advancing visual understanding tasks in surveillance imagery. The video sequences in both benchmarks cover 10 different scenes of human surveillance recorded from significant height and distance. The proposed benchmarks offer rich annotations of bounding boxes and categories, where OD-VIRAT Large has 8.7 million annotated instances in 599,996 images and OD-VIRAT Tiny has 288,901 annotated instances in 19,860 images. This work also focuses on benchmarking state-of-the-art object detection architectures, including RETMDET, YOLOX, RetinaNet, DETR, and Deformable-DETR on this object detection-specific variant of VIRAT dataset. To the best of our knowledge, it is the first work to examine the performance of these recently published state-of-the-art object detection architectures on realistic surveillance imagery under challenging conditions such as complex backgrounds, occluded objects, and small-scale objects. The proposed benchmarking and experimental settings will help in providing insights concerning the performance of selected object detection models and set the base for developing more efficient and robust object detection architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。