arXiv:2511.02563cs.CV2025-11被引 4

首个印度交通场景大尺度标注数据集,提升本地化视觉模型精度

The Urban Vision Hackathon Dataset and Models: Towards Image Annotations and Accurate Vision Models for Indian Traffic

  • 从班加罗尔2800个摄像头采集2.6万张高清图像并众包标注
  • 14类本土车辆共180万框标注,30万+框通过投票与算法生成真值
  • 基于该数据训练的模型比COCO基线提升31.5%检测准确率

本报告介绍由印度科学研究所(AIM@IISc)发布的首个公开大型数据集——UVH-26,包含来自印度班加罗尔2800个安全城市摄像头在4周内采集的26,646张1080p高分辨率图像。通过一场面向全国565名大学生的众包黑客松活动,对这些图像进行了标注,共完成180万条边界框标注,覆盖14类印度特有车辆:自行车、两轮车(摩托车)、三轮车(自动三轮车)、轻型商用车、小货车、图姆波旅行车、掀背车、轿车、SUV、MUV、小型巴士、公交车、卡车及其他。其中,28.3万至31.6万条共识真值边界框和标签通过多数投票与STAPLE算法生成。我们基于该数据集训练了多种主流检测器(如YOLO11-S/X、RT-DETR-S/X、DAMO-YOLO-T/L),并在mAP50、mAP75、mAP50:95上报告性能。在共同类别(汽车、公交车、卡车)上,使用UVH-26训练的模型相比在COCO上训练的基线模型,mAP50:95提升8.4%-31.5%,其中RT-DETR-X达到0.67,优于COCO权重的0.40。这表明针对印度交通场景进行领域特定训练的有效性。发布包包含基于多数投票(UVH-26-MV)和STAPLE(UVH-26-ST)生成的2.6万张图像共识标注,以及在每种标注数据上微调的6个YOLO和DETR模型。该数据集直接从实际交通摄像头流中捕捉印度城市交通的多样性,填补了现有全球基准的空白,为发展新兴国家复杂交通条件下的智能交通系统提供了基础。

原文摘要 · Abstract (English)

This report describes the UVH-26 dataset, the first public release by AIM@IISc of a large-scale dataset of annotated traffic-camera images from India. The dataset comprises 26,646 high-resolution (1080p) images sampled from 2800 Bengaluru's Safe-City CCTV cameras over a 4-week period, and subsequently annotated through a crowdsourced hackathon involving 565 college students from across India. In total, 1.8 million bounding boxes were labeled across 14 vehicle classes specific to India: Cycle, 2-Wheeler (Motorcycle), 3-Wheeler (Auto-rickshaw), LCV (Light Commercial Vehicles), Van, Tempo-traveller, Hatchback, Sedan, SUV, MUV, Mini-bus, Bus, Truck and Other. Of these, 283k-316k consensus ground truth bounding boxes and labels were derived for distinct objects in the 26k images using Majority Voting and STAPLE algorithms. Further, we train multiple contemporary detectors, including YOLO11-S/X, RT-DETR-S/X, and DAMO-YOLO-T/L using these datasets, and report accuracy based on mAP50, mAP75 and mAP50:95. Models trained on UVH-26 achieve 8.4-31.5% improvements in mAP50:95 over equivalent baseline models trained on COCO dataset, with RT-DETR-X showing the best performance at 0.67 (mAP50:95) as compared to 0.40 for COCO-trained weights for common classes (Car, Bus, and Truck). This demonstrates the benefits of domain-specific training data for Indian traffic scenarios. The release package provides the 26k images with consensus annotations based on Majority Voting (UVH-26-MV) and STAPLE (UVH-26-ST) and the 6 fine-tuned YOLO and DETR models on each of these datasets. By capturing the heterogeneity of Indian urban mobility directly from operational traffic-camera streams, UVH-26 addresses a critical gap in existing global benchmarks, and offers a foundation for advancing detection, classification, and deployment of intelligent transportation systems in emerging nations with complex traffic conditions.

交通检测数据集计算机视觉印度场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。