arXiv:2603.18192cs.CV2026-03

构建首个从行人视角采集的微出行交通参与者检测数据集。

MicroVision: An Open Dataset and Benchmark Models for Detecting Vulnerable Road Users and Micromobility Vehicles

  • 从行人视角采集8000+张高清图像,聚焦行人的真实出行环境。
  • 标注超3万条交通参与者,包含骑行者与电动滑板车等新型交通工具。
  • 提供基准模型,检测准确率最高达72.3%,适用于城市安全监测。

微出行正成为主流交通方式,但其与弱势道路使用者(VRUs)在共享基础设施中的交互带来新的交通安全挑战。现有公开图像数据集对VRUs和微出行车辆(MMVs)缺乏足够关注,常将骑行人与行人统归为“人”,且缺少如电动滑板车等新车型,同时多从汽车视角采集,缺乏仅供行人通行区域(人行道、自行车道)的数据。为此,我们推出MicroVision数据集:一个从弱势道路使用者视角采集的开放图像数据集,涵盖哥德堡(瑞典)超过8,000张匿名化全高清图像,包含超过30,000个精心标注的VRU与MMV实例,覆盖近2,000个独特互动场景,记录周期长达一年。同时提供基于先进架构的基准目标检测模型,在未见测试集上达到最高0.723的平均精度。该数据集与模型可助力交通安全管理中区分不同类型的用户,或用于监控系统识别微出行使用情况。数据集与模型权重可通过https://doi.org/10.71870/eepz-jd52获取。

原文摘要 · Abstract (English)

Micromobility is a growing mode of transportation, raising new challenges for traffic safety and planning due to increased interactions in areas where vulnerable road users (VRUs) share the infrastructure with micromobility, including parked micromobility vehicles (MMVs). Approaches to support traffic safety and planning increasingly rely on detecting road users in images -- a computer-vision task relying heavily on the quality of the images to train on. However, existing open image datasets for training such models lack focus and diversity in VRUs and MMVs, for instance, by categorizing both pedestrians and MMV riders as "person", or by not including new MMVs like e-scooters. Furthermore, datasets are often captured from a car perspective and lack data from areas where only VRUs travel (sidewalks, cycle paths). To help close this gap, we introduce the MicroVision dataset: an open image dataset and annotations for training and evaluating models for detecting the most common VRUs (pedestrians, cyclists, e-scooterists) and stationary MMVs (bicycles, e-scooters), from a VRU perspective. The dataset, recorded in Gothenburg (Sweden), consists of more than 8,000 anonymized, full-HD images with more than 30,000 carefully annotated VRUs and MMVs, captured over an entire year and part of almost 2,000 unique interaction scenes. Along with the dataset, we provide first benchmark object-detection models based on state-of-the-art architectures, which achieved a mean average precision of up to 0.723 on an unseen test set. The dataset and model can support traffic safety to distinguish between different VRUs and MMVs, or help monitoring systems identify the use of micromobility. The dataset and model weights can be accessed at https://doi.org/10.71870/eepz-jd52.

交通感知目标检测数据集微出行

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。