用政府数据自动生成21万张街景图标注,提升盲道检测准确率。
RampNet: A Two-Stage Pipeline for Bootstrapping Curb Ramp Detection in Streetscape Images from Open Government Metadata
- 通过政府数据自动转换坐标,生成大规模街景标注
- 检测模型达0.9236 AP,精度94.0%,召回92.5%
- 适合城市无障碍建设与计算机视觉研究者
盲道对城市无障碍至关重要,但因缺乏大规模高质量数据集,其图像检测仍具挑战。现有工作依赖众包或人工标注,常受限于质量或规模。本文提出两阶段管道RampNet,以扩展盲道检测数据集并提升模型性能。第一阶段通过自动转换政府提供的盲道位置数据至全景图像像素坐标,生成超过21万张标注的Google街景(GSV)图像。第二阶段基于该数据集训练改进的ConvNeXt V2检测模型,达到当前最优表现。通过与人工标注图像对比评估,生成数据集实现94.0%精度和92.5%召回率,检测模型达0.9236 AP,显著超越已有方法。本工作首次提供大规模、高质量的盲道检测数据集、基准及模型。
原文摘要 · Abstract (English)
Curb ramps are critical for urban accessibility, but robustly detecting them in images remains an open problem due to the lack of large-scale, high-quality datasets. While prior work has attempted to improve data availability with crowdsourced or manually labeled data, these efforts often fall short in either quality or scale. In this paper, we introduce and evaluate a two-stage pipeline called RampNet to scale curb ramp detection datasets and improve model performance. In Stage 1, we generate a dataset of more than 210,000 annotated Google Street View (GSV) panoramas by auto-translating government-provided curb ramp location data to pixel coordinates in panoramic images. In Stage 2, we train a curb ramp detection model (modified ConvNeXt V2) from the generated dataset, achieving state-of-the-art performance. To evaluate both stages of our pipeline, we compare to manually labeled panoramas. Our generated dataset achieves 94.0% precision and 92.5% recall, and our detection model reaches 0.9236 AP -- far exceeding prior work. Our work contributes the first large-scale, high-quality curb ramp detection dataset, benchmark, and model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。