首个面向行车记录仪的路旁垃圾检测数据集,解决小物体、低密度难题。
RoLID-11K: A Dashcam Dataset for Small-Object Roadside Litter Detection
- 构建11,000+张真实驾驶场景图像,覆盖英国多样路况。
- 变压器模型定位最准,但实时模型因特征层次粗受限。
- 适合做小物体检测与智能环卫系统研发的科研人员使用。
路旁垃圾带来环境、安全和经济挑战,但现有监测依赖人工调查和公众报告,空间覆盖有限。现有视觉数据集聚焦街景静态图像、航拍或水下环境,未能反映行车记录仪画面中垃圾极小、稀疏且嵌入复杂路边背景的独特特性。本文提出RoLID-11K,首个基于行车记录仪的路旁垃圾检测大规模数据集,包含超过11,000张标注图像,涵盖多样英国驾驶条件,呈现显著长尾分布与小物体特征。我们对多种现代检测器进行基准测试,从注重精度的变换器架构到实时YOLO模型,分析其在该任务中的优劣。结果表明,尽管CO-DETR等变换器实现最佳定位精度,实时模型仍受限于粗粒度特征层级。RoLID-11K为动态驾驶场景中的极端小物体检测建立挑战性基准,旨在推动可扩展、低成本的路旁垃圾监测系统发展。数据集已开源:https://github.com/xq141839/RoLID-11K。
原文摘要 · Abstract (English)
Roadside litter poses environmental, safety and economic challenges, yet current monitoring relies on labour-intensive surveys and public reporting, providing limited spatial coverage. Existing vision datasets for litter detection focus on street-level still images, aerial scenes or aquatic environments, and do not reflect the unique characteristics of dashcam footage, where litter appears extremely small, sparse and embedded in cluttered road-verge backgrounds. We introduce RoLID-11K, the first large-scale dataset for roadside litter detection from dashcams, comprising over 11k annotated images spanning diverse UK driving conditions and exhibiting pronounced long-tail and small-object distributions. We benchmark a broad spectrum of modern detectors, from accuracy-oriented transformer architectures to real-time YOLO models, and analyse their strengths and limitations on this challenging task. Our results show that while CO-DETR and related transformers achieve the best localisation accuracy, real-time models remain constrained by coarse feature hierarchies. RoLID-11K establishes a challenging benchmark for extreme small-object detection in dynamic driving scenes and aims to support the development of scalable, low-cost systems for roadside-litter monitoring. The dataset is available at https://github.com/xq141839/RoLID-11K.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。