通过识别难以标注的模糊场景,高效筛选高价值数据提升路边3D检测性能。
Learnability-Driven Submodular Optimization for Active Roadside 3D Detection
- 基于可学习性筛选兼具信息量与可标注性的样本,避免无效标注。
- 仅用25%标注预算即达到全数据量86%以上的检测精度。
- 适合需要降低标注成本的自动驾驶路边感知系统研发人员。
路边感知数据集通常通过同步车辆与路边帧对的协作标注构建。然而实际部署中,受硬件与隐私限制,常需仅依赖路边数据进行标注。人类专家在缺乏车辆侧数据(图像、激光雷达)时难以准确标注,不仅增加标注难度与成本,更暴露根本性可学习性问题:许多路边单视角场景中的物体距离远、模糊或被遮挡,其3D属性从单一视角不明确,必须通过配对车辆-路边帧交叉验证才能可靠标注。我们称此类样本为固有模糊样本。为减少对这类样本的无效标注投入,同时保持模型高性能,本文提出一种学习驱动的主动学习框架,旨在选择既具信息量又可可靠标注的场景,抑制固有模糊样本并保障覆盖度。实验表明,所提方法LH3D在DAIR-V2X-I数据集上仅使用25%标注预算,即实现车辆、行人、自行车分别达86.06%、67.32%、78.67%的全性能,显著优于基于不确定性的基线方法。结果证明,对于路边3D感知,可学习性比不确定性更为关键。
原文摘要 · Abstract (English)
Roadside perception datasets are typically constructed via cooperative labeling between synchronized vehicle and roadside frame pairs. However, real deployment often requires annotation of roadside-only data due to hardware and privacy constraints. Even human experts struggle to produce accurate labels without vehicle-side data (image, LIDAR), which not only increases annotation difficulty and cost, but also reveals a fundamental learnability problem: many roadside-only scenes contain distant, blurred, or occluded objects whose 3D properties are ambiguous from a single view and can only be reliably annotated by cross-checking paired vehicle--roadside frames. We refer to such cases as inherently ambiguous samples. To reduce wasted annotation effort on inherently ambiguous samples while still obtaining high-performing models, we turn to active learning. This work focuses on active learning for roadside monocular 3D object detection and proposes a learnability-driven framework that selects scenes which are both informative and reliably labelable, suppressing inherently ambiguous samples while ensuring coverage. Experiments demonstrate that our method, LH3D, achieves 86.06%, 67.32%, and 78.67% of full-performance for vehicles, pedestrians, and cyclists respectively, using only 25% of the annotation budget on DAIR-V2X-I, significantly outperforming uncertainty-based baselines. This confirms that learnability, not uncertainty, matters for roadside 3D perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。