arXiv:2603.16742cs.CV2026-03

用路边传感器无监督训练自动驾驶3D感知,降低标注成本。

When the City Teaches the Car: Label-Free 3D Perception from Infrastructure

  • 路边单元做教师,通过无标签数据生成伪标签供车辆学习。
  • 在CARLA仿真中达82.3%车辆检测准确率,接近有监督上限94.4%。
  • 适合大规模部署、需减少人工标注的自动驾驶研发团队。

构建鲁棒的自动驾驶3D感知仍严重依赖大规模数据采集与人工标注,但该范式在跨城市部署时变得不切实际。现代城市正广泛部署路侧单元(RSUs),即沿道路和路口设置的静态传感器,用于交通监控。这引出一个自然问题:城市能否帮助训练车辆?我们提出基础设施引导的无标签3D感知,其中RSUs作为固定视角的无监督教师,对路过的车辆进行监督。利用其固定视角和重复观测能力,RSUs从无标签数据中学习局部3D检测器,并将预测结果广播给经过的车辆,车辆聚合这些预测作为伪标签,用于训练独立的自身检测器。最终模型在测试时无需基础设施或通信。我们在CARLA多智能体环境中实现全无标签三阶段流程,使用CenterPoint,车辆检测平均精度(AP)达到82.3%,接近完全监督下自我中心方法的上界94.4%。我们系统分析各阶段性能,评估可扩展性,并验证其与现有自车主导无标签方法的互补性。结果表明,城市基础设施本身可能为自动驾驶提供可扩展的监督信号,使基础设施引导学习成为降低3D感知标注成本的有前景的替代范式。

原文摘要 · Abstract (English)

Building robust 3D perception for self-driving still relies heavily on large-scale data collection and manual annotation, yet this paradigm becomes impractical as deployment expands across diverse cities and regions. Meanwhile, modern cities are increasingly instrumented with roadside units (RSUs), static sensors deployed along roads and at intersections to monitor traffic. This raises a natural question: can the city itself help train the vehicle? We propose infrastructure-taught, label-free 3D perception, a paradigm in which RSUs act as stationary, unsupervised teachers for ego vehicles. Leveraging their fixed viewpoints and repeated observations, RSUs learn local 3D detectors from unlabeled data and broadcast predictions to passing vehicles, which are aggregated as pseudo-label supervision for training a standalone ego detector. The resulting model requires no infrastructure or communication at test time. We instantiate this idea as a fully label-free three-stage pipeline and conduct a concept-and-feasibility study in a CARLA-based multi-agent environment. With CenterPoint, our pipeline achieves 82.3% AP for detecting vehicles, compared to a fully supervised ego upper bound of 94.4%. We further systematically analyze each stage, evaluate its scalability, and demonstrate complementarity with existing ego-centric label-free methods. Together, these results suggest that city infrastructure itself can potentially provide a scalable supervisory signal for autonomous vehicles, positioning infrastructure-taught learning as a promising orthogonal paradigm for reducing annotation cost in 3D perception.

3D感知无监督学习自动驾驶城市智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。