用无标注行车视频训练自动驾驶模型,无需标签或激光雷达。
Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos
- 用多模态教师指导单目视频,端到端预测点云、位姿和语义分割。
- 在未标定的YouTube视频上训练,实现媲美多摄像头和激光雷达的规划性能。
- 适合想用低成本单目摄像头构建自动驾驶系统的研究者与工程师。
在线获取的视角驾驶视频提供了丰富的视觉数据,但缺乏标注使其难以学习同时包含语义结构和3D几何特征的表征。近期大型前馈空间模型表明,可在一次前向传播中推断点图和自车运动,为可扩展的驾驶感知指明方向。为此,我们提出一种无标签、教师引导的框架,直接从无姿态视频中学习自动驾驶表征。不同于以往主要关注帧间一致性的自监督方法,我们认为安全且响应式的驾驶高度依赖于时间上下文。因此,我们采用带有轻量级自回归模块的前馈架构,利用多模态监督信号,联合预测当前与未来点图、相机位姿、语义分割和运动掩码。多模态教师提供序列级伪监督,使LFG能从原始YouTube视频中学习统一的伪4D表征,无需姿态、标签或激光雷达。所获编码器在NAVSIM基准上有效迁移至下游自动驾驶规划任务,仅使用单目摄像头即超越多摄像头和激光雷达基线;在多种语义、几何及定性运动预测任务中也表现优异。这些具备几何与运动感知能力的特征使LFG成为自动驾驶领域极具潜力的视频中心基础模型。
原文摘要 · Abstract (English)
Ego-centric driving videos available online provide an abundant source of visual data for autonomous driving, yet their lack of annotations makes it difficult to learn representations that capture both semantic structure and 3D geometry. Recent advances in large feedforward spatial models demonstrate that point maps and ego-motion can be inferred in a single forward pass, suggesting a promising direction for scalable driving perception. We therefore propose a label-free, teacher-guided framework for learning autonomous driving representations directly from unposed videos. Unlike prior self-supervised approaches that focus primarily on frame-to-frame consistency, we posit that safe and reactive driving depends critically on temporal context. To this end, we leverage a feedforward architecture equipped with a lightweight autoregressive module, trained using multi-modal supervisory signals that guide the model to jointly predict current and future point maps, camera poses, semantic segmentation, and motion masks. Multi-modal teachers provide sequence-level pseudo-supervision, enabling LFG to learn a unified pseudo-4D representation from raw YouTube videos without poses, labels, or LiDAR. The resulting encoder not only transfers effectively to downstream autonomous driving planning on the NAVSIM benchmark, surpassing multi-camera and LiDAR baselines with only a single monocular camera, but also yields strong performance when evaluated on a range of semantic, geometric, and qualitative motion prediction tasks. These geometry and motion-aware features position LFG as a compelling video-centric foundation model for autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。