用单目视频自动生成车辆可行轨迹,无需人工标注
GHOST: Ground-projected Hypotheses from Observed Structure-from-Motion Trajectories
- 通过单目SfM恢复相机轨迹并投影到地面,生成无标注的空间掩码
- 在NuScenes上实现可靠轨迹预测,支持跨设备迁移
- 适合自动驾驶路径规划与零样本场景泛化研究
我们提出一种可扩展的自监督方法,从单目图像中分割复杂城市环境下的可行车辆轨迹。利用大规模行车记录仪视频,将记录的本车运动作为隐式监督信号,通过单目结构从运动(monocular structure-from-motion)恢复相机轨迹,并将其投影至地面平面,生成未遍历区域的空间掩码,无需人工标注。这些自动生成的标签用于训练深度分割网络,可在运行时仅凭单张RGB图像预测运动条件下的路径提案,无需显式建模道路或车道线。模型在多样化、无约束的互联网数据上训练,隐式学习场景布局、车道拓扑和交叉口结构,且能适应不同相机配置。我们在NuScenes上评估该方法,展示了可靠的轨迹预测能力,并通过轻量微调成功迁移至电动滑板车平台。结果表明,大规模自我运动蒸馏可生成结构化且可泛化的路径提案,超越已观测轨迹,实现基于图像分割的轨迹假设估计。
原文摘要 · Abstract (English)
We present a scalable self-supervised approach for segmenting feasible vehicle trajectories from monocular images for autonomous driving in complex urban environments. Leveraging large-scale dashcam videos, we treat recorded ego-vehicle motion as implicit supervision and recover camera trajectories via monocular structure-from-motion, projecting them onto the ground plane to generate spatial masks of traversed regions without manual annotation. These automatically generated labels are used to train a deep segmentation network that predicts motion-conditioned path proposals from a single RGB image at run time, without explicit modeling of road or lane markings. Trained on diverse, unconstrained internet data, the model implicitly captures scene layout, lane topology, and intersection structure, and generalizes across varying camera configurations. We evaluate our approach on NuScenes, demonstrating reliable trajectory prediction, and further show transfer to an electric scooter platform through light fine-tuning. Our results indicate that large-scale ego-motion distillation yields structured and generalizable path proposals beyond the demonstrated trajectory, enabling trajectory hypothesis estimation via image segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。