arXiv:2505.24139cs.CVcs.AI2025-05CVPR被引 22

无需人工标注,用3D视觉表示实现可扩展的自动驾驶轨迹规划

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation

  • 将多模态大模型的2D视觉表示转换为3D空间,通过稀疏体素策略融合多视角多帧信息
  • 在nuScenes和Waymo数据集上性能优于有监督方法,且不依赖人工标注
  • 适合追求无监督、可扩展自动驾驶系统的研究者与工程师

多模态大语言模型(MLLM)的最新进展推动了端到端自动驾驶运动规划的新热潮。现有方法大多依赖人工标注学习感知与预测任务,而纯自监督方法虽无需标注但性能常落后于前沿水平。我们发现关键瓶颈在于输入表征:基于MLLM的端到端方法通常在2D图像空间预训练,而非自动驾驶实际使用的3D空间。为此,我们提出S4-Driver,一种基于PaLI模型的可扩展自监督运动规划算法,采用新型稀疏体素策略,将MLLM强大的2D视觉表示无缝转化为3D空间,无需微调视觉编码器。该表征聚合多视角、多帧视觉输入,显著提升3D空间中的轨迹预测能力。我们在nuScenes和Waymo Open Motion Dataset(含自建相机数据)上进行实验,结果表明S4-Driver在无需任何人工标注的情况下,性能优于现有监督多任务方法,并在大规模未标注驾驶日志上展现出良好可扩展性。

原文摘要 · Abstract (English)

The latest advancements in multi-modal large language models (MLLMs) have spurred a strong renewed interest in end-to-end motion planning approaches for autonomous driving. Many end-to-end approaches rely on human annotations to learn intermediate perception and prediction tasks, while purely self-supervised approaches--which directly learn from sensor inputs to generate planning trajectories without human annotations often underperform the state of the art. We observe a key gap in the input representation space: end-to-end approaches built on MLLMs are often pretrained with reasoning tasks in 2D image space rather than the native 3D space in which autonomous vehicles plan. To this end, we propose S4-Driver, a scalable self-supervised motion planning algorithm with spatio-temporal visual representation, based on the popular PaLI multimodal large language model. S4-Driver uses a novel sparse volume strategy to seamlessly transform the strong visual representation of MLLMs from perspective view to 3D space without the need to finetune the vision encoder. This representation aggregates multi-view and multi-frame visual inputs and enables better prediction of planning trajectories in 3D space. To validate our method, we run experiments on both nuScenes and Waymo Open Motion Dataset (with in-house camera data). Results show that S4-Driver performs favorably against existing supervised multi-task approaches while requiring no human annotations. It also demonstrates great scalability when pretrained on large volumes of unannotated driving logs.

自动驾驶自监督3D表征多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。