通过预测时空中的几何与语义结构,提升自动驾驶预训练效果。
GASP: Unifying Geometric and Semantic Self-Supervised Pre-training for Autonomous Driving
- 联合预测4D场景占用、车辆路径和视觉特征,学习统一表征。
- 在多个基准上显著提升语义占用预测与轨迹预测性能。
- 适合需要强泛化能力的自动驾驶系统研发人员。
基于下一词预测的自监督预训练使大语言模型能够捕捉文本内在结构,并在大规模应用下实现卓越性能。类似地,自动驾驶生成海量时空数据,暗示可通过规模学习环境及其随时间演变的几何与语义结构。为此,我们提出一种几何与语义自监督预训练方法GASP,通过在任意查询的时空点预测:(1) 通用占据,捕捉3D场景的动态结构;(2) 自身占据,建模车辆行驶路径;(3) 来自视觉基础模型的提炼高层特征。通过建模几何与语义4D占据场而非原始传感器数据,模型学习到结构化且可泛化的环境演化表征。我们在多个自动驾驶基准上验证了GASP,显著提升了语义占据预测、在线地图构建和自身轨迹预测性能。结果表明,连续4D几何与语义占据预测为自动驾驶提供了一种可扩展且高效的预训练范式。
原文摘要 · Abstract (English)
Self-supervised pre-training based on next-token prediction has enabled large language models to capture the underlying structure of text, and has led to unprecedented performance on a large array of tasks when applied at scale. Similarly, autonomous driving generates vast amounts of spatiotemporal data, alluding to the possibility of harnessing scale to learn the underlying geometric and semantic structure of the environment and its evolution over time. In this direction, we propose a geometric and semantic self-supervised pre-training method, GASP, that learns a unified representation by predicting, at any queried future point in spacetime, (1) general occupancy, capturing the evolving structure of the 3D scene; (2) ego occupancy, modeling the ego vehicle path through the environment; and (3) distilled high-level features from a vision foundation model. By modeling geometric and semantic 4D occupancy fields instead of raw sensor measurements, the model learns a structured, generalizable representation of the environment and its evolution through time. We validate GASP on multiple autonomous driving benchmarks, demonstrating significant improvements in semantic occupancy forecasting, online mapping, and ego trajectory prediction. Our results demonstrate that continuous 4D geometric and semantic occupancy prediction provides a scalable and effective pre-training paradigm for autonomous driving. For code and additional visualizations, see \href{https://research.zenseact.com/publications/gasp/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。