DriveX通过自监督学习建模通用驾驶场景,提升模型泛化能力。
DriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous Driving
- 采用多模态统一监督,融合3D点云、2D语义和图像生成建模场景演化。
- 在3D未来点云预测上超越现有方法,多任务表现达顶尖水平。
- 适合追求鲁棒性与通用性的自动驾驶系统研发者使用。
数据驱动的学习推动了自动驾驶发展,但任务专用模型在分布外场景中表现不佳,受限于狭窄优化目标和对昂贵标注数据的依赖。我们提出DriveX,一种自监督世界模型,能从大规模驾驶视频中学习可泛化的场景动态与整体表征(几何、语义、运动)。DriveX引入全域场景建模(OSM)模块,统一3D点云预测、2D语义表示与图像生成的多模态监督,捕捉全面的场景演变。为简化复杂动态学习,提出解耦潜空间世界建模策略,将世界表征学习与未来状态解码分离,并结合动态感知射线采样增强运动建模。针对下游适应,设计未来空间注意力(FSA),动态聚合来自DriveX预测的时空特征以增强任务特定推理。大量实验表明,DriveX在3D未来点云预测上显著优于先前方法,并在占据预测、流估计和端到端驾驶等多样任务中达到最先进水平。这些结果验证了DriveX作为通用世界模型的能力,为构建鲁棒且统一的自动驾驶框架铺平道路。
原文摘要 · Abstract (English)
Data-driven learning has advanced autonomous driving, yet task-specific models struggle with out-of-distribution scenarios due to their narrow optimization objectives and reliance on costly annotated data. We present DriveX, a self-supervised world model that learns generalizable scene dynamics and holistic representations (geometric, semantic, and motion) from large-scale driving videos. DriveX introduces Omni Scene Modeling (OSM), a module that unifies multimodal supervision-3D point cloud forecasting, 2D semantic representation, and image generation-to capture comprehensive scene evolution. To simplify learning complex dynamics, we propose a decoupled latent world modeling strategy that separates world representation learning from future state decoding, augmented by dynamic-aware ray sampling to enhance motion modeling. For downstream adaptation, we design Future Spatial Attention (FSA), a unified paradigm that dynamically aggregates spatiotemporal features from DriveX's predictions to enhance task-specific inference. Extensive experiments demonstrate DriveX's effectiveness: it achieves significant improvements in 3D future point cloud prediction over prior work, while attaining state-of-the-art results on diverse tasks including occupancy prediction, flow estimation, and end-to-end driving. These results validate DriveX's capability as a general-purpose world model, paving the way for robust and unified autonomous driving frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。