arXiv:2512.24385cs.CVcs.RO2025-12被引 6

构建多模态预训练框架,让自动驾驶系统理解空间信息。

Forging Spatial Intelligence: A Roadmap of Multi-Modal Data Pre-Training for Autonomous Systems

  • 提出统一的多模态预训练分类体系,涵盖单模态到融合框架。
  • 在3D目标检测与语义占据预测任务中提升感知能力。
  • 适合研究自动驾驶感知与多传感器融合的学者与工程师。

自主系统(如自动驾驶汽车和无人机)的快速发展,迫切需要从车载多模态传感器数据中构建真正的空间智能。尽管基础模型在单模态场景下表现优异,但如何将相机、激光雷达等异构传感器的能力整合,形成统一的理解仍面临巨大挑战。本文提出一个全面的多模态预训练框架,识别推动该方向进展的核心技术。我们分析了基础传感器特性与学习策略之间的相互作用,评估了平台特定数据集的作用。核心贡献是建立统一的预训练范式分类体系:从单模态基线到复杂融合框架,可学习整体表征以支持3D目标检测、语义占据预测等高级任务。此外,我们探索文本输入与占据表示的融合,以促进开放世界感知与规划。最后,指出计算效率与模型可扩展性等关键瓶颈,并提出迈向通用多模态基础模型的路线图,以实现真实场景中鲁棒的空间智能。

原文摘要 · Abstract (English)

The rapid advancement of autonomous systems, including self-driving vehicles and drones, has intensified the need to forge true Spatial Intelligence from multi-modal onboard sensor data. While foundation models excel in single-modal contexts, integrating their capabilities across diverse sensors like cameras and LiDAR to create a unified understanding remains a formidable challenge. This paper presents a comprehensive framework for multi-modal pre-training, identifying the core set of techniques driving progress toward this goal. We dissect the interplay between foundational sensor characteristics and learning strategies, evaluating the role of platform-specific datasets in enabling these advancements. Our central contribution is the formulation of a unified taxonomy for pre-training paradigms: ranging from single-modality baselines to sophisticated unified frameworks that learn holistic representations for advanced tasks like 3D object detection and semantic occupancy prediction. Furthermore, we investigate the integration of textual inputs and occupancy representations to facilitate open-world perception and planning. Finally, we identify critical bottlenecks, such as computational efficiency and model scalability, and propose a roadmap toward general-purpose multi-modal foundation models capable of achieving robust Spatial Intelligence for real-world deployment.

多模态空间智能自动驾驶预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。