arXiv:2607.04541cs.CVcs.AI2026-07

用雷达与摄像头预测未来点云,实现自动驾驶的通用感知预训练。

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining

论文配图:CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining
图 1 · 摘自论文原文
  • 通过预测未来激光雷达点云,融合图像与雷达数据构建统一鸟瞰表征。
  • 在nuScenes上提升长时距点云预测性能,下游任务全链路效果显著增强。
  • 适合做多模态感知、自动驾驶系统研发的工程师和研究者参考。

相机-雷达(CR)融合是自动驾驶中实用的感知配置,但现有模型通常依赖任务特定监督,限制了可复用表征学习。本文提出CRISP,一种基于预测式世界建模预训练的时空相机-雷达骨干网络。给定历史多视角图像和雷达扫描数据,CRISP通过预测未来激光雷达点云来学习统一的鸟瞰图(BEV)表示。激光雷达仅在预训练阶段作为特权监督信号;部署模型仅需相机与雷达。为使基于预测的预训练在CR融合中有效,CRISP引入增强雷达编码器、雷达增强的时间自注意力机制,以及具有模态创新门控的多模态特征渲染。这些组件将雷达距离与多普勒信息注入BEV时间传播,并允许BEV令牌有选择性地融合相机与雷达证据。在nuScenes上的实验表明,CRISP提升了长时距点云预测能力,并有效迁移至3D检测、跟踪、在线建图、运动预测、未来占用预测及规划等下游任务,表明预测式CR预训练是实现实际传感器配置下可扩展驾驶表征的有前景路径。

原文摘要 · Abstract (English)

Camera-radar (CR) fusion is a practical sensing configuration for autonomous driving, but existing models are typically trained with task-specific supervision, limiting reusable representation learning. We present CRISP, a spatiotemporal CR backbone pretrained through forecasting-based representation learning. Given historical multi-view images and radar sweeps, CRISP learns a unified bird's-eye-view (BEV) representation by predicting future LiDAR point clouds. LiDAR is used only as privileged supervision during pretraining; the deployed model requires only camera and radar. To make forecasting-based pretraining effective for CR fusion, CRISP introduces an enhanced radar encoder, radar-enhanced temporal self-attention, and multimodal feature rendering with modality innovation gating. These components inject radar range and Doppler cues into BEV temporal propagation and allow BEV tokens to selectively incorporate camera and radar evidence. Experiments on nuScenes show that CRISP improves long-horizon point cloud forecasting and transfers effectively to downstream tasks, including 3D detection, tracking, online mapping, motion forecasting, future occupancy prediction, and planning, suggesting that predictive CR pretraining is a promising path toward scalable driving representations under practical sensor configurations. The project website is https://umfieldrobotics.github.io/CRISP.

多模态融合自动驾驶预训练感知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。