用自监督视频预训练+多模态轨迹蒸馏,提升端到端自动驾驶表现
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
- 基于V-JEPA框架,用大规模驾驶视频预训练视觉编码器
- 融合仿真与真实轨迹,实现93.3 PDMS的最新性能
- 适合关注自动驾驶规划与表征学习的研究者
端到端自动驾驶越来越多地利用自监督视频预训练来学习可迁移的规划表征。然而,目前针对场景理解的视频世界模型预训练带来的改进有限。这一局限性因驾驶任务本身固有的模糊性而加剧:每幅场景通常只对应一条人类行驶轨迹,难以学习多模式行为。本文提出Drive-JEPA,将视频联合嵌入预测架构(V-JEPA)与多模态轨迹蒸馏相结合,用于端到端驾驶。首先,我们将V-JEPA适配至端到端驾驶任务,在大规模驾驶视频上预训练ViT编码器,生成与轨迹规划对齐的预测表征。其次,引入以提议为中心的规划器,蒸馏仿真生成的多样化轨迹与真实人类轨迹,并通过动量感知选择机制促进稳定安全的行为。在NAVSIM上,仅使用V-JEPA表征与简单Transformer解码器,即在无感知设置下优于先前方法3 PDMS。完整框架在v1上达到93.3 PDMS,v2上达87.8 EPDMS,创下新纪录。
原文摘要 · Abstract (English)
End-to-end autonomous driving increasingly leverages self-supervised video pretraining to learn transferable planning representations. However, pretraining video world models for scene understanding has so far brought only limited improvements. This limitation is compounded by the inherent ambiguity of driving: each scene typically provides only a single human trajectory, making it difficult to learn multimodal behaviors. In this work, we propose Drive-JEPA, a framework that integrates Video Joint-Embedding Predictive Architecture (V-JEPA) with multimodal trajectory distillation for end-to-end driving. First, we adapt V-JEPA for end-to-end driving, pretraining a ViT encoder on large-scale driving videos to produce predictive representations aligned with trajectory planning. Second, we introduce a proposal-centric planner that distills diverse simulator-generated trajectories alongside human trajectories, with a momentum-aware selection mechanism to promote stable and safe behavior. When evaluated on NAVSIM, the V-JEPA representation combined with a simple transformer-based decoder outperforms prior methods by 3 PDMS in the perception-free setting. The complete Drive-JEPA framework achieves 93.3 PDMS on v1 and 87.8 EPDMS on v2, setting a new state-of-the-art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。