只预测影响驾驶决策的关键视觉特征,实现端到端自动驾驶。
Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving

- 通过联合嵌入预测学习连续驾驶意图,聚焦规划相关特征。
- 在NAVSIM上达91.3 PDMS和89.1 EPDMS,优于传统密集预测方法。
- 无需感知标注或轨迹生成器,适合轻量化高阶规划系统。
现有自动驾驶世界模型通常对未来的视频、占据状态、鸟瞰图表示或车辆运动进行密集预测。本文认为规划无需重建完整未来世界,只需关注影响自身行为的场景特征。基于此,提出Auto-JEPA——一种面向动作的潜在世界模型,通过联合嵌入预测学习连续未来驾驶意图。给定视觉观测、自车运动历史与导航指令,Auto-JEPA预测与未来自车轨迹潜在表征对齐的意图嵌入。该意图从固定轨迹记忆中检索可执行轨迹,并由场景条件化的候选选择模块排序。Auto-JEPA保持视觉编码器冻结,无需显式感知标注,也不使用学习型轨迹生成器。仅优化任务特定模块(轨迹表示、意图预测、候选选择),在NAVSIM v1上取得91.3 PDMS,NAVSIM v2上取得89.1 EPDMS。语义遮蔽实验显示,遮蔽动态目标区域导致意图变化幅度为等面积随机遮蔽的2.97倍;遮蔽影响驾驶的车辆显著改变预测意图与选择轨迹,而遮蔽无关车辆则基本无影响。结果表明,未来意图预测促使模型聚焦规划相关视觉特征,支持高质量规划且无需密集未来世界建模。
原文摘要 · Abstract (English)
Existing autonomous-driving world models typically perform dense prediction of future videos, occupancy states, BEV representations, or agent motion. We argue that planning need not reconstruct the complete future world, but only focus on scene features that affect future ego action. Based on this perspective, we propose Auto-JEPA, an action-oriented latent world model that learns continuous future driving intent through joint-embedding prediction. Given visual observations, egomotion history, and navigation commands, Auto-JEPA predicts an intent embedding aligned with the latent representation of the future ego trajectory. The predicted intent retrieves executable trajectories from a fixed trajectory memory, which are then ranked by a scene-conditioned candidate selection module. Auto-JEPA keeps the visual encoder frozen, requires no explicit perception annotations, and uses no learned trajectory generator. By optimizing only task-specific modules for trajectory representation, intent prediction, and candidate selection, Auto-JEPA achieves 91.3 PDMS on NAVSIM v1 and 89.1 EPDMS on NAVSIM v2. Semantic occlusion experiments show that masking dynamic-agent regions induces an average intent change 2.97x that of equal-area random masking. Moreover, occluding vehicles that affect future driving substantially changes the predicted intent and selected trajectory, whereas both remain essentially unchanged when non-influential vehicles are occluded. These results show that future-intent prediction encourages the model to focus on planning-relevant visual features and supports high-quality planning without dense future-world modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。