arXiv:2602.12540cs.CVcs.RO2026-02被引 4

用自监督学习构建车载激光雷达环境模型,预测未来场景演化。

Self-Supervised JEPA-based World Models for LiDAR Occupancy Completion and Forecasting

  • 基于JEPA框架,利用无标注数据训练世界模型。
  • 预训练编码器使激光雷达占位完成与预测任务性能提升。
  • 适合自动驾驶长期规划研究者参考。

自动驾驶作为在物理世界中运行的智能体,需具备构建时空演化环境模型的能力以支持长期规划。同时,为满足可扩展性,模型需通过自监督方式学习;联合嵌入预测架构(JEPA)可通过大量未标注数据实现世界模型学习,无需依赖昂贵的人工标注。本文提出AD-LiST-JEPA,一种基于JEPA框架的自监督世界模型,利用激光雷达数据预测未来的时空演化。通过下游的激光雷达占位完成与预测(OCF)任务评估所学表征质量,该任务联合评估感知与预测能力。概念验证实验表明,经JEPA训练后的预训练编码器能显著提升OCF性能。

原文摘要 · Abstract (English)

Autonomous driving, as an agent operating in the physical world, requires the fundamental capability to build \textit{world models} that capture how the environment evolves spatiotemporally in order to support long-term planning. At the same time, scalability demands learning such models in a self-supervised manner; \textit{joint-embedding predictive architecture (JEPA)} enables learning world models via leveraging large volumes of unlabeled data without relying on expensive human annotations. In this paper, we propose \textbf{AD-LiST-JEPA}, a self-supervised world model for autonomous driving that predicts future spatiotemporal evolution from LiDAR data using a JEPA framework. We evaluate the quality of the learned representations through a downstream LiDAR-based occupancy completion and forecasting (OCF) task, which jointly assesses perception and prediction. Proof of concept experiments show better OCF performance with pretrained encoder after JEPA-based world model learning.

自动驾驶自监督世界模型激光雷达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。