让自动驾驶视频自监督学习更关注关键区域,提升感知精度。
Mask What Matters: Saliency-Guided Video Self-Supervised Learning for Autonomous Driving

- 根据语义重要性和时间相关性动态遮蔽视频片段,而非随机遮挡。
- 在BDD100k上减少25%身份切换,在Cityscapes上达73.2 mIoU。
- 适合需要高精度视觉理解的自动驾驶场景,尤其关注安全关键目标。
通过掩码时空预测的视频自监督学习已成为从无标签数据中学习特征表示的有前景范式。然而,现有方法通常采用随机掩码,不加区分地移除区域,无论其语义或时间相关性如何。在自车视角驾驶视频中,安全关键线索如行人、车辆、车道线和动态交互往往仅占画面小部分,却对下游感知至关重要。我们提出V-JEPA4A,一种面向自动驾驶的V-JEPA定制版本,基于公开驾驶视频预训练,并引入新颖的显著性驱动掩码策略。该策略根据语义重要性和时间相关性保留并预测上下文,从而获得更具信息量的表征学习,同时保持掩码预测的效率。我们在四个涵盖跟踪、语义分割和深度估计的驾驶基准上评估所得编码器。结果表明,与随机掩码的V-JEPA相比,V-JEPA4A在BDD100k MOT上将身份切换减少25%,在Cityscapes上达到73.2 mIoU,KITTI-2015深度估计达到3.75 RMSE,且预训练迭代开销仅增加约14%。
原文摘要 · Abstract (English)
Video self-supervised learning through masked spatiotemporal prediction has emerged as a promising paradigm for learning feature representations from unlabeled data. However, existing methods typically rely on random masking, which indiscriminately removes regions irrespective of their semantic or temporal relevance. In ego-centric driving videos, this can weaken the pretext signal since safety-critical cues such as pedestrians, vehicles, lane boundaries, and dynamic interactions often occupy only a small portion of the frame, yet are central to downstream perception. We introduce V-JEPA4A, a domain-specialized variant of V-JEPA for autonomous driving that is pre-trained on publicly available driving videos with a novel saliency-driven masking policy. It accounts for semantically and temporally relevant context. The proposed policy preserves and predicts context according to semantic importance and temporal relevance, yielding more informative representation learning while retaining the efficiency of masked prediction. We evaluate the resulting encoders on four driving benchmarks spanning tracking, semantic segmentation, and depth estimation. The results demonstrate that V-JEPA4A reduces identity switches on BDD100k MOT by 25% over V-JEPA with random masking, achieves 73.2 mIoU on Cityscapes, and 3.75 RMSE on KITTI-2015 depth, while incurring only ~14% additional pre-training iteration overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。