arXiv:2512.19934cs.CVcs.AI2025-12被引 3

用多模态结构先验提升车辆感知预训练效果

Vehicle-centric Perception via Multimodal Structured Pre-training

  • 引入对称、轮廓、语义三类结构先验指导遮蔽重建
  • 在5个下游任务中超越现有方法,性能显著提升
  • 适合自动驾驶与智能交通系统中的车辆理解场景

车辆中心感知在大规模监控、智能交通和自动驾驶中至关重要。现有方法在预训练阶段缺乏有效的车辆知识学习,导致难以建模通用的车辆感知表示。为此,我们提出VehicleMAE-V2,一种新型车辆中心预训练大模型。通过挖掘并利用车辆相关的多模态结构先验引导遮蔽标记重建,显著增强模型学习通用车辆感知表示的能力。具体设计了对称性引导掩码模块(SMM)、轮廓引导表征模块(CRM)和语义引导表征模块(SRM),分别融入车辆的对称性、轮廓和语义结构先验。SMM利用车辆对称约束避免保留对称块,从而选择高质量遮蔽图像块并减少信息冗余;CRM最小化轮廓特征与重建特征间的概率分布差异,以在像素级重建中保持整体车辆结构;SRM通过对比学习与跨模态蒸馏对齐图文特征,缓解遮蔽重建中因语义理解不足导致的特征混淆。为支持VehicleMAE-V2的预训练,我们构建了Autobot4M数据集,包含约400万张车辆图像和12,693条文本描述。在五个下游任务上的大量实验表明,VehicleMAE-V2表现优异。

原文摘要 · Abstract (English)

Vehicle-centric perception plays a crucial role in many intelligent systems, including large-scale surveillance systems, intelligent transportation, and autonomous driving. Existing approaches lack effective learning of vehicle-related knowledge during pre-training, resulting in poor capability for modeling general vehicle perception representations. To handle this problem, we propose VehicleMAE-V2, a novel vehicle-centric pre-trained large model. By exploring and exploiting vehicle-related multimodal structured priors to guide the masked token reconstruction process, our approach can significantly enhance the model's capability to learn generalizable representations for vehicle-centric perception. Specifically, we design the Symmetry-guided Mask Module (SMM), Contour-guided Representation Module (CRM) and Semantics-guided Representation Module (SRM) to incorporate three kinds of structured priors into token reconstruction including symmetry, contour and semantics of vehicles respectively. SMM utilizes the vehicle symmetry constraints to avoid retaining symmetric patches and can thus select high-quality masked image patches and reduce information redundancy. CRM minimizes the probability distribution divergence between contour features and reconstructed features and can thus preserve holistic vehicle structure information during pixel-level reconstruction. SRM aligns image-text features through contrastive learning and cross-modal distillation to address the feature confusion caused by insufficient semantic understanding during masked reconstruction. To support the pre-training of VehicleMAE-V2, we construct Autobot4M, a large-scale dataset comprising approximately 4 million vehicle images and 12,693 text descriptions. Extensive experiments on five downstream tasks demonstrate the superior performance of VehicleMAE-V2.

车辆感知多模态预训练结构先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。