arXiv:2605.25262cs.CV2026-05中稿 · ICRA

用语义引导的掩码策略提升3D感知,让自动驾驶更精准。

Semantics-Guided Multimodal Masked Autoencoder Pretraining for 3D BEV Object Detection

论文配图:Semantics-Guided Multimodal Masked Autoencoder Pretraining for 3D BEV Object Detection
图 1 · 摘自论文原文
  • 基于语义信息动态掩码雷达点云,保留重要区域
  • 引入点级语义解码分支,提升检测精度1.49%以上
  • 适合追求高精度3D目标检测的自动驾驶研究者

准确的3D鸟瞰图(BEV)目标检测对自动驾驶至关重要,依赖于相机与激光雷达等互补传感器的有效多模态表征。多模态掩码自编码器在学习此类表征方面展现出巨大潜力。然而,现有方法通常对相机和激光雷达输入采用均匀随机掩码,同等对待所有区域,仅通过掩码重建学习表征。本文提出一种语义引导的多模态掩码自编码器框架,通过两个独立组件引入语义信息:(i) 语义引导的激光雷达体素掩码,更强地保留语义重要区域;(ii) 辅助的点级激光雷达语义解码分支,额外注入语义引导。在BEVFusion 3D目标检测任务上,相较于标准的UniM2AE基线,在nuScenes mini验证集上,语义引导的激光雷达体素掩码带来+1.49% mAP和+1.66% NDS提升,解码分支的点语义监督带来+1.39% mAP和+3.22% NDS提升。

原文摘要 · Abstract (English)

Accurate 3D bird's-eye view (BEV) object detection is essential for autonomous driving, and depends strongly on effective multimodal representations from complementary sensors such as cameras and LiDAR. Multimodal masked autoencoders have shown strong potential for learning such representations for downstream 3D BEV object detection. However, existing methods typically apply uniform random masking to camera and LiDAR inputs, treating all regions equally, and learn representations only through masked reconstruction. We propose a semantics-guided multimodal masked autoencoder framework that introduces semantic information during pretraining through two separate components: (i) semantics-guided LiDAR voxel masking, which preserves semantically important LiDAR regions more strongly, and (ii) an auxiliary point-wise LiDAR semantic decoder branch that injects semantic guidance in addition to reconstruction. On BEVFusion 3D object detection, our semantics-guided pretraining strategy improves performance on the nuScenes mini validation set compared to the standard UniM2AE baseline: semantics-guided LiDAR voxel masking yields +1.49% mean Average Precision (mAP) and +1.66% nuScenes Detection Score (NDS), while decoder-side point semantic supervision yields +1.39% mAP and +3.22% NDS over the baseline.

3D检测多模态自编码器自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。