通过分离语义子空间提升多模态地图预测鲁棒性
SEF-MAP: Subspace-Decomposed Expert Fusion for Robust Multimodal HD Map Prediction
- 将鸟瞰图特征分解为四类子空间,每类由专用专家处理
- 在nuScenes和Argoverse2上分别提升4.2%和4.8%的mAP
- 适合需要应对低光、遮挡等恶劣条件的自动驾驶系统
高精地图对自动驾驶至关重要,但相机与激光雷达的多模态融合常因模态不一致导致性能下降,尤其在低光照、遮挡或点云稀疏条件下。为此,我们提出SEF-MAP框架,通过显式将鸟瞰图(BEV)特征分解为四个语义子空间:激光雷达私有、图像私有、共享与交互。每个子空间分配专用专家,以保留模态特异性线索并捕捉跨模态共识。引入基于不确定性的门控机制,在BEV单元层级动态降低不可靠专家权重,并加入使用均衡正则项防止专家坍塌。为增强退化场景下的鲁棒性并促进角色分化,提出分布感知掩码:训练时用EMA统计代理特征模拟模态缺失,并通过专业化损失强制私有、共享与交互专家在完整与掩码输入下表现不同。在nuScenes和Argoverse2基准测试中,SEF-MAP分别超越先前方法4.2%和4.8%的mAP,提供了一种在复杂退化条件下有效的多模态高精地图预测方案。
原文摘要 · Abstract (English)
High-definition (HD) maps are essential for autonomous driving, yet multi-modal fusion often suffers from inconsistency between camera and LiDAR modalities, leading to performance degradation under low-light conditions, occlusions, or sparse point clouds. To address this, we propose SEFMAP, a Subspace-Expert Fusion framework for robust multimodal HD map prediction. The key idea is to explicitly disentangle BEV features into four semantic subspaces: LiDAR-private, Image-private, Shared, and Interaction. Each subspace is assigned a dedicated expert, thereby preserving modality-specific cues while capturing cross-modal consensus. To adaptively combine expert outputs, we introduce an uncertainty-aware gating mechanism at the BEV-cell level, where unreliable experts are down-weighted based on predictive variance, complemented by a usage balance regularizer to prevent expert collapse. To enhance robustness in degraded conditions and promote role specialization, we further propose distribution-aware masking: during training, modality-drop scenarios are simulated using EMA-statistical surrogate features, and a specialization loss enforces distinct behaviors of private, shared, and interaction experts across complete and masked inputs. Experiments on nuScenes and Argoverse2 benchmarks demonstrate that SEFMAP achieves state-of-the-art performance, surpassing prior methods by +4.2% and +4.8% in mAP, respectively. SEF-MAPprovides a robust and effective solution for multi-modal HD map prediction under diverse and degraded conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。