arXiv:2608.24366cs.CVcs.RO2026-08

通过空间可靠性门控融合多模态感知,提升传感器失效下的自动驾驶鲁棒性。

Variance-Guided Spatial Attention Fusion for Robust End-to-End Driving under Asymmetric Sensor Degradation

论文配图:Variance-Guided Spatial Attention Fusion for Robust End-to-End Driving under Asymmetric Sensor Degradation
图 1 · 摘自论文原文
  • 利用物理模拟生成密集可靠性掩码,指导模型判断各区域可信度。
  • 在CARLA Longest6上实现93.2分驾驶得分,优于基线模型。
  • 适合研究自动驾驶多模态融合与故障容错的工程师与研究员。

端到端多模态驾驶系统在融合摄像头与激光雷达数据方面进展迅速。然而,在传感器异构退化场景下——即某一模态整体或局部区域受损而其余区域仍可用时,现有方法仍显脆弱。核心挑战不仅在于添加不确定性头,更在于获取稠密的可靠性监督、将该可靠性校准至物理故障严重程度,并在不可靠特征影响规划前进行处理。为此,本文提出方差引导的空间注意力融合(VG-SAF),其中稠密的异方差可靠性估计作为可解释的空间门控。框架包含三部分:首先,一个基于物理的增强器模拟典型摄像头和激光雷达故障,输出连续空间掩码,提供无需额外标注的稠密监督;其次,模态特异性专家通过对数空间中的跨分支密集蒸馏预测像素级可靠性尺度,强制实现严重度到尺度的单调响应;第三,校准后的可靠性图驱动混合注意力机制,通过局部空间门控抑制不可靠单元,并以跨模态信任Softmax协调模态间决策。拉普拉斯不确定性头输出系统性路径点不确定性尺度,可识别训练范围外的严重或复合传感器退化。在CARLA Longest6基准上,VG-SAF在纯摄像头、纯激光雷达及联合退化场景下均持续提升闭环鲁棒性,体现在驾驶得分、路线完成率和违规分数上的综合优化。

原文摘要 · Abstract (English)

End-to-end multimodal driving has progressed rapidly by fusing camera and LiDAR streams. Existing pipelines remain fragile under asymmetric sensor degradation, where either an entire modality or only a localized region is corrupted while other regions remain useful. The key difficulty is not simply to add an uncertainty head, but to obtain dense reliability supervision, calibrate this reliability against physical fault severity, and use it before unreliable features bias the planner. We propose Variance-Guided Spatial Attention Fusion (VG-SAF), in which dense heteroscedastic reliability estimates act as interpretable spatial gates. The framework couples three components. First, a physically grounded augmentor simulates representative camera and LiDAR failures and emits a continuous spatial mask, providing dense supervision without additional annotation. Second, modality-specific experts predict per-pixel reliability scales through cross-branch dense distillation in log space, enforcing a monotone severity-to-scale response. Third, calibrated reliability maps drive a hybrid attention mechanism that suppresses unreliable cells with a local spatial gate and arbitrates between modalities through a cross-modal trust softmax. A Laplace uncertainty head emits a systemic waypoint uncertainty scale that signals severe or combined sensor degradation, including severities outside the training ranges. On the CARLA Longest6 benchmark, VG-SAF consistently improves closed-loop robustness over the baselines across camera-only, LiDAR-only, and joint degradation regimes, as measured by driving score, route completion, and infraction score.

自动驾驶多模态融合可靠性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。