arXiv:2511.00859cs.CV2025-11NeurIPS被引 1

提出分层解耦方法,让自动驾驶多传感器融合模型更透明

Layer-Wise Modality Decomposition for Interpretable Multimodal Sensor Fusion

论文配图:Layer-Wise Modality Decomposition for Interpretable Multimodal Sensor Fusion
图 1 · 摘自论文原文
  • 逐层分离不同传感器信息,解析各模态贡献
  • 在摄像头-雷达/激光雷达等组合下验证有效
  • 适合关注自动驾驶决策可解释性的研究者

在自动驾驶中,感知模型决策的透明性至关重要,单一误判可能引发严重后果。然而,多传感器输入导致信息在融合网络中相互纠缠,难以判断各模态对预测的具体贡献。本文提出分层模态解耦(LMD),一种后处理、模型无关的可解释性方法,可在预训练融合模型的所有层级上解耦特定模态的信息。据我们所知,LMD是首个在自动驾驶传感器融合系统中,将感知模型预测归因于单一输入模态的方法。我们在摄像头-雷达、摄像头-LiDAR及摄像头-雷达-LiDAR三种设置下的预训练融合模型上评估了LMD。通过结构化扰动指标和模态级可视化分解验证其有效性,证明该方法可实际应用于高容量多模态架构的解释。代码已开源:https://github.com/detxter-jvb/Layer-Wise-Modality-Decomposition。

原文摘要 · Abstract (English)

In autonomous driving, transparency in the decision-making of perception models is critical, as even a single misperception can be catastrophic. Yet with multi-sensor inputs, it is difficult to determine how each modality contributes to a prediction because sensor information becomes entangled within the fusion network. We introduce Layer-Wise Modality Decomposition (LMD), a post-hoc, model-agnostic interpretability method that disentangles modality-specific information across all layers of a pretrained fusion model. To our knowledge, LMD is the first approach to attribute the predictions of a perception model to individual input modalities in a sensor-fusion system for autonomous driving. We evaluate LMD on pretrained fusion models under camera-radar, camera-LiDAR, and camera-radar-LiDAR settings for autonomous driving. Its effectiveness is validated using structured perturbation-based metrics and modality-wise visual decompositions, demonstrating practical applicability to interpreting high-capacity multimodal architectures. Code is available at https://github.com/detxter-jvb/Layer-Wise-Modality-Decomposition.

可解释性多模态融合自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。