提出不确定性感知的多模态融合框架,提升复杂场景下目标检测鲁棒性。
Cocoon: Robust Multi-Modal Perception with Uncertainty-Aware Sensor Fusion
- 在物体与特征层面量化异构表征不确定性,实现跨模态公平比较。
- 在正常与挑战性场景中均优于静态和自适应融合方法,包括自然与人工噪声。
- 适用于多模态感知任务,尤其适合长尾分布与高噪声环境。
3D目标检测中,利用多模态信息可提升正常及挑战性条件下的精度,尤其在长尾场景中。现有研究主要分为两类:基于MoE的自适应融合难以处理不同物体配置带来的不确定性;晚期融合依赖独立检测流水线,限制了整体理解。本文提出Cocoon,一种面向物体与特征层级的不确定性感知融合框架。其核心创新在于对异构表征进行不确定性量化,通过引入特征对齐器与可学习的代理真值(称为特征印象),实现跨模态公平比较。同时定义训练目标,确保该关系构成有效的不确定性度量。Cocoon在正常与挑战性条件下均显著优于现有静态与自适应方法,包括存在自然与人为扰动的情况。此外,其不确定性度量在多种数据集上表现出有效性与鲁棒性。
原文摘要 · Abstract (English)
An important paradigm in 3D object detection is the use of multiple modalities to enhance accuracy in both normal and challenging conditions, particularly for long-tail scenarios. To address this, recent studies have explored two directions of adaptive approaches: MoE-based adaptive fusion, which struggles with uncertainties arising from distinct object configurations, and late fusion for output-level adaptive fusion, which relies on separate detection pipelines and limits comprehensive understanding. In this work, we introduce Cocoon, an object- and feature-level uncertainty-aware fusion framework. The key innovation lies in uncertainty quantification for heterogeneous representations, enabling fair comparison across modalities through the introduction of a feature aligner and a learnable surrogate ground truth, termed feature impression. We also define a training objective to ensure that their relationship provides a valid metric for uncertainty quantification. Cocoon consistently outperforms existing static and adaptive methods in both normal and challenging conditions, including those with natural and artificial corruptions. Furthermore, we show the validity and efficacy of our uncertainty metric across diverse datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。