用霍尔德散度提升多模态学习的不确定性估计可靠性
Uncertainty Quantification via Hölder Divergence for Multi-View Representation Learning
- 基于霍尔德散度度量多模态分布差异,捕捉不同模态间的领域差距
- 在多个基准上超越现有方法,尤其在不完整或噪声数据下表现更优
- 适用于需要可靠不确定性的多模态任务,如医疗影像分析
基于证据的深度学习为不确定性估计提供了新兴范式,能在几乎无额外计算开销的情况下提供可靠预测。现有方法通常采用相对熵估计网络预测不确定性,却忽视了不同模态间的领域差异。为此,本文提出一种基于霍尔德散度(Hölder Divergence, HD)的新算法,通过解决不完整或噪声数据带来的固有不确定性挑战,增强多视图学习的可靠性。该方法通过并行网络分支提取多模态表征,并利用HD估计预测不确定性。结合戴姆斯特-沙弗理论融合各模态的不确定性,生成综合结果。数学上,HD能更准确衡量真实数据分布与模型预测分布之间的“距离”,显著提升多分类识别性能。具体而言,本方法在所有评估基准上均优于现有最先进方法。通过多种骨干网络的广泛实验验证了其优越鲁棒性。进一步在更具挑战性的不完整或噪声数据场景下测试,结果表明本方法对数据退化具有高度容忍性。
原文摘要 · Abstract (English)
Evidence-based deep learning represents a burgeoning paradigm for uncertainty estimation, offering reliable predictions with negligible extra computational overheads. Existing methods usually adopt Kullback-Leibler divergence to estimate the uncertainty of network predictions, ignoring domain gaps among various modalities. To tackle this issue, this paper introduces a novel algorithm based on Hölder Divergence (HD) to enhance the reliability of multi-view learning by addressing inherent uncertainty challenges from incomplete or noisy data. Generally, our method extracts the representations of multiple modalities through parallel network branches, and then employs HD to estimate the prediction uncertainties. Through the Dempster-Shafer theory, integration of uncertainty from different modalities, thereby generating a comprehensive result that considers all available representations. Mathematically, HD proves to better measure the ``distance'' between real data distribution and predictive distribution of the model and improve the performances of multi-class recognition tasks. Specifically, our method surpass the existing state-of-the-art counterparts on all evaluating benchmarks. We further conduct extensive experiments on different backbones to verify our superior robustness. It is demonstrated that our method successfully pushes the corresponding performance boundaries. Finally, we perform experiments on more challenging scenarios, \textit{i.e.}, learning with incomplete or noisy data, revealing that our method exhibits a high tolerance to such corrupted data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。