通过多层级自适应去混淆,提升多模态模型分类可靠性
Reliable Multimodal Learning Via Multi-Level Adaptive DeConfusion
- 设计动态退出编码器与跨类残差重建,消除全局层面类别混淆
- 利用低混淆模态先验,实现样本级跨模态修正,显著降低误判
- 在多个数据集上优于现有方法,尤其适用于噪声数据场景
多模态学习通过融合不同模态的互补信息提升各类机器学习任务性能。然而,现有方法常保留大量类别间混淆,导致难以获得高置信度预测,尤其在低质量或噪声数据的真实场景中。为此,本文提出多层级自适应去混淆(MLAD),在全局和样本两个层面消除多模态数据中的类别混淆,显著提升多模态模型的分类可靠性。具体而言,MLAD首先通过动态退出模态编码器和跨类残差重建机制,学习去除全局混淆的类级潜在分布;随后,基于低混淆模态特征构建无混淆模态先验,通过高斯混合模型筛选低混淆特征,并指导样本级跨模态修正。实验表明,MLAD在多个基准测试中均优于当前最优方法,展现出更强的可靠性。
原文摘要 · Abstract (English)
Multimodal learning enhances the performance of various machine learning tasks by leveraging complementary information across different modalities. However, existing methods often learn multimodal representations that retain substantial inter-class confusion, making it difficult to achieve high-confidence predictions, particularly in real-world scenarios with low-quality or noisy data. To address this challenge, we propose Multi-Level Adaptive DeConfusion (MLAD), which eliminates inter-class confusion in multimodal data at both global and sample levels, significantly enhancing the classification reliability of multimodal models. Specifically, MLAD first learns class-wise latent distributions with global-level confusion removed via dynamic-exit modality encoders that adapt to the varying discrimination difficulty of each class and a cross-class residual reconstruction mechanism. Subsequently, MLAD further removes sample-specific confusion through sample-adaptive cross-modality rectification guided by confusion-free modality priors. These priors are constructed from low-confusion modality features, identified by evaluating feature confusion using the learned class-wise latent distributions and selecting those with low confusion via a Gaussian mixture model. Experiments demonstrate that MLAD outperforms state-of-the-art methods across multiple benchmarks and exhibits superior reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。