提出多层级自适应网络,让多模态分类更可靠。
Multi-QuAD: Multi-Level Quality-Adaptive Dynamic Network for Reliable Multimodal Classification
- 用无噪声原型和免分类器设计,精准评估样本质量。
- 根据样本质量动态调整网络深度与参数,提升可靠性。
- 适合处理不同质量数据的多模态任务,如医疗影像分析。
多模态机器学习在诸多场景中取得显著进展,但其可靠性受样本质量差异影响。本文发现现有可靠多模态分类方法既无法有效估计数据质量,也缺乏针对样本的动态网络结构以实现可靠推理。为此,提出一种名为多层级质量自适应动态网络(Multi-QuAD)的新框架。Multi-QuAD首先采用基于无噪声原型和免分类器设计的方法,在模态与特征层面可靠估计每个样本的质量。随后通过全局置信度归一化深度(GCND)机制实现样本级网络深度自适应,通过跨模态与样本的深度归一化,有效缓解困难模态输入对动态深度可靠性的影响。此外,通过由特征级质量驱动的逐层贪婪参数(LGP)机制,实现样本自适应网络参数。LGP采用跨模态逐层贪婪策略,首次为可变架构的多模态网络设计了可靠的参数预测范式。在四个数据集上的实验表明,Multi-QuAD在分类性能与可靠性上显著优于当前最优方法,对多样质量数据表现出强适应性。
原文摘要 · Abstract (English)
Multimodal machine learning has achieved remarkable progress in many scenarios, but its reliability is undermined by varying sample quality. This paper finds that existing reliable multimodal classification methods not only fail to provide robust estimation of data quality, but also lack dynamic networks for sample-specific depth and parameters to achieve reliable inference. To this end, a novel framework for multimodal reliable classification termed \textit{Multi-level Quality-Adaptive Dynamic multimodal network} (Multi-QuAD) is proposed. Multi-QuAD first adopts a novel approach based on noise-free prototypes and a classifier-free design to reliably estimate the quality of each sample at both modality and feature levels. It then achieves sample-specific network depth via the \textbf{\textit{Global Confidence Normalized Depth (GCND)}} mechanism. By normalizing depth across modalities and samples, \textit{\textbf{GCND}} effectively mitigates the impact of challenging modality inputs on dynamic depth reliability. Furthermore, Multi-QuAD provides sample-adaptive network parameters via the \textbf{\textit{Layer-wise Greedy Parameter (LGP)}} mechanism driven by feature-level quality. The cross-modality layer-wise greedy strategy in \textbf{\textit{LGP}} designs a reliable parameter prediction paradigm for multimodal networks with variable architecture for the first time. Experiments conducted on four datasets demonstrate that Multi-QuAD significantly outperforms state-of-the-art methods in classification performance and reliability, exhibiting strong adaptability to data with diverse quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。