通过信息论方法提升多模态学习在数据缺失时的鲁棒性
Robult: Leveraging Redundancy and Modality Specific Features for Robust Multimodal Learning
- 用软PU对比损失最大化任务相关特征对齐,高效利用少量标注数据
- 引入潜在重建损失保留各模态特有信息,应对模态缺失问题
- 轻量设计易集成,适合真实场景下的多模态系统
解决模态缺失和标注数据有限问题是推进鲁棒多模态学习的关键。我们提出Robult,一种可扩展框架,通过保留模态特有信息并利用冗余,采用新颖的信息论方法缓解上述挑战。Robult优化两个核心目标:(1) 软正-未标记(PU)对比损失,在半监督设置中最大化任务相关特征对齐,有效利用有限标注数据;(2) 潜在重建损失,确保独特模态特有信息得以保留。这些策略嵌入模块化设计,显著提升多种下游任务表现,并在推理阶段对不完整模态具备强韧性。跨多个数据集的实验验证,Robult在半监督学习和模态缺失场景下均优于现有方法。其轻量级设计促进可扩展性,便于与现有架构无缝集成,适用于真实世界多模态应用。
原文摘要 · Abstract (English)
Addressing missing modalities and limited labeled data is crucial for advancing robust multimodal learning. We propose Robult, a scalable framework designed to mitigate these challenges by preserving modality-specific information and leveraging redundancy through a novel information-theoretic approach. Robult optimizes two core objectives: (1) a soft Positive-Unlabeled (PU) contrastive loss that maximizes task-relevant feature alignment while effectively utilizing limited labeled data in semi-supervised settings, and (2) a latent reconstruction loss that ensures unique modality-specific information is retained. These strategies, embedded within a modular design, enhance performance across various downstream tasks and ensure resilience to incomplete modalities during inference. Experimental results across diverse datasets validate that Robult achieves superior performance over existing approaches in both semi-supervised learning and missing modality contexts. Furthermore, its lightweight design promotes scalability and seamless integration with existing architectures, making it suitable for real-world multimodal applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。