arXiv:2602.04920cs.LGcs.SD2026-02NeurIPS被引 3

提出循环信息隐空间,让模型在模态缺失时仍能稳定表现。

CyIN: Cyclic Informative Latent Space for Bridging Complete and Incomplete Multimodal Learning

  • 用循环信息瓶颈构建任务相关隐空间,提升跨模态交互效率。
  • 在4个数据集上,完整与不完整场景下均优于现有方法。
  • 适合真实世界中模态动态缺失的多模态应用。

多模态机器学习模仿人脑融合多种感知能力,但多数模型依赖完全配对的数据训练,在真实场景中因模态缺失导致性能显著下降。本文提出循环信息学习框架(CyIN),通过在不同模态间交替使用标记级和词元级信息瓶颈(IB),构建任务相关的可变性近似隐空间,净化信息以增强跨模态交互与融合。为弥补不完整输入造成的缺失信息,设计了正向与反向传播的跨模态循环重建机制,恢复缺失模态。借助提取与重构的信息隐变量,CyIN 在统一模型中联合优化完整与不完整多模态学习。在4个多模态数据集上的实验表明,该方法在完整及多样缺失场景下均表现出色。

原文摘要 · Abstract (English)

Multimodal machine learning, mimicking the human brain's ability to integrate various modalities has seen rapid growth. Most previous multimodal models are trained on perfectly paired multimodal input to reach optimal performance. In real-world deployments, however, the presence of modality is highly variable and unpredictable, causing the pre-trained models in suffering significant performance drops and fail to remain robust with dynamic missing modalities circumstances. In this paper, we present a novel Cyclic INformative Learning framework (CyIN) to bridge the gap between complete and incomplete multimodal learning. Specifically, we firstly build an informative latent space by adopting token- and label-level Information Bottleneck (IB) cyclically among various modalities. Capturing task-related features with variational approximation, the informative bottleneck latents are purified for more efficient cross-modal interaction and multimodal fusion. Moreover, to supplement the missing information caused by incomplete multimodal input, we propose cross-modal cyclic translation by reconstruct the missing modalities with the remained ones through forward and reverse propagation process. With the help of the extracted and reconstructed informative latents, CyIN succeeds in jointly optimizing complete and incomplete multimodal learning in one unified model. Extensive experiments on 4 multimodal datasets demonstrate the superior performance of our method in both complete and diverse incomplete scenarios.

多模态学习隐空间建模模态缺失信息瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。