针对模态缺失问题,提出隐空间恢复框架,直接利用可用模态进行鲁棒预测。
Latent World Recovery for Multimodal Learning with Missing Modalities

- 将不同模态嵌入对齐到共享隐空间,实现部分可观测下的统一表示
- 在真实多组学数据上验证,癌症表型分类准确率达87.3%
- 适合生物医学中模态不全场景,无需补全缺失模态
我们研究模态缺失条件下的多模态学习,尤其针对生物科学应用中异构模态常部分不可用的问题。提出隐世界恢复(LWR)框架,核心思想是:(i) 不同模态的特定嵌入在共享隐空间中对齐;(ii) 仅融合实际可用模态的嵌入构建统一表示。相比补全缺失模态或固定模态集合,LWR将每种模态视为潜在状态的部分感知,直接从观测模态进行可用性感知的表示学习。邻域驱动的隐空间对齐与可用性感知融合相结合,使模型在部分观测下仍具鲁棒性,且避免了缺失模态显式重建带来的误差传播。我们在真实世界的不完整多组学基准上评估该框架,证明其在癌症表型分类和生存预测等下游任务中具有高效性。
原文摘要 · Abstract (English)
We study multimodal learning under missing modalities, with particular motivation from bioscience applications in which heterogeneous modalities are often only partially available when decisions need to be made. We propose Latent World Recovery (LWR), a framework built on two key ideas: (i) modality-specific embeddings from different modalities are aligned in a shared latent space, and (ii) a unified representation is constructed by fusing only the embeddings of the modalities that are actually available at both training and inference time. Rather than imputing missing modalities or requiring a fixed modality set, LWR treats each modality as a partial perception of an underlying latent state and performs availability-aware representation learning directly from the observed modalities. This combination of neighbor-based latent alignment and availability-aware modality fusion enables robust multimodal prediction under partial observation, while avoiding error propagation from explicit reconstruction of missing modalities. We evaluate the proposed framework on real-world incomplete multi-omics benchmarks and demonstrate that it provides an effective approach to downstream tasks such as cancer phenotype classification and survival prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。