arXiv:2604.05462stat.MLcs.LG2026-04

提出分层对比学习,更精准分离多模态数据中的共享与私有信息。

Hierarchical Contrastive Learning for Multimodal Data

论文配图:Hierarchical Contrastive Learning for Multimodal Data
图 1 · 摘自论文原文
  • 构建分层潜在变量模型,区分全局、部分共享和模态特有表示
  • 在电子病历数据上提升预测性能,准确恢复层次结构
  • 适合需要精细理解多模态交互的医疗、跨模态分析场景

多模态表征学习通常基于共享-私有分解,将潜在信息视为所有模态共有或单一模态特有。这种二元视角常不充分:许多因素仅被部分模态共享,忽略此类部分共享会过度对齐无关信号并掩盖互补信息。我们提出分层对比学习(HCL),在统一模型中学习全局共享、部分共享及模态特定表示。HCL结合分层潜在变量形式、结构稀疏性与结构感知对比目标,仅对真正共享潜因子的模态进行对齐。在潜在变量无相关条件下,我们证明了分层分解的可辨识性,建立了载荷矩阵的恢复保证,并推导了下游预测的参数估计与过失风险界。模拟实验显示能准确恢复分层结构并有效选择任务相关成分。在多模态电子健康记录数据上,HCL生成更具信息量的表示,并持续提升预测性能。

原文摘要 · Abstract (English)

Multimodal representation learning is commonly built on a shared-private decomposition, treating latent information as either common to all modalities or specific to one. This binary view is often inadequate: many factors are shared by only subsets of modalities, and ignoring such partial sharing can over-align unrelated signals and obscure complementary information. We propose Hierarchical Contrastive Learning (HCL), a framework that learns globally shared, partially shared, and modality-specific representations within a unified model. HCL combines a hierarchical latent-variable formulation with structural sparsity and a structure-aware contrastive objective that aligns only modalities that genuinely share a latent factor. Under uncorrelated latent variables, we prove identifiability of the hierarchical decomposition, establish recovery guarantees for the loading matrices, and derive parameter estimation and excess-risk bounds for downstream prediction. Simulations show accurate recovery of hierarchical structure and effective selection of task-relevant components. On multimodal electronic health records, HCL yields more informative representations and consistently improves predictive performance.

多模态学习对比学习分层结构医疗数据分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。