解决多视图数据缺失标签问题,用生成模型融合有标签和无标签数据提升预测与补全效果。
A Semi-supervised Generative Model for Incomplete Multi-view Data Integration with Missing Labels
- 基于生成模型统一处理有标签与无标签数据,共享潜在空间提升泛化能力。
- 在图像和多组学数据上,相比现有方法,预测准确率与缺失值补全效果均更优。
- 适合标签稀疏、视图不完整的真实生物医学与图像数据分析场景。
多视图学习广泛应用于真实数据集,如多组学生物数据,但常面临视图缺失与标签缺失问题。先前的概率方法通过专家乘积机制聚合可用视图表示,在信息瓶颈(IB)原理下表现优于确定性分类器,但其框架本质为全监督,无法利用无标签数据。本文提出一种半监督生成模型,将有标签与无标签样本统一建模。该方法最大化无标签样本的似然,学习与带标签数据共享的潜在空间;同时在潜在空间中进行跨视图互信息最大化,增强跨视图共享信息的提取。相比现有方法,本模型在具有缺失视图和少量标签样本的图像与多组学数据上,实现了更优的预测性能与数据补全效果。
原文摘要 · Abstract (English)
Multi-view learning is widely applied to real-life datasets, such as multiple omics biological data, but it often suffers from both missing views and missing labels. Prior probabilistic approaches addressed the missing view problem by using a product-of-experts scheme to aggregate representations from present views and achieved superior performance over deterministic classifiers, using the information bottleneck (IB) principle. However, the IB framework is inherently fully supervised and cannot leverage unlabeled data. In this work, we propose a semi-supervised generative model that utilizes both labeled and unlabeled samples in a unified framework. Our method maximizes the likelihood of unlabeled samples to learn a latent space shared with the IB on labeled data. We also perform cross-view mutual information maximization in the latent space to enhance the extraction of shared information across views. Compared to existing approaches, our model achieves better predictive and imputation performance on both image and multi-omics data with missing views and limited labeled samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。