解决多模态数据缺失时的融合分析问题,提升聚类效果。
Generalized probabilistic canonical correlation analysis for multi-modal data integration with full or partial observations
- 基于概率模型统一处理多模态数据,支持任意数量模态融合。
- 在多种缺失模式下表现稳健,在模拟和真实数据中聚类准确率更高。
- 适合生物信息学、医学图像等存在数据缺失的多模态研究者使用。
多模态数据分析在生物信息学等领域日益重要。随着数据量和复杂度上升,亟需能整合不同模态并利用其互补信息以提高聚类精度与洞察力的计算模型,尤其在存在部分观测缺失的情况下。本文提出广义概率典型相关分析(GPCCA),一种无监督的多模态数据整合与联合降维方法。该模型通过在建模中直接处理缺失值,支持超过两种模态的整合,并识别各模态内的相关特征,同时保留模态内相关性。实验表明,该方法对各类缺失模式具有鲁棒性,在多个仿真设置中优于现有方法,能有效捕捉跨模态的核心模式。此外,我们在TCGA癌症组学数据集和多视角图像数据集上验证了其适用性。结果表明,GPCCA可生成高信息量的低维嵌入,适用于下游聚类与分析。为促进应用,我们已发布R包GPCCA,开源地址:https://github.com/Kaversoniano/GPCCA。
原文摘要 · Abstract (English)
Background: The integration and analysis of multi-modal data are increasingly essential across various domains including bioinformatics. As the volume and complexity of such data grow, there is a pressing need for computational models that not only integrate diverse modalities but also leverage their complementary information to improve clustering accuracy and insights, especially when dealing with partial observations with missing data. Results: We propose Generalized Probabilistic Canonical Correlation Analysis (GPCCA), an unsupervised method for the integration and joint dimensionality reduction of multi-modal data. GPCCA addresses key challenges in multi-modal data analysis by handling missing values within the model, enabling the integration of more than two modalities, and identifying informative features while accounting for correlations within individual modalities. The model demonstrates robustness to various missing data patterns and provides low-dimensional embeddings that facilitate downstream clustering and analysis. In a range of simulation settings, GPCCA outperforms existing methods in capturing essential patterns across modalities. Additionally, we demonstrate its applicability to multi-omics data from TCGA cancer datasets and a multi-view image dataset. Conclusion: GPCCA offers a useful framework for multi-modal data integration, effectively handling missing data and providing informative low-dimensional embeddings. Its performance across cancer genomics and multi-view image data highlights its robustness and potential for broad application. To make the method accessible to the wider research community, we have released an R package, GPCCA, which is available at https://github.com/Kaversoniano/GPCCA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。