整合多源分子嵌入,构建通用生物表示平台。
Platform for Representation and Integration of multimodal Molecular Embeddings
- 用自编码器融合基因组、文本和知识图谱嵌入
- 不同模态嵌入信号重叠度低,集成后性能更优
- 适合需跨数据源建模的生物医药研究者
现有分子(如基因)嵌入方法受限于特定任务或数据模态,难以全面捕捉基因功能与相互作用。本研究系统评估了来自三大数据源(组学实验数据、文献文本数据、知识图谱)的生物分子表征,采用改进的奇异向量典型相关分析(SVCCA)量化不同模态间信号冗余与互补性。结果表明,现有嵌入捕获的分子信号高度非重叠,凸显嵌入集成的价值。据此提出PRISME平台,基于自编码器将异构嵌入统一为多模态表示。在多个基准任务中验证,该方法表现稳定,且在缺失值补全任务中优于单一嵌入方法。该框架支持生物分子的综合性建模,推动面向下游生物医学机器学习应用的鲁棒、通用嵌入发展。
原文摘要 · Abstract (English)
Existing machine learning methods for molecular (e.g., gene) embeddings are restricted to specific tasks or data modalities, limiting their effectiveness within narrow domains. As a result, they fail to capture the full breadth of gene functions and interactions across diverse biological contexts. In this study, we have systematically evaluated knowledge representations of biomolecules across multiple dimensions representing a task-agnostic manner spanning three major data sources, including omics experimental data, literature-derived text data, and knowledge graph-based representations. To distinguish between meaningful biological signals from chance correlations, we devised an adjusted variant of Singular Vector Canonical Correlation Analysis (SVCCA) that quantifies signal redundancy and complementarity across different data modalities and sources. These analyses reveal that existing embeddings capture largely non-overlapping molecular signals, highlighting the value of embedding integration. Building on this insight, we propose Platform for Representation and Integration of multimodal Molecular Embeddings (PRISME), a machine learning based workflow using an autoencoder to integrate these heterogeneous embeddings into a unified multimodal representation. We validated this approach across various benchmark tasks, where PRISME demonstrated consistent performance, and outperformed individual embedding methods in missing value imputations. This new framework supports comprehensive modeling of biomolecules, advancing the development of robust, broadly applicable multimodal embeddings optimized for downstream biomedical machine learning applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。