提出SCONE方法,实现大规模多组学数据的高效图嵌入分析。
Subset-Contrastive Multi-Omics Network Embedding
- 基于子图对比学习,提升多组学网络嵌入的可扩展性
- 在单细胞数据中实现细胞类型聚类的协同整合效果
- 适用于大样本数据,适合研究多组学集成与高维数据
动机:基于网络的组学数据分析广泛应用,尽管部分方法已适配单细胞场景,但通常仍存在内存和空间占用过高的问题,更适合批次数据或小规模数据集。此外,多组学中网络方法常依赖相似性构建网络,缺乏结构清晰的拓扑特性,这可能削弱原本为明确结构设计的图方法的有效性。结果:我们提出子图对比多组学网络嵌入(SCONE),通过可扩展的子图对比策略,在大规模数据上应用对比学习技术。利用众多网络方法固有的成对相似性基础,将其转化为优势,实现了高效且可扩展的分析。该方法在单细胞数据中表现出优异的细胞类型聚类整合能力;在批量多组学集成场景下,即使仅使用原始数据的有限视图,性能也达到当前最佳水平。我们预计这些发现将推动子图对比方法在组学数据中的进一步研究。
原文摘要 · Abstract (English)
Motivation: Network-based analyses of omics data are widely used, and while many of these methods have been adapted to single-cell scenarios, they often remain memory- and space-intensive. As a result, they are better suited to batch data or smaller datasets. Furthermore, the application of network-based methods in multi-omics often relies on similarity-based networks, which lack structurally-discrete topologies. This limitation may reduce the effectiveness of graph-based methods that were initially designed for topologies with better defined structures. Results: We propose Subset-Contrastive multi-Omics Network Embedding (SCONE), a method that employs contrastive learning techniques on large datasets through a scalable subgraph contrastive approach. By exploiting the pairwise similarity basis of many network-based omics methods, we transformed this characteristic into a strength, developing an approach that aims to achieve scalable and effective analysis. Our method demonstrates synergistic omics integration for cell type clustering in single-cell data. Additionally, we evaluate its performance in a bulk multi-omics integration scenario, where SCONE performs comparable to the state-of-the-art despite utilising limited views of the original data. We anticipate that our findings will motivate further research into the use of subset contrastive methods for omics data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。