首个融合病理图像与多组学数据的自监督预训练模型
Masked Omics Modeling for Multimodal Representation Learning across Histopathology and Molecular Profiles

- 用掩码组学建模让模型学习图像与组学间的跨模态关联
- 在多种癌症任务中表现超越监督与自监督基线方法
- 支持任意组学数据从病理图像中重建,适合多模态研究者
自监督学习推动了计算病理学发展,使从病理图像中学习丰富表征成为可能。然而,仅依靠组织分析难以捕捉更广泛的分子复杂性,关键互补信息存在于转录组、甲基化组和基因组等高维组学数据中。为此,我们提出MORPHEUS,首个将病理图像与多组学数据整合到共享Transformer架构中的多模态预训练策略。其核心是新颖的掩码组学建模目标,促使模型学习有意义的跨模态关系。该模型可作为通用预训练编码器,单独用于病理图像或结合任意组学子集。除推理外,还支持任意组学数据从包含病理图像的任意模态子集重建。在大规模泛癌队列上预训练后,MORPHEUS在多种任务和模态组合中显著优于监督及自监督基线。这些能力使其成为肿瘤学多模态基础模型的重要方向。代码已公开于https://github.com/Lucas-rbnt/MORPHEUS。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) has driven major advances in computational pathology by enabling the learning of rich representations from histopathology data. Yet, tissue analysis alone may fall short in capturing broader molecular complexity, as key complementary information resides in high-dimensional omics profiles such as transcriptomics, methylomics, and genomics. To address this gap, we introduce MORPHEUS, the first multimodal pre-training strategy that integrates histopathology images and multi-omics data within a shared transformer-based architecture. At its core, MORPHEUS relies on a novel masked omics modeling objective that encourages the model to learn meaningful cross-modal relationships. This yields a general-purpose pre-trained encoder that can be applied to histopathology alone or in combination with any subset of omics modalities. Beyond inference, MORPHEUS also supports flexible any-to-any omics reconstruction, enabling one or more omics profiles to be reconstructed from any modality subset that includes histopathology. Pre-trained on a large pan-cancer cohort, MORPHEUS shows substantial improvements over supervised and SSL baselines across diverse tasks and modality combinations. Together, these capabilities position it as a promising direction for the development of multimodal foundation models in oncology. Code is publicly available at https://github.com/Lucas-rbnt/MORPHEUS
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。