基于中心法则设计的多组学模型,提升癌症研究中跨任务、跨数据的泛化能力。
DoGMA: A Central-Dogma-Guided Foundation Model for Multi-Omics Alignment and Multi-Task Learning in Oncology

- 采用有向注意力机制,模拟中心法则的信息流动方向。
- 在生存预测与转移预测任务中均表现优异,优于传统双向模型。
- 适合需要生物可解释性的癌症多组学分析场景。
注意力机制在现代深度学习中广泛应用,现有大多数多组学模型沿用传统的双向交互方式,忽略了生命活动的基本方向性。本文提出DoGMA,一种面向泛癌多组学分析的中心法则引导型基础模型,认为鲁棒的迁移能力依赖于具有领域特定归纳偏置的表征。具体而言,模型基于Transformer-MoE架构,通过有向注意力限制组间通信方向,使其符合中心法则信息流。进一步采用分层组学掩码重建进行预训练,引导模型学习与中心法则一致的跨组学交互。在癌症表征学习、生存预测和转移预测等多样化下游任务中,DoGMA始终表现出强劲的预测性能。消融实验表明,性能提升源于有向注意力与重建预训练之间的协同作用,促进更符合生物学逻辑的跨组学信息交换。结果表明,引入领域特定归纳偏置可显著提升多组学基础模型的鲁棒性与迁移能力,为注意力机制在多组学表征学习中的设计提供新思路。
原文摘要 · Abstract (English)
Attention mechanisms have been widely utilized in modern deep learning, and many existing multi-omics models inherit their conventional use to allow unrestricted bidirectional interactions. However, the fundamental logic of life is directional. Existing designs often overlook the directionality suggested by the central dogma, potentially limiting transfer across heterogeneous cancers, downstream tasks, and incomplete modality settings.In this work, we present DoGMA, a central-dogma-guided foundation model for pan-cancer multi-omics analysis, arguing that robust transfer requires representations with domain-specific inductive bias. Concretely, we build it on a Transformer-MoE architecture where directed attention biases inter-omics communication toward central-dogma information flow. We further pretrain our model with masked hierarchical omics reconstruction to guide it toward learning central-dogma-consistent interactions. Across diverse downstream tasks, including cancer representation learning, survival prediction, and metastasis prediction, DoGMA consistently demonstrates strong predictive performance. Ablations and analyses further suggest that the performance gains arise from the synergy between central-dogma-guided directed attention and reconstruction-based pretraining, which together promote more biologically consistent cross-omics information exchange. Overall, DoGMA demonstrates that domain-specific inductive biases can improve the robustness and transferability of multi-omics foundation models, offering new insights into the design of attention mechanisms for multi-omics representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。