arXiv:2511.20382cs.LGq-bio.GN2025-11

用冻结的预训练模型实现多组学数据的鲁棒对齐与嵌入。

MoRE: Batch-Robust Multi-Omics Representations from Frozen Pre-trained Transformers

  • 通过轻量适配器和融合层,参数高效微调预训练模型。
  • 在跨样本、跨模态任务中表现优异,减少可训练参数。
  • 适合需要低资源、高泛化的多组学分析场景。

多组学数据表征学习面临维度极高、模态异质性强及批次效应等挑战。尽管预训练的Transformer在生物序列建模中展现出良好泛化能力,但其在多组学整合中的应用仍不充分。本文提出MoRE(Multi-Omics Representation Embedding)框架,将冻结的预训练Transformer重用于对齐异构检测数据至共享潜在空间。不同于纯生成方法,MoRE采用参数高效的微调策略,侧重于跨样本与跨模态对齐,而非简单序列重建。具体地,它在冻结主干上附加轻量级、模态特定的适配器和任务自适应融合层,联合优化掩码建模目标、监督对比损失与批次不变对齐损失,生成保持结构特性的嵌入表示,在未见细胞类型和测序平台间具有强泛化能力。在scGPT、scVI、Harmony+Scrublet等基线上的评估显示,MoRE在整合保真度、稀有群体识别与模态迁移方面表现优异,且相比全微调模型显著降低可训练参数量。该工作为通用多组学基础模型提供了一条可行路径。

原文摘要 · Abstract (English)

Representation learning on multi-omics data is challenging due to extreme dimensionality, modality heterogeneity, and cohort-specific batch effects. While pre-trained transformer backbones have shown broad generalization capabilities in biological sequence modeling, their application to multi-omics integration remains underexplored. We present MoRE (Multi-Omics Representation Embedding), a framework that repurposes frozen pre-trained transformers to align heterogeneous assays into a shared latent space. Unlike purely generative approaches, MoRE employs a parameter-efficient fine-tuning (PEFT) strategy, prioritizing cross-sample and cross-modality alignment over simple sequence reconstruction. Specifically, MoRE attaches lightweight, modality-specific adapters and a task-adaptive fusion layer to the frozen backbone. It optimizes a masked modeling objective jointly with supervised contrastive and batch-invariant alignment losses, yielding structure-preserving embeddings that generalize across unseen cell types and platforms. We benchmark MoRE against established baselines, including scGPT, scVI, and Harmony with Scrublet, evaluating integration fidelity, rare population detection, and modality transfer. Our results demonstrate that MoRE achieves competitive batch robustness and biological conservation while significantly reducing trainable parameters compared to fully fine-tuned models. This work positions MoRE as a practical step toward general-purpose omics foundation models.

多组学表示学习预训练模型批效应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。