用潜在细胞表征预测代替基因重建,提升空间转录组模型泛化能力。
CellWorld: From Gene-Level Reconstruction to Latent Cell Prediction in Spatial Transcriptomics Foundation Models

- 将预测目标从基因表达改为潜在细胞表征,避免复制技术噪声。
- 小模型(574万参数)在11项线性探测和7项微调任务中全面超越基线。
- 仅用5%数据预训练的大模型冻结后仍优于全微调基线,适合低资源场景。
本文表明,潜在空间预测预训练可为空间转录组学提供可扩展的基础模型。现有模型多通过重建被遮蔽的基因身份或表达值,可能诱发特定检测技术的变异,并限制表征迁移性。为避免直接重建此类变异,我们改用可见空间上下文与有限部分表达提示,预测遮蔽细胞的潜在表征,提出CellWorld模型。我们在包含4600万个人类细胞的语料库上预训练了四个变体,参数量从574万到9456万不等。控制性缩放实验显示,性能随模型容量提升,尤其在空间任务上;空间迁移更依赖充分优化与广泛的生物来源多样性,而非单纯细胞数量。在四个保留数据集上,即使最小的CellWorld-Small(574万参数)也在全部11个线性探测基准和7个微调空间基准上超越所有基线。尤为显著的是,一个仅用5%语料库、覆盖广泛生物来源的冻结版CellWorld-Large,在所有7个空间基准上均优于完全微调的基线。代码已开源。
原文摘要 · Abstract (English)
This paper shows that latent-space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics. Existing spatial transcriptomics foundation models primarily reconstruct masked gene identities or expression values, potentially encouraging the reproduction of assay-specific technical variation and limiting representation transferability. To avoid directly reconstructing such variation, we shift the prediction target from observed gene measurements to latent cell representations and introduce CellWorld, which predicts the latent representations of masked cells from visible spatial context and a limited partial-expression hint. We pretrain four CellWorld variants, spanning 5.74M to 94.56M trainable parameters, on a corpus of 46 million human cells. Our controlled scaling experiments show that performance improves with model capacity, particularly on spatial tasks, while spatial transfer depends more on sufficient optimization and broad biological source diversity than on cell count alone. Across four held-out datasets, even CellWorld-Small, with 5.74M trainable parameters, outperforms every baseline on all 11 linear-probe benchmarks and all seven fine-tuned spatial benchmarks. Most notably, a frozen CellWorld-Large pretrained on only 5\% of the corpus with broad biological source coverage outperforms every fully fine-tuned baseline across all seven spatial benchmarks. Code is available at https://github.com/UoM-HealthAI/CellWorld.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。