将空间转录组数据当作可裁剪的图像,实现大规模预训练。
Spatial Transcriptomics as Images for Large-Scale Pretraining
- 把组织切片裁成固定大小的图像块作为训练样本,保留空间信息。
- 在基因通道上按规则选子集,提升预训练稳定性与效率。
- 适合做空间转录组下游任务的模型开发者,尤其是病理研究者。
空间转录组(Spatial Transcriptomics, ST)在组织切片的离散位置上测定数千个基因的表达值,并保留精确的空间坐标,对临床和病理研究至关重要。随着测序通量提升和平台进步,数据量激增,推动了大规模ST预训练的需求。然而,预训练的基本单位——即单个训练样本的定义——仍不明确。现有方法分为两类:(1) 将每个点视为独立样本,忽略空间依赖性,使ST退化为单细胞转录组;(2) 将整张切片视为单一样本,导致输入过大、训练样本过少,难以有效预训练。为此,我们提出将空间转录组视为可裁剪的图像。具体地,通过从原始切片中裁剪出固定空间尺寸的图像块,构建多通道图像表示,既保留空间上下文,又显著增加训练样本数量。在通道维度上,设计基因子集选择规则以控制输入维度,提升预训练稳定性。大量实验表明,该图像式数据构造方法显著提升下游任务性能,优于传统预训练方案。消融实验验证了空间裁剪和通道设计均不可或缺,确立了一种统一且实用的ST数据组织范式,支持大规模预训练。
原文摘要 · Abstract (English)
Spatial Transcriptomics (ST) profiles thousands of gene expression values at discrete spots with precise coordinates on tissue sections, preserving spatial context essential for clinical and pathological studies. With rising sequencing throughput and advancing platforms, the expanding data volumes motivate large-scale ST pretraining. However, the fundamental unit for pretraining, i.e., what constitutes a single training sample, remains ill-posed. Existing choices fall into two camps: (1) treating each spot as an independent sample, which discards spatial dependencies and collapses ST into single-cell transcriptomics; and (2) treating an entire slide as a single sample, which produces prohibitively large inputs and drastically fewer training examples, undermining effective pretraining. To address this gap, we propose treating spatial transcriptomics as croppable images. Specifically, we define a multi-channel image representation with fixed spatial size by cropping patches from raw slides, thereby preserving spatial context while substantially increasing the number of training samples. Along the channel dimension, we define gene subset selection rules to control input dimensionality and improve pretraining stability. Extensive experiments show that the proposed image-like dataset construction for ST pretraining consistently improves downstream performance, outperforming conventional pretraining schemes. Ablation studies verify that both spatial patching and channel design are necessary, establishing a unified, practical paradigm for organizing ST data and enabling large-scale pretraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。