arXiv:2603.26827cs.LGcs.AI2026-03

用少量数据生成逼真组织图像,提升稀缺空间转录组预测精度

Central-to-Local Adaptive Generative Diffusion Framework for Improving Gene Expression Prediction in Data-Limited Spatial Transcriptomics

  • 先用大样本病理图像预训练全局模型,再用少量基因数据微调本地模型
  • 生成图像与真实数据在多个器官中嵌入重合度高,细胞组成还原准确
  • 适合空间转录组数据少的研究者,可显著降低实验采样成本

空间转录组(ST)可在完整组织结构中提供空间解析的基因表达谱,实现组织学背景下的分子分析。然而,高成本、低通量和数据共享受限导致严重数据稀缺,制约了计算模型的发展。为此,我们提出中心到本地自适应生成扩散框架(C2L-ST),将大规模形态先验与有限分子引导相结合。首先在广泛病理数据集上预训练全局中心模型以学习可迁移的形态表征,随后通过轻量级基因条件调制,在少量配对的图像-基因点下适配机构特异的局部模型。该策略在数据受限条件下合成出视觉与结构保真度高的真实组织图像块,能有效还原细胞组成,并在多个器官中与真实数据呈现强嵌入重合性,体现真实性和多样性。将合成图像-基因对用于下游训练,可显著提升基因表达预测准确率与空间一致性,性能接近真实数据,且仅需极少量采样点。C2L-ST为分子层面的数据增强提供可扩展、高效的数据方案,是一种适用于空间生物学及相关领域的领域自适应、泛化性强的组织学与转录组整合方法。

原文摘要 · Abstract (English)

Spatial Transcriptomics (ST) provides spatially resolved gene expression profiles within intact tissue architecture, enabling molecular analysis in histological context. However, the high cost, limited throughput, and restricted data sharing of ST experiments result in severe data scarcity, constraining the development of robust computational models. To address this limitation, we present a Central-to-Local adaptive generative diffusion framework for ST (C2L-ST) that integrates large-scale morphological priors with limited molecular guidance. A global central model is first pretrained on extensive histopathology datasets to learn transferable morphological representations, and institution-specific local models are then adapted through lightweight gene-conditioned modulation using a small number of paired image-gene spots. This strategy enables the synthesis of realistic and molecularly consistent histology patches under data-limited conditions. The generated images exhibit high visual and structural fidelity, reproduce cellular composition, and show strong embedding overlap with real data across multiple organs, reflecting both realism and diversity. When incorporated into downstream training, synthetic image-gene pairs improve gene expression prediction accuracy and spatial coherence, achieving performance comparable to real data while requiring only a fraction of sampled spots. C2L-ST provides a scalable and data-efficient framework for molecular-level data augmentation, offering a domain-adaptive and generalizable approach for integrating histology and transcriptomics in spatial biology and related fields.

空间转录组生成模型数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。