通过双视角自蒸馏,让显微图像学会捕捉细胞微环境信息。
MAD: Microenvironment-Aware Distillation -- A Pretraining Strategy for Virtual Spatial Omics from Microscopy
- 联合形态与微环境视图,统一学习细胞嵌入表示。
- 在多种组织和成像模式下,下游任务表现领先现有方法。
- 适合希望从海量显微图像中挖掘生物信息的研究者。
将显微成像与组学结合,可在单细胞分辨率和组织尺度上无成本地读取分子状态,突破传统组学技术的代价与通量限制。自监督预训练提供了一种低标签依赖的可扩展方案,但如何编码细胞在组织微环境中的身份及其可捕获的生物学信息仍不明确。本文提出MAD(微环境感知蒸馏),一种通过联合自蒸馏同一细胞的形态视图与微环境视图,将其映射至统一嵌入空间的预训练策略。在多种组织类型与成像模态下,MAD在细胞亚型分类、转录组预测及生信推断等下游任务中均达到当前最优性能;其表现甚至优于参数量相近但在更大数据集上训练的基础模型。结果表明,双视图联合自蒸馏能有效捕捉组织内细胞的复杂性与多样性。MAD因此成为显微图像表示学习的通用工具,为虚拟空间组学与海量显微数据中的生物洞见提供支持。
原文摘要 · Abstract (English)
Bridging microscopy and omics would allow us to read molecular states from images-at single-cell resolution and tissue scale-without the cost and throughput limits of omics technologies. Self-supervised pretraining offers a scalable approach with minimal labels, yet how to encode single-cell identity within tissue environments-and the extent of biological information such models can capture-remains an open question. Here, we introduce MAD (microenvironment-aware distillation), a pretraining strategy that learns cell-centric embeddings by jointly self-distilling the morphology view and the microenvironment view of the same indexed cell into a unified embedding space. Across diverse tissues and imaging modalities, MAD achieves state-of-the-art prediction performance on downstream tasks including cell subtyping, transcriptomic prediction, and bioinformatic inference. MAD even outperforms foundation models with a similar number of model parameters that have been trained on substantially larger datasets. These results demonstrate that MAD's dual-view joint self-distillation effectively captures the complexity and diversity of cells within tissues. Together, this establishes MAD as a general tool for representation learning in microscopy, enabling virtual spatial omics and biological insights from vast microscopy datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。