arXiv:2608.14355cs.AI2026-08

分离共享与特有信息,提升空间转录组与形态学数据融合效果

Disentangled Shared Representations Improve Morpho-Transcriptomic Integration

论文配图:Disentangled Shared Representations Improve Morpho-Transcriptomic Integration
图 1 · 摘自论文原文
  • 通过显式分解共享与模态特有特征,优化多模态表示学习
  • 对比学习模型在下游任务中表现优于变分自编码器
  • 该方法适合开发空间转录组领域的基础模型,提升可解释性

空间转录组学(ST)可同时分析基因表达与组织形态,为学习捕捉共现结构的多模态表示提供了机会。然而,标准多模态模型常将不同模态压缩至同一潜在空间,未显式区分共享与模态特有变异源,可能限制下游应用。本文研究在配对的苏木精-伊红(H&E)与空间转录组数据上,显式解耦共享与私有潜在成分是否能改善多模态表示学习。我们在两个癌症队列中,于相同实验条件下比较基于变分自编码器(VAE)与对比学习的方法,每类均包含标准与解耦变体。表示性能通过跨模态重建、下游探测及跨模态探针迁移进行评估。结果表明:第一,对比学习目标在下游探测任务中表现优于基于VAE的模型;第二,解耦变体在选定重建与探测指标上表现更优,但增益程度依赖于模型类型、任务、方向及解耦强度。总体而言,显式分解共享与模态特有信息有助于提升空间转录组多模态表示学习,并为未来基础模型提供有效评估框架。

原文摘要 · Abstract (English)

Spatial transcriptomics (ST) enables the simultaneous profiling of gene expression and tissue morphology, creating an opportunity to learn multimodal representations capturing shared morpho-transcriptomic structure. However, standard multimodal models often compress modalities into a common latent space without explicitly separating shared and modality-specific sources of variation, which may limit downstream utility. We investigate whether explicit disentanglement of shared and private latent components improves multimodal representation learning for paired Hematoxylin \& Eosin (H\&E) and ST data. We compare VAE-based and contrastive approaches, each in standard and disentangled variants, across two cancer cohorts under matched experimental conditions. Representations are evaluated using cross-modal reconstruction, downstream probing and cross-modal probe transfer. The experiments suggest two main trends. First, contrastive objectives yield higher downstream probing performance than VAE-based models. Second, disentangled variants improve the selected reconstruction and probing metrics, although the gains depend on the model family, task, direction, and disentanglement strength. Overall, our results suggest that explicitly factorizing shared and modality-specific information can improve multimodal representation learning for spatial transcriptomics and provides a useful evaluation framework for future foundation models.

多模态学习空间转录组表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。