从无配对数据中学习结构化潜在变量,实现半监督聚类与配对样本生成
CASL-VAE: Learning Structured Latent Variables from Unpaired Data for Semi-supervised Clustering and Paired Sample Generation
- 通过对比学习分离共性因子与目标特异性子类型,建模异质变异
- 在无配对数据下实现跨域联合似然优化,提升亚型识别与生成能力
- 适用于阿尔茨海默病等复杂疾病中的生物异质性分析
量化目标群体相对于参考群体的变异是许多科学与临床问题的核心(如疾病与健康对照)。然而,在缺乏配对数据且目标群体存在异质性变异的情况下,现有方法难以分离多种目标特异性变异模式。本文提出CASL-VAE,一种深度对比潜在变量模型,可从无配对数据中学习结构化潜在生成因子。该模型将变异分解为跨群体共享的连续共性潜因子,以及分层的显著潜因子,后者以离散亚型和亚型内连续变异的形式建模目标特异性异质性。利用变分推断,我们展示了如何在无配对数据下对参考与目标域进行近似联合似然优化,为配对样本生成与跨域分析提供理论基础。我们在半合成神经影像数据上验证了CASL-VAE,结果表明其在亚型恢复与配对样本生成方面优于基线聚类与生成模型。此外,该方法还能揭示阿尔茨海默病中具有生物学合理性的异质性特征。
原文摘要 · Abstract (English)
Quantifying variability in a target population relative to a reference population is central to many scientific and clinical problems (e.g., diseased vs. healthy). Yet, without paired data and in the presence of heterogeneous target variation, existing methods struggle to separate multiple modes of target-specific variation. We propose \textit{CASL-VAE}, a deep contrastive latent variable model that learns structured latent generative factors from unpaired data. CASL-VAE factorizes variation into continuous common latent factors shared across populations and hierarchical salient latent factors that model target-specific heterogeneity as discrete subtypes and continuous within-subtype variation. Using variational inference, we show how approximate joint likelihood optimization over reference and target domains can be performed using unpaired data, providing a principled basis for paired-sample generation and cross-domain analysis. We validate CASL-VAE on semi-synthetic neuroimaging data, demonstrating improved subtype recovery and paired-sample generation compared to baseline clustering and generative models. We also validate its ability to reveal biologically plausible heterogeneity in Alzheimer's disease.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。