提出新框架DisCoVR,更好分离共享与特有特征。
Variational Learning of Disentangled Representations
- 基于概率生成结构设计目标函数,加入对抗项防止信息泄露。
- 在图像和单细胞数据上实现更强解耦,共享与特有表示均保持信息量。
- 适合需要高质量解耦表示的生物、医学或跨域泛化任务。
解耦表示需将跨条件共享因素与条件特有因素分离,这对新领域、治疗、患者或物种的泛化至关重要。现有变分方法虽部分实现解耦,但常存在三类问题:未能完全移除条件特有信息、导致条件特有表示无意义,或强加不符合生成过程的独立性假设。本文提出DisCoVR,一种对齐数据生成概率结构的变分框架。其目标函数包含对抗项,防止条件特有信息被编码到条件特有表示中;通过共享与条件特有表示联合重构数据,确保两者均具信息量;并引入结构化先验进一步强化表示有效性。在合成数据、图像及单细胞RNA测序数据集上,DisCoVR均显著优于先前方法,实现更强解耦。
原文摘要 · Abstract (English)
Disentangled representations separate factors that are shared across conditions from those that are condition-specific. Such separation is needed for generalization to new domains, treatments, patients, or species. A dominant line of work pursues this goal through variational formulations. While these approaches achieve partial disentanglement, they often exhibit three common limitations: they either do not remove all condition-specific information from the condition-specific representation, allow the condition-specific representation to become uninformative, or impose independence assumptions that do not reflect the underlying generative process. In this work, we introduce DisCoVR, a variational framework that addresses these limitations. Its objective is aligned with the probabilistic structure of the data-generating process, and includes an adversarial term that prevents condition-specific information from being encoded in the condition-specific representation.DisCoVR reconstructs the data from both shared and condition-specific representations, ensuring that each remains informative, and uses a structured prior that further reinforces the informativeness of both representations. We show that across synthetic, image, and single-cell RNA-sequencing datasets, DisCoVR achieves stronger disentanglement compared to previous approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。