提出两种生成模型,有效学习多维流形并列的数据分布。
A Deep Generative Approach to Stratified Learning

- 用分层VAE和扩散模型捕捉不同维度流形的联合分布
- 理论证明收敛速率依赖于内在维度与光滑性
- 能准确估计流形数量和各自维度,适合复杂数据建模
尽管流形假设在现代机器学习中被广泛采用,但复杂数据更宜建模为不同维度流形(即‘层’)的并集。分层学习因维度差异、交叠奇点及缺乏高效模型而困难。本文提出两种深度生成框架来学习分层空间上的分布:一是基于维度感知的变分自编码器混合的筛子最大似然方法;二是探索混合分布得分场结构的扩散模型。我们建立了对环境分布和内在分布的学习收敛率,其依赖于各层的内在维度与光滑性。利用得分场几何,还建立了对每层内在维度估计的一致性,并提出可一致估计层数与维度的算法。两类框架的理论结果揭示了底层几何、环境噪声水平与深度生成模型间的相互作用。大量模拟及真实数据应用(如分子动力学)验证了方法有效性。
原文摘要 · Abstract (English)
While the manifold hypothesis is widely adopted in modern machine learning, complex data is often better modeled as stratified spaces -- unions of manifolds (strata) of varying dimensions. Stratified learning is challenging due to varying dimensionality, intersection singularities, and lack of efficient models in learning the underlying distributions. We provide a deep generative approach to stratified learning by developing two generative frameworks for learning distributions on stratified spaces. The first is a sieve maximum likelihood approach realized via a dimension-aware mixture of variational autoencoders. The second is a diffusion-based framework that explores the score field structure of a mixture. We establish the convergence rates for learning both the ambient and intrinsic distributions, which are shown to be dependent on the intrinsic dimensions and smoothness of the underlying strata. Utilizing the geometry of the score field, we also establish consistency for estimating the intrinsic dimension of each stratum and propose an algorithm that consistently estimates both the number of strata and their dimensions. Theoretical results for both frameworks provide fundamental insights into the interplay of the underlying geometry, the ambient noise level, and deep generative models. Extensive simulations and real dataset applications, such as molecular dynamics, demonstrate the effectiveness of our methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。