arXiv:2410.23594cs.LGcs.AI2024-10被引 20

研究生成模型如何在数据子空间内记忆与泛化,揭示其保真机制。

How Do Flow Matching Models Memorize and Generalize in Sample Data Subspaces?

  • 用流匹配模型分析数据子空间内的生成机理,推导最优速度场
  • 生成样本精确复现真实数据点,保持子空间结构不变
  • 提出OSDNet分解速度场,实现子空间内泛化与多样性保持

现实数据通常位于高维空间中的低维结构中。实际应用中仅能观测到有限样本,构成所谓的样本数据子空间,该子空间对降维与生成任务至关重要。核心挑战在于生成模型能否可靠地合成仍位于该子空间内的样本,而非偏离底层结构。本文借助流匹配模型,通过将真实数据分布视为离散分布,在高斯先验下推导出最优速度场的解析表达式,表明生成样本会记忆真实数据点并精确表征样本数据子空间。为处理次优情形,提出正交子空间分解网络(OSDNet),系统性地将速度场分解为子空间与非子空间成分。分析显示,非子空间成分衰减,而子空间成分在样本数据子空间内泛化,确保生成样本保持邻近性与多样性。

原文摘要 · Abstract (English)

Real-world data is often assumed to lie within a low-dimensional structure embedded in high-dimensional space. In practical settings, we observe only a finite set of samples, forming what we refer to as the sample data subspace. It serves an essential approximation supporting tasks such as dimensionality reduction and generation. A major challenge lies in whether generative models can reliably synthesize samples that stay within this subspace rather than drifting away from the underlying structure. In this work, we provide theoretical insights into this challenge by leveraging Flow Matching models, which transform a simple prior into a complex target distribution via a learned velocity field. By treating the real data distribution as discrete, we derive analytical expressions for the optimal velocity field under a Gaussian prior, showing that generated samples memorize real data points and represent the sample data subspace exactly. To generalize to suboptimal scenarios, we introduce the Orthogonal Subspace Decomposition Network (OSDNet), which systematically decomposes the velocity field into subspace and off-subspace components. Our analysis shows that the off-subspace component decays, while the subspace component generalizes within the sample data subspace, ensuring generated samples preserve both proximity and diversity.

生成模型流匹配子空间泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。