针对少样本生成中模型记忆数据的问题,提出几何约束的迭代优化框架提升多样性。
Latent Iterative Refinement Flow: A Geometric Constrained Approach for Few-Shot Generation
- 通过生成-修正-增强闭环,在语义对齐潜空间中逐步稠密化数据流形。
- 在FFHQ子集和低样本数据集上,生成多样性与召回率显著优于现有方法。
- 适用于少样本图像生成场景,尤其适合追求高多样性的应用。
在有限数据下训练的扩散模型和流匹配模型常出现过拟合记忆现象,导致生成多样性严重下降。本文从动态视角揭示这一‘坍缩至记忆’现象源于速度场坍缩,即学习到的速度场退化为孤立点吸引子,使采样轨迹陷入其中。受此启发,我们提出一种几何感知的少样本生成框架——潜在迭代精炼流(LIRF),通过利用语义对齐潜空间的内在几何结构,构建‘生成-修正-增强’闭环,逐步稠密化训练数据流形,有效缓解速度场坍缩问题。理论证明了该流形稠密化过程的收敛性。在FFHQ子集和低样本数据集上的实验表明,相较于现有扩散模型,LIRF在保持良好生成质量的同时,显著提升了生成多样性与召回率。
原文摘要 · Abstract (English)
Diffusion and flow-matching models trained with limited data often tend to memorize the training data instead of generalization, leading to severely reduced diversity. In this paper, we provide a dynamical perspective and identify this ``collapse-to-memorization'' phenomenon as a consequence of the \emph{velocity field collapse}, where the learned field degenerates into isolated point attractors and trap the sampling trajectories. Inspired by this novel view, we introduce \textbf{{\BLUE L}atent {\BLUE I}terative {\BLUE R}efinement {\BLUE F}low ({\BLUE LIRF})}, a geometry-aware framework for from-scratch training of diffusion models in the limited-data regime. By exploiting the intrinsic geometry of a semantically aligned latent space, LIRF progressively densifies the training data manifold via a \emph{generation--correction--augmentation} closed loop, thereby effectively resolving the velocity field collapse. Theoretical guarantee on the convergence of this manifold densification procedure is also provided. Experiments on FFHQ subsets and Low-Shot datasets demonstrate the advantageous performance of LIRF over existing diffusion models for limited-data generation, achieving significantly higher diversity and recall, with comparably good generative performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。