arXiv:2508.07581cs.LGmath.DS2025-08NeurIPS被引 5

即使生成模型有误差,为何仍能生成数据分布内的样本?

When and how can inexact generative models still sample from the data manifold?

论文配图:When and how can inexact generative models still sample from the data manifold?
图 1 · 摘自论文原文
  • 从动力系统角度分析生成过程的微小误差影响
  • 误差仅在数据流形上改变密度,不导致样本偏离流形
  • 流形边界切空间对齐是保持鲁棒性的关键机制

某些动态生成模型存在一种奇特现象:尽管评分函数或漂移向量场存在学习误差,生成样本仍会沿数据分布支持集移动而非脱离。本文通过动力系统方法研究该‘支持集鲁棒性’现象。对概率流的微扰分析表明,在一大类生成模型中,无穷小学习误差仅导致目标密度在数据流形上产生差异。进一步揭示,生成过程的动态机制使最敏感扰动方向(主李雅普诺夫向量)与数据流形边界的切空间对齐,从而实现鲁棒性,并给出了达成此对齐的充分条件。该条件可高效计算,且在实际中自动提供数据流形切丛的准确估计。通过有限时间线性扰动分析样本路径与概率流,本工作补充并拓展了基于随机分析、统计学习与不确定性量化对生成模型理论保证的研究。结果适用于条件流匹配、基于评分的生成模型等不同动态生成模型,以及满足或不满足流形假设的目标分布。

原文摘要 · Abstract (English)

A curious phenomenon observed in some dynamical generative models is the following: despite learning errors in the score function or the drift vector field, the generated samples appear to shift \emph{along} the support of the data distribution but not \emph{away} from it. In this work, we investigate this phenomenon of \emph{robustness of the support} by taking a dynamical systems approach on the generating stochastic/deterministic process. Our perturbation analysis of the probability flow reveals that infinitesimal learning errors cause the predicted density to be different from the target density only on the data manifold for a wide class of generative models. Further, what is the dynamical mechanism that leads to the robustness of the support? We show that the alignment of the top Lyapunov vectors (most sensitive infinitesimal perturbation directions) with the tangent spaces along the boundary of the data manifold leads to robustness and prove a sufficient condition on the dynamics of the generating process to achieve this alignment. Moreover, the alignment condition is efficient to compute and, in practice, for robust generative models, automatically leads to accurate estimates of the tangent bundle of the data manifold. Using a finite-time linear perturbation analysis on samples paths as well as probability flows, our work complements and extends existing works on obtaining theoretical guarantees for generative models from a stochastic analysis, statistical learning and uncertainty quantification points of view. Our results apply across different dynamical generative models, such as conditional flow-matching and score-based generative models, and for different target distributions that may or may not satisfy the manifold hypothesis.

生成模型动力系统流形学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。