arXiv:2409.20423stat.MLcs.AI2024-09ICML被引 2

用高斯过程建模流路径,降低生成样本方差。

Stream-level flow matching with Gaussian processes

  • 以流为单位定义高斯过程路径,保持无需模拟的训练特性。
  • 在图像与神经时间序列数据上,生成样本质量显著提升。
  • 适合处理具有时序关联的数据,如连续神经信号。

流匹配(FM)是一类用于训练连续归一化流(CNF)的算法。条件流匹配(CFM)利用了CNF的边际向量场可通过最小二乘回归拟合给定流路径端点的条件向量场这一性质。本文通过在‘流’(即连接数据对源与目标的潜在随机路径)上定义条件概率路径,并采用高斯过程(GP)进行建模,扩展了CFM算法。高斯过程独特的分布性质有助于保持CFM训练中‘无需模拟’的特性。我们证明,该方法在适度增加计算成本的情况下,能有效降低估计边际向量场的方差,从而在常用指标下提升生成样本质量。此外,将GP应用于流路径可灵活关联多个相关训练数据点(如时间序列)。我们在仿真及图像、神经时间序列数据应用中实证验证了该方法的有效性。

原文摘要 · Abstract (English)

Flow matching (FM) is a family of training algorithms for fitting continuous normalizing flows (CNFs). Conditional flow matching (CFM) exploits the fact that the marginal vector field of a CNF can be learned by fitting least-squares regression to the conditional vector field specified given one or both ends of the flow path. In this paper, we extend the CFM algorithm by defining conditional probability paths along ``streams'', instances of latent stochastic paths that connect data pairs of source and target, which are modeled with Gaussian process (GP) distributions. The unique distributional properties of GPs help preserve the ``simulation-free" nature of CFM training. We show that this generalization of the CFM can effectively reduce the variance in the estimated marginal vector field at a moderate computational cost, thereby improving the quality of the generated samples under common metrics. Additionally, adopting the GP on the streams allows for flexibly linking multiple correlated training data points (e.g., time series). We empirically validate our claim through both simulations and applications to image and neural time series data.

流匹配高斯过程生成模型时间序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。