用流形视角统一理解生成模型、神经网络优化与大模型演化
Optimal and Diffusion Transports in Machine Learning
- 通过粒子流场描述概率分布演变,突破传统密度表示局限
- 对比扩散与最优传输两种插值方法,揭示其在生成与优化中的共性
- 适用于生成模型、神经网络训练及大模型动态建模的研究者
机器学习中的多个问题可自然表述为随时间演化的概率分布设计与分析,涵盖通过扩散方法采样、神经网络权重优化以及大语言模型各层标记分布的演化。尽管应用对象不同(样本、权重、标记),其数学结构具有共性。核心思想是通过描述粒子运动的向量场,将欧拉密度表示转换为拉格朗日粒子表示。这一双重视角虽带来拉格朗日向量场非唯一性的挑战,但也为构建具有良好正则性、稳定性和计算可处理性的分布演化流程提供了机会。本文综述了这些方法,重点介绍两种互补路径:基于随机插值过程的扩散方法(支撑现代生成式AI),以及通过最小化位移代价定义插值的最优传输方法。文中展示了两者在采样、神经网络优化及大语言模型中变换器动态建模等场景中的应用。
原文摘要 · Abstract (English)
Several problems in machine learning are naturally expressed as the design and analysis of time-evolving probability distributions. This includes sampling via diffusion methods, optimizing the weights of neural networks, and analyzing the evolution of token distributions across layers of large language models. While the targeted applications differ (samples, weights, tokens), their mathematical descriptions share a common structure. A key idea is to switch from the Eulerian representation of densities to their Lagrangian counterpart through vector fields that advect particles. This dual view introduces challenges, notably the non-uniqueness of Lagrangian vector fields, but also opportunities to craft density evolutions and flows with favorable properties in terms of regularity, stability, and computational tractability. This survey presents an overview of these methods, with emphasis on two complementary approaches: diffusion methods, which rely on stochastic interpolation processes and underpin modern generative AI, and optimal transport, which defines interpolation by minimizing displacement cost. We illustrate how both approaches appear in applications ranging from sampling, neural network optimization, to modeling the dynamics of transformers for large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。