用平衡态动力学建模生成,实现更高效采样与更强性能。
Equilibrium Matching: Generative Modeling with Implicit Energy-Based Models
- 基于隐式能量景观的平衡态动力学建模,无需时间条件。
- 在ImageNet 256×256上达成FID 1.90,优于扩散/流模型。
- 支持去噪、异常检测等任务,适合需要可调计算的场景。
我们提出均衡匹配(EqM),一种从平衡态动力学视角构建的生成建模框架。与传统扩散和流模型依赖非平衡、时间条件的动力学不同,EqM学习隐式能量景观的平衡梯度。推理时采用基于优化的采样过程,通过梯度下降在学习到的景观上迭代求解,可调节步长、自适应优化器和计算量。实验显示,EqM在ImageNet 256×256上取得1.90的FID,超越现有扩散与流模型。理论证明其能学习并采样数据流形。此外,该框架灵活,自然支持部分噪声图像去噪、分布外检测与图像合成。通过用统一的平衡景观替代时间条件速度,EqM加强了流模型与能量基模型的联系,并为优化驱动推理提供了简洁路径。
原文摘要 · Abstract (English)
We introduce Equilibrium Matching (EqM), a generative modeling framework built from an equilibrium dynamics perspective. EqM discards the non-equilibrium, time-conditional dynamics in traditional diffusion and flow-based generative models and instead learns the equilibrium gradient of an implicit energy landscape. Through this approach, we can adopt an optimization-based sampling process at inference time, where samples are obtained by gradient descent on the learned landscape with adjustable step sizes, adaptive optimizers, and adaptive compute. EqM surpasses the generation performance of diffusion/flow models empirically, achieving an FID of 1.90 on ImageNet 256$\times$256. EqM is also theoretically justified to learn and sample from the data manifold. Beyond generation, EqM is a flexible framework that naturally handles tasks including partially noised image denoising, OOD detection, and image composition. By replacing time-conditional velocities with a unified equilibrium landscape, EqM offers a tighter bridge between flow and energy-based models and a simple route to optimization-driven inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。