arXiv:2509.04394cs.LGcs.CV2025-09被引 26

提出新型生成模型,一步到多步都能高效生成高质量图像。

Transition Models: Rethinking the Generative Learning Objective

  • 用连续时间方程建模任意时间间隔的状态转移,突破传统生成目标限制。
  • 865M参数模型在所有采样步数下超越8B/12B参数的顶尖模型,且质量随步数提升。
  • 支持最高4096x4096分辨率生成,适合需要高质高效图像生成的场景。

生成建模中存在根本性矛盾:迭代扩散模型虽保真度高但计算成本大,而高效少步方案则受限于质量天花板。根源在于训练目标仅关注极小时间动态(PF-ODE)或直接端点预测。本文提出一个精确的连续时间动力学方程,可解析定义任意有限时间间隔内的状态转移。由此构建的新一代生成范式——过渡模型(TiM),能自适应任意步数的生成过程,无缝覆盖从单步跳跃到精细优化的整个轨迹。尽管仅有865M参数,TiM在所有评估步数下均超越主流模型如SD3.5(8B参数)和FLUX.1(12B参数)。更重要的是,不同于以往少步生成器,TiM在增加采样预算时表现出单调的质量提升。此外,采用原生分辨率策略后,其在4096x4096分辨率下仍保持卓越保真度。

原文摘要 · Abstract (English)

A fundamental dilemma in generative modeling persists: iterative diffusion models achieve outstanding fidelity, but at a significant computational cost, while efficient few-step alternatives are constrained by a hard quality ceiling. This conflict between generation steps and output quality arises from restrictive training objectives that focus exclusively on either infinitesimal dynamics (PF-ODEs) or direct endpoint prediction. We address this challenge by introducing an exact, continuous-time dynamics equation that analytically defines state transitions across any finite time interval. This leads to a novel generative paradigm, Transition Models (TiM), which adapt to arbitrary-step transitions, seamlessly traversing the generative trajectory from single leaps to fine-grained refinement with more steps. Despite having only 865M parameters, TiM achieves state-of-the-art performance, surpassing leading models such as SD3.5 (8B parameters) and FLUX.1 (12B parameters) across all evaluated step counts. Importantly, unlike previous few-step generators, TiM demonstrates monotonic quality improvement as the sampling budget increases. Additionally, when employing our native-resolution strategy, TiM delivers exceptional fidelity at resolutions up to 4096x4096.

生成模型扩散模型高效生成高分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。