提出新型生成框架Transition Matching,统一扩散与自回归模型,提升图像生成质量与效率。
Transition Matching: Scalable and Flexible Generative Modeling
- 将生成过程分解为马尔可夫转移,支持灵活的非确定性概率核和任意监督方式。
- DTM在图像质量和文本对齐上达到当前最优,采样效率显著提升。
- FHTM是首个在连续域上超越流模型的完全因果生成模型,适合与现有文本生成技术融合。
扩散模型和流匹配模型在媒体生成中取得显著进展,但其设计空间已趋成熟,制约进一步突破。与此同时,生成连续标记的自回归(AR)模型正成为统一文本与媒体生成的有前景方向。本文提出一种新的离散时间、连续状态生成范式——过渡匹配(Transition Matching, TM),统一并推进了扩散/流模型与连续自回归生成方法。TM将复杂生成任务分解为更简单的马尔可夫转移,允许表达性强的非确定性概率转移核及任意非连续监督过程,从而开辟新的灵活设计路径。我们通过三种TM变体进行探索:(i) 差分过渡匹配(DTM)通过直接学习转移概率,将流匹配推广至离散时间,实现了当前最优的图像质量、文本一致性与采样效率;(ii) 自回归过渡匹配(ARTM)和 (iii) 全历史过渡匹配(FHTM)分别为部分因果与全因果模型,推广了连续自回归方法。它们在连续因果生成方面达到与非因果方法相当的性能,并可能实现与现有自回归文本生成技术的无缝集成。值得注意的是,FHTM是首个在连续域上于文生图任务中达到或超越流模型性能的完全因果模型。我们通过大规模对比实验验证了这些贡献,所有实验保持固定架构、训练数据与超参数。
原文摘要 · Abstract (English)
Diffusion and flow matching models have significantly advanced media generation, yet their design space is well-explored, somewhat limiting further improvements. Concurrently, autoregressive (AR) models, particularly those generating continuous tokens, have emerged as a promising direction for unifying text and media generation. This paper introduces Transition Matching (TM), a novel discrete-time, continuous-state generative paradigm that unifies and advances both diffusion/flow models and continuous AR generation. TM decomposes complex generation tasks into simpler Markov transitions, allowing for expressive non-deterministic probability transition kernels and arbitrary non-continuous supervision processes, thereby unlocking new flexible design avenues. We explore these choices through three TM variants: (i) Difference Transition Matching (DTM), which generalizes flow matching to discrete-time by directly learning transition probabilities, yielding state-of-the-art image quality and text adherence as well as improved sampling efficiency. (ii) Autoregressive Transition Matching (ARTM) and (iii) Full History Transition Matching (FHTM) are partially and fully causal models, respectively, that generalize continuous AR methods. They achieve continuous causal AR generation quality comparable to non-causal approaches and potentially enable seamless integration with existing AR text generation techniques. Notably, FHTM is the first fully causal model to match or surpass the performance of flow-based methods on text-to-image task in continuous domains. We demonstrate these contributions through a rigorous large-scale comparison of TM variants and relevant baselines, maintaining a fixed architecture, training data, and hyperparameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。