arXiv:2505.20147cs.CV2025-05NeurIPS被引 43

用离散流模型替代自回归架构,实现图像生成与理解统一

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities

  • 基于离散流匹配构建无自回归的统一多模态模型
  • 在视觉理解和图像生成上达到主流自回归模型水平
  • 支持生成过程自我修正,适合追求高效生成的开发者

大型语言模型的快速发展推动了多模态大模型(MLLM)的兴起,这类模型将视觉理解与图像生成统一于单一框架中。然而,现有多数MLLM依赖自回归(AR)架构,存在图像生成受栅格扫描顺序限制、因果上下文建模能力受限等固有缺陷。本文提出FUDOKI,一种完全基于离散流匹配的统一多模态模型,挑战传统AR范式。通过引入度量诱导的概率路径与动能最优速度,该框架超越了传统的掩码式破坏过程,支持迭代精炼与自我修正,并实现更丰富的双向上下文融合。为降低从头训练成本,FUDOKI以预训练的AR型MLLM为起点,渐进式过渡至离散流匹配范式。实验表明,FUDOKI在视觉理解与图像生成任务上表现与最先进AR模型相当,展现出下一代统一多模态模型的潜力。此外,测试时缩放技术可显著提升其性能,进一步证明其可通过强化学习持续优化。

原文摘要 · Abstract (English)

The rapid progress of large language models (LLMs) has catalyzed the emergence of multimodal large language models (MLLMs) that unify visual understanding and image generation within a single framework. However, most existing MLLMs rely on autoregressive (AR) architectures, which impose inherent limitations on future development, such as the raster-scan order in image generation and restricted reasoning abilities in causal context modeling. In this work, we challenge the dominance of AR-based approaches by introducing FUDOKI, a unified multimodal model purely based on discrete flow matching, as an alternative to conventional AR paradigms. By leveraging metric-induced probability paths with kinetic optimal velocities, our framework goes beyond the previous masking-based corruption process, enabling iterative refinement with self-correction capability and richer bidirectional context integration during generation. To mitigate the high cost of training from scratch, we initialize FUDOKI from pre-trained AR-based MLLMs and adaptively transition to the discrete flow matching paradigm. Experimental results show that FUDOKI achieves performance comparable to state-of-the-art AR-based MLLMs across both visual understanding and image generation tasks, highlighting its potential as a foundation for next-generation unified multimodal models. Furthermore, we show that applying test-time scaling techniques to FUDOKI yields significant performance gains, further underscoring its promise for future enhancement through reinforcement learning.

多模态模型离散流生成统一

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。