arXiv:2601.04056cs.CL2026-01被引 3

统一文本图像生成,用连续与离散扩散结合新方法。

Bridging the Discrete-Continuous Gap: Unified Multimodal Generation via Coupled Manifold Discrete Absorbing Diffusion

  • 用连续潜空间规划语义,离散过程生成token。
  • 变量噪声调度提升生成稳定性,避免对齐难题。
  • 适合做多模态统一生成的科研与工程人员。

生成模型在离散数据(如文本)和连续数据(如图像)上分道扬镳,阻碍了真正统一的多模态系统发展。尽管掩码语言模型(MLMs)能高效捕捉双向上下文,但其生成保真度不及自回归模型,也缺乏扩散模型的语义连续性。将掩码生成扩展到多模态场景还会带来严重的对齐挑战和训练不稳定性。本文提出新型概率框架 CoM-DAD(Coupled Manifold Discrete Absorbing Diffusion),将多模态生成建模为分层双过程:首先通过连续潜在扩散过程建模语义流形;其次将标记生成视为受演化语义先验调控的离散吸收扩散过程,采用可变噪声调度机制。关键创新是引入随机混合模态传输策略,无需重型对比双编码器即可实现跨模态对齐。该方法在标准掩码建模基础上显著提升稳定性,为可扩展的统一文本-图像生成树立新范式。

原文摘要 · Abstract (English)

The bifurcation of generative modeling into autoregressive approaches for discrete data (text) and diffusion approaches for continuous data (images) hinders the development of truly unified multimodal systems. While Masked Language Models (MLMs) offer efficient bidirectional context, they traditionally lack the generative fidelity of autoregressive models and the semantic continuity of diffusion models. Furthermore, extending masked generation to multimodal settings introduces severe alignment challenges and training instability. In this work, we propose \textbf{CoM-DAD} (\textbf{Co}upled \textbf{M}anifold \textbf{D}iscrete \textbf{A}bsorbing \textbf{D}iffusion), a novel probabilistic framework that reformulates multimodal generation as a hierarchical dual-process. CoM-DAD decouples high-level semantic planning from low-level token synthesis. First, we model the semantic manifold via a continuous latent diffusion process; second, we treat token generation as a discrete absorbing diffusion process, regulated by a \textbf{Variable-Rate Noise Schedule}, conditioned on these evolving semantic priors. Crucially, we introduce a \textbf{Stochastic Mixed-Modal Transport} strategy that aligns disparate modalities without requiring heavy contrastive dual-encoders. Our method demonstrates superior stability over standard masked modeling, establishing a new paradigm for scalable, unified text-image generation.

多模态生成扩散模型统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。