用离散扩散模型实现多视角图像一致生成,效果超越连续方法。
ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models
- 将多视角生成转为离散序列建模,用视觉令牌逐步解码。
- 在3D-FUTURE上比连续扩散模型高10.6%的交并比,综合指标第一。
- 无需3D先验,随机掩码+自注意力自动保证视角一致性,适合多模态任务。
受离散扩散在图文建模中成功启发,我们探索其在多视角生成中的潜力,该任务长期由连续方法主导。提出ViewMask-1-to-3,将多视角生成建模为离散序列问题,每个视点以MAGVIT-v2生成的视觉令牌表示。通过掩码令牌预测的离散扩散,实现迭代式多视角生成,统一语言与视觉于共享令牌空间。重要的是,仅用简单随机掩码结合自注意力即可自然促进跨视角一致性,无需专用架构或3D几何先验。在GSO和3D-FUTURE基准上优于基线,平均图像指标排名第一,在3D-FUTURE上交并比(IoU)比连续扩散模型高10.6%。此外,该框架可自然扩展至文本到图像生成与多模态理解,展现向统一多模态生成与理解范式的潜力。
原文摘要 · Abstract (English)
Motivated by discrete diffusion's success in language-vision modeling, we explore its potential for multi-view generation, a task dominated by continuous approaches. We introduce ViewMask-1-to-3, formulating multi-view generation as a discrete sequence modeling problem where each viewpoint is represented as visual tokens from MAGVIT-v2. Through discrete diffusion via masked token prediction, our approach enables progressive multi-view generation via iterative token unmasking, unifying language and vision in a shared token space. Importantly, simple random masking combined with self-attention naturally encourages cross-view consistency without specialized architectures or 3D geometric priors. Our method outperforms the baseline on the GSO and 3D-FUTURE benchmarks, ranking first on average across standard image metrics, and achieving a 10.6% higher IoU than continuous diffusion models on 3D-FUTURE. Furthermore, the proposed framework can be naturally extended to support text-to-image generation and multimodal understanding, highlighting its potential toward a more unified paradigm for multimodal understanding and generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。