Lavida-O统一实现图文理解与高分辨率生成,支持对象定位和编辑。
Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
- 采用弹性混合变换器架构,分轻量生成与大型理解分支。
- 1024px高清文生图、对象定位与图像编辑均达顶尖水平。
- 融合规划与自省机制,生成质量显著提升,推理更快。
我们提出 Lavida-O,一种统一的多模态掩码扩散模型(MDM),用于多模态理解与生成。与现有模型如 MMaDa 和 Muddit 仅支持图像级理解任务和低分辨率图像生成不同,Lavida-O 在单一框架内实现了图像级理解、对象定位、图像编辑及 1024px 高分辨率文本到图像合成。其创新性地引入弹性混合变压器(Elastic-MoT)架构,将轻量生成分支与大型理解分支结合,通过标记压缩、通用文本条件与分层采样实现高效高质量生成。此外,模型在生成与编辑任务中引入规划与迭代自省机制,无缝利用理解能力提升生成质量。Lavida-O 在 RefCOCO 对象定位、GenEval 文本到图像生成、ImgEdit 图像编辑等多个基准上达到当前最佳性能,优于 Qwen2.5-VL 和 FluxKontext-dev 等自回归与连续扩散模型,同时实现显著推理加速。这些进展确立了 Lavida-O 作为可扩展多模态推理与生成的新范式。
原文摘要 · Abstract (English)
We propose Lavida-O, a unified Masked Diffusion Model (MDM) for multimodal understanding and generation. Unlike existing multimodal MDMs such as MMaDa and Muddit which only support simple image-level understanding tasks and low-resolution image generation, Lavida-O presents a single framework that enables image-level understanding, object grounding, image editing, and high-resolution (1024px) text-to-image synthesis. Lavida-O incorporates a novel Elastic Mixture-of-Transformers (Elastic-MoT) architecture that couples a lightweight generation branch with a larger understanding branch, supported by token compression, universal text conditioning and stratified sampling for efficient and high-quality generation. Lavida-O further incorporates planning and iterative self-reflection in image generation and editing tasks, seamlessly boosting generation quality with its understanding capabilities. Lavida-O achieves state-of-the-art performance on a wide range of benchmarks including RefCOCO object grounding, GenEval text-to-image generation, and ImgEdit image editing, outperforming existing autoregressive models and continuous diffusion models such as Qwen2.5-VL and FluxKontext-dev, while offering considerable speedup at inference. These advances establish Lavida-O as a new paradigm for scalable multimodal reasoning and generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。