让多模态大模型直接在图像潜空间规划,提升生成精度。
Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- 用多模态大模型在潜空间做视觉规划,而非仅作文本编码
- 六项任务中均优于全局条件基线,尤其在布局控制上表现突出
- 适合需要精确结构控制的图像/视频生成场景
多模态大语言模型(MLLM)在视觉理解上进步显著,但当前在图像生成中仅被用作扩散模型的全局文本编码器,其推理与规划能力未被充分利用。这导致理解能力强但生成控制精度不足的问题。本文提出轻量级框架MetaCanvas,使MLLM能直接在空间和时空潜空间中进行推理与规划,并与扩散生成器紧密协同。我们在三种不同扩散模型上实现MetaCanvas,评估了六项任务:文本到图像生成、文本/图像到视频生成、图像/视频编辑及上下文视频生成,这些任务均需精准布局、强属性绑定和复杂推理控制。结果表明,MetaCanvas持续优于全局条件基线,证明将MLLM视为潜空间规划者是弥合多模态理解与生成差距的可行方向。
原文摘要 · Abstract (English)
Multimodal learning has rapidly advanced visual understanding, largely via multimodal large language models (MLLMs) that use powerful LLMs as cognitive cores. In visual generation, however, these powerful core models are typically reduced to global text encoders for diffusion models, leaving most of their reasoning and planning ability unused. This creates a gap: current multimodal LLMs can parse complex layouts, attributes, and knowledge-intensive scenes, yet struggle to generate images or videos with equally precise and structured control. We propose MetaCanvas, a lightweight framework that lets MLLMs reason and plan directly in spatial and spatiotemporal latent spaces and interface tightly with diffusion generators. We empirically implement MetaCanvas on three different diffusion backbones and evaluate it across six tasks, including text-to-image generation, text/image-to-video generation, image/video editing, and in-context video generation, each requiring precise layouts, robust attribute binding, and reasoning-intensive control. MetaCanvas consistently outperforms global-conditioning baselines, suggesting that treating MLLMs as latent-space planners is a promising direction for narrowing the gap between multimodal understanding and generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。