统一多粒度视觉生成,让大模型一次搞定图文创作与精细编辑
PUMA: Empowering Unified MLLM with Multi-granular Visual Generation
- 将多粒度视觉特征作为输入输出,统一图像生成与编辑任务
- 在多种视觉任务上表现优异,支持从文本生图到精准编辑
- 适合需要灵活生成与控制的AI绘画、设计类应用
多模态大模型在视觉-语言理解方面已取得显著进展。尽管已有研究探索了多模态大语言模型(MLLM)在视觉内容生成中的潜力,但现有方法未能充分应对统一框架下不同图像生成任务对粒度需求的差异——从文本到图像生成所需的多样性,到图像编辑所需的精确控制。本文提出PUMA,通过将多粒度视觉特征作为多模态大语言模型(MLLM)的输入和输出,优雅地解决统一框架中各类视觉任务的粒度适配问题。经过多模态预训练与任务特定指令微调,PUMA在广泛多模态任务中表现出色。该工作标志着迈向真正统一的、能适应不同视觉任务粒度需求的多模态大模型的重要一步。代码与模型将发布于 https://github.com/rongyaofang/PUMA。
原文摘要 · Abstract (English)
Recent advancements in multimodal foundation models have yielded significant progress in vision-language understanding. Initial attempts have also explored the potential of multimodal large language models (MLLMs) for visual content generation. However, existing works have insufficiently addressed the varying granularity demands of different image generation tasks within a unified MLLM paradigm - from the diversity required in text-to-image generation to the precise controllability needed in image manipulation. In this work, we propose PUMA, emPowering Unified MLLM with Multi-grAnular visual generation. PUMA unifies multi-granular visual features as both inputs and outputs of MLLMs, elegantly addressing the different granularity requirements of various image generation tasks within a unified MLLM framework. Following multimodal pretraining and task-specific instruction tuning, PUMA demonstrates proficiency in a wide range of multimodal tasks. This work represents a significant step towards a truly unified MLLM capable of adapting to the granularity demands of various visual tasks. The code and model will be released in https://github.com/rongyaofang/PUMA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。