用自然语言指令实现多种图像操作的通用视觉助手。
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
- 统一框架支持文本到图像、编辑、修复等多样任务。
- 可处理任意分辨率输入,生成效果贴近人类感知。
- 适合需要灵活图像处理的开发者与设计师使用。
本文提出通用图像到图像视觉助手PixWizard,基于自由语言指令实现图像生成、操作与转换。将多种视觉任务统一为图像-文本-图像生成框架,并构建了涵盖文本生成、图像修复、图像定位、密集预测、图像编辑、可控生成、修补/外补等多样任务的全景像素级指令微调数据集。采用扩散变换器(DiT)作为基础模型,引入灵活的任意分辨率机制,可动态适应输入图像的宽高比,更贴近人类感知。模型还集成结构感知与语义感知引导,有效融合输入图像信息。实验表明,PixWizard在多种分辨率下具备出色的生成与理解能力,并对未见任务和人类指令表现出良好泛化性。代码与资源已开源。
原文摘要 · Abstract (English)
This paper presents a versatile image-to-image visual assistant, PixWizard, designed for image generation, manipulation, and translation based on free-from language instructions. To this end, we tackle a variety of vision tasks into a unified image-text-to-image generation framework and curate an Omni Pixel-to-Pixel Instruction-Tuning Dataset. By constructing detailed instruction templates in natural language, we comprehensively include a large set of diverse vision tasks such as text-to-image generation, image restoration, image grounding, dense image prediction, image editing, controllable generation, inpainting/outpainting, and more. Furthermore, we adopt Diffusion Transformers (DiT) as our foundation model and extend its capabilities with a flexible any resolution mechanism, enabling the model to dynamically process images based on the aspect ratio of the input, closely aligning with human perceptual processes. The model also incorporates structure-aware and semantic-aware guidance to facilitate effective fusion of information from the input image. Our experiments demonstrate that PixWizard not only shows impressive generative and understanding abilities for images with diverse resolutions but also exhibits promising generalization capabilities with unseen tasks and human instructions. The code and related resources are available at https://github.com/AFeng-x/PixWizard
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。