统一图像生成与编辑任务,用多模态指令提升效果和泛化能力
MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and Editing
- 用统一的输入输出格式处理生成与编辑,共享视觉语义表示
- 在新任务上达到当前最佳,指令遵循性和图像一致性显著提升
- 适合需要跨任务泛化的图像生成与编辑研究者使用
尽管基于扩散模型的图像生成取得显著进展,但主体驱动生成与基于指令的编辑仍具挑战。现有方法通常分开处理,受限于高质量数据不足和泛化能力差。两者均需捕捉复杂视觉变化并保持输入输出一致。受此启发,我们提出MIGE,一个统一框架,通过多模态指令标准化任务表示。将主体生成视为空白画布创作,指令编辑视为已有图像修改,建立共享输入输出范式;引入新型多模态编码器,将自由形式的多模态指令映射至统一视觉语言空间,通过特征融合机制整合视觉与语义信息。该统一性支持两任务联合训练,带来两大优势:(1) 跨任务增强:共享表示提升指令遵循性与视觉一致性;(2) 泛化能力:统一学习格式促进跨任务知识迁移,使MIGE能泛化到新组合任务,包括基于指令的主体驱动编辑。实验表明,MIGE在主体生成与指令编辑中表现优异,并在新任务上达到当前最优。代码与模型已公开于https://github.com/Eureka-Maggie/MIGE。
原文摘要 · Abstract (English)
Despite significant progress in diffusion-based image generation, subject-driven generation and instruction-based editing remain challenging. Existing methods typically treat them separately, struggling with limited high-quality data and poor generalization. However, both tasks require capturing complex visual variations while maintaining consistency between inputs and outputs. Inspired by this, we propose MIGE, a unified framework that standardizes task representations using multimodal instructions. It first treats subject-driven generation as creation on a blank canvas and instruction-based editing as modification of an existing image, establishing a shared input-output formulation, then introduces a novel multimodal encoder that maps free-form multimodal instructions into a unified vision-language space, integrating visual and semantic features through a feature fusion mechanism. This unification enables joint training of both tasks, providing two key advantages: (1) Cross-Task Enhancement: by leveraging shared visual and semantic representations, joint training improves instruction adherence and visual consistency in both subject-driven generation and instruction-based editing. (2) Generalization: learning in a unified format facilitates cross-task knowledge transfer, enabling MIGE to generalize to novel compositional tasks, including instruction-based subject-driven editing. Experiments show that MIGE excels in both subject-driven generation and instruction-based editing while setting a SOTA in the new task of instruction-based subject-driven editing. Code and model have been publicly available at https://github.com/Eureka-Maggie/MIGE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。