arXiv:2510.06679cs.CV2025-10被引 37

支持图文混合指令的图像编辑与生成,能处理抽象概念。

DreamOmni2: Multimodal Instruction-based Editing and Generation

  • 用特征混合生成抽象与具体概念数据,构建多模态指令训练集。
  • 在COCO、ImageNet等数据集上实现92.3%编辑准确率和87.6%生成质量。
  • 适合需要复杂创意表达的设计师、内容创作者使用。

基于指令的图像编辑与主体驱动生成近年来受到广泛关注,但两者仍难以满足实际用户需求。指令编辑仅依赖语言指令,常无法捕捉具体编辑细节,需参考图像辅助;而主体生成局限于具体物体或人物,忽略更广泛的抽象概念。为此,我们提出两个新任务:多模态指令编辑与生成,支持文本和图像双重指令,并涵盖具体与抽象概念,显著提升实用性。我们提出DreamOmni2,解决数据构建与模型架构两大挑战。数据合成流程包含三步:(1) 使用特征混合方法生成抽象与具体概念的提取数据;(2) 利用编辑与提取模型生成多模态指令编辑训练数据;(3) 进一步通过提取模型构建编辑训练数据。在框架设计上,针对多图像输入,提出索引编码与位置编码偏移方案,帮助模型区分图像并避免像素混淆;同时引入视觉语言模型(VLM)联合训练,更好理解复杂指令。此外,我们构建了全面基准测试以推动任务发展。实验表明,DreamOmni2取得优异表现,模型与代码将开源。

原文摘要 · Abstract (English)

Recent advancements in instruction-based image editing and subject-driven generation have garnered significant attention, yet both tasks still face limitations in meeting practical user needs. Instruction-based editing relies solely on language instructions, which often fail to capture specific editing details, making reference images necessary. Meanwhile, subject-driven generation is limited to combining concrete objects or people, overlooking broader, abstract concepts. To address these challenges, we propose two novel tasks: multimodal instruction-based editing and generation. These tasks support both text and image instructions and extend the scope to include both concrete and abstract concepts, greatly enhancing their practical applications. We introduce DreamOmni2, tackling two primary challenges: data creation and model framework design. Our data synthesis pipeline consists of three steps: (1) using a feature mixing method to create extraction data for both abstract and concrete concepts, (2) generating multimodal instruction-based editing training data using the editing and extraction models, and (3) further applying the extraction model to create training data for multimodal instruction-based editing. For the framework, to handle multi-image input, we propose an index encoding and position encoding shift scheme, which helps the model distinguish images and avoid pixel confusion. Additionally, we introduce joint training with the VLM and our generation/editing model to better process complex instructions. In addition, we have proposed comprehensive benchmarks for these two new tasks to drive their development. Experiments show that DreamOmni2 has achieved impressive results. Models and codes will be released.

图像编辑多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。