让图像编辑理解用户隐含意图,实现复杂语义的智能生成
R-Genie: Reasoning-Guided Generative Image Editing
- 结合扩散模型与多模态大模型,通过推理注意力机制连接语言理解与视觉合成
- 在1000+图像-指令-编辑三元组数据集上验证,可精准响应抽象意图和上下文关联
- 适合需要理解深层语义、跨场景推理的智能图像创作人群
尽管近期图像编辑技术实现了出色的视觉合成能力,现有方法仍受限于显式文本指令和有限的编辑操作,缺乏对隐含用户意图和上下文推理的深入理解。本文提出一种新的图像编辑范式:推理引导的生成式编辑,能够基于包含世界知识与意图推断的复杂多维度文本查询生成图像。为此,我们构建了一个包含超过1,000个图像-指令-编辑三元组的综合性数据集,融合丰富的推理上下文与真实世界知识。随后提出R-Genie:一种推理引导的生成式图像编辑器,将扩散模型的生成能力与多模态大语言模型的先进推理能力相结合。R-Genie引入推理注意力机制,实现语言理解与视觉合成之间的有效衔接,可处理涉及抽象用户意图及上下文推理关系的复杂编辑请求。大量实验结果表明,R-Genie能赋予扩散模型基于推理的高级编辑能力,开拓了智能图像合成的新潜力。
原文摘要 · Abstract (English)
While recent advances in image editing have enabled impressive visual synthesis capabilities, current methods remain constrained by explicit textual instructions and limited editing operations, lacking deep comprehension of implicit user intentions and contextual reasoning. In this work, we introduce a new image editing paradigm: reasoning-guided generative editing, which synthesizes images based on complex, multi-faceted textual queries accepting world knowledge and intention inference. To facilitate this task, we first construct a comprehensive dataset featuring over 1,000 image-instruction-edit triples that incorporate rich reasoning contexts and real-world knowledge. We then propose R-Genie: a reasoning-guided generative image editor, which synergizes the generation power of diffusion models with advanced reasoning capabilities of multimodal large language models. R-Genie incorporates a reasoning-attention mechanism to bridge linguistic understanding with visual synthesis, enabling it to handle intricate editing requests involving abstract user intentions and contextual reasoning relations. Extensive experimental results validate that R-Genie can equip diffusion models with advanced reasoning-based editing capabilities, unlocking new potentials for intelligent image synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。