用大模型视觉理解能力提升图像编辑精准度,让指令更懂用户意图。
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection
- 基于大模型语义推理优化模糊指令,提升理解准确性。
- 利用大模型内在视觉能力生成嵌入,指导扩散模型编辑过程。
- 联合训练双策略,复杂场景下编辑更精准、视觉更连贯。
近年来,AI生成内容(AIGC)的发展显著加速了图像编辑技术,推动对多样化和细粒度编辑的需求。然而,现有方法在复杂场景中仍难以实现高精度与语义准确性。尽管已有研究将多模态大语言模型(MLLM)引入图像编辑流程,但当前方法主要依赖文本指令解析,未充分挖掘大模型的内在视觉理解能力,导致文本语义与视觉结果之间对齐不足。为此,我们提出MIND-Edit,一个融合预训练扩散模型与MLLM的端到端图像编辑框架。该框架引入两种互补策略:(1) 基于MLLM语义推理优化模糊用户指令的文本指令优化策略;(2) 显式利用MLLM内在视觉理解能力,通过生成的视觉嵌入推断编辑意图并引导扩散过程的洞察驱动编辑策略。此外,我们提出一种联合训练方法,使两种策略相互强化,实现更准确的指令理解与符合用户意图的视觉一致编辑。大量实验表明,MIND-Edit在定量指标与视觉质量上均优于现有先进方法,尤其在复杂挑战性场景下表现优异。
原文摘要 · Abstract (English)
Recent advances in AI-generated content (AIGC) have significantly accelerated image editing techniques, driving increasing demand for diverse and fine-grained edits. Despite these advances, existing image editing methods still face challenges in achieving high precision and semantic accuracy in complex scenarios. Recent studies address this issue by incorporating multimodal large language models (MLLMs) into image editing pipelines. However, current MLLM-based methods mainly rely on interpreting textual instructions, leaving the intrinsic visual understanding of large models largely unexplored, thus resulting in insufficient alignment between textual semantics and visual outcomes. To overcome these limitations, we propose MIND-Edit, an end-to-end image-editing framework integrating pretrained diffusion model with MLLM. MIND-Edit introduces two complementary strategies: (1) a text instruction optimization strategy that clarifies ambiguous user instructions based on semantic reasoning from the MLLM, and (2) an MLLM insight-driven editing strategy that explicitly leverages the intrinsic visual understanding capability of the MLLM to infer editing intent and guide the diffusion process via generated visual embeddings. Furthermore, we propose a joint training approach to effectively integrate both strategies, allowing them to reinforce each other for more accurate instruction interpretation and visually coherent edits aligned with user intent. Extensive experiments demonstrate that MIND-Edit outperforms state-of-the-art image editing methods in both quantitative metrics and visual quality, particularly under complex and challenging scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。