无需遮罩的图像编辑,能精准理解复杂指令并保持画面一致性。
SmartFreeEdit: Mask-Free Spatial-Aware Image Editing with Complex Instruction Understanding
- 用语言指令直接生成编辑区域,结合空间感知标记增强场景理解。
- 在Reason-Edit数据集上,编辑准确率与指令遵循度均优于现有方法。
- 适合需要高精度、无遮罩图像修改的研究者和设计师使用。
近期图像编辑进展借助大规模多模态模型实现了直观、自然语言驱动的交互。然而,传统方法仍面临空间推理、精确区域分割及复杂场景中语义一致性等挑战。为此,我们提出SmartFreeEdit,一种端到端框架,融合多模态大语言模型(MLLM)与超图增强的修复架构,实现仅依赖自然语言指令的精准、无遮罩图像编辑。其关键创新包括:(1) 引入区域感知标记与遮罩嵌入机制,提升复杂场景的空间理解;(2) 设计基于自然语言指令优化的推理分割流程,生成更准确的编辑掩码;(3) 采用超图增强的修复模块,在复杂编辑中保障结构完整性和语义连贯性,克服局部生成的局限。在Reason-Edit基准上的大量实验表明,SmartFreeEdit在分割准确率、指令遵循度和视觉质量保留等多个指标上超越当前最优方法,解决了局部信息聚焦问题,显著提升编辑图像的全局一致性。
原文摘要 · Abstract (English)
Recent advancements in image editing have utilized large-scale multimodal models to enable intuitive, natural instruction-driven interactions. However, conventional methods still face significant challenges, particularly in spatial reasoning, precise region segmentation, and maintaining semantic consistency, especially in complex scenes. To overcome these challenges, we introduce SmartFreeEdit, a novel end-to-end framework that integrates a multimodal large language model (MLLM) with a hypergraph-enhanced inpainting architecture, enabling precise, mask-free image editing guided exclusively by natural language instructions. The key innovations of SmartFreeEdit include:(1)the introduction of region aware tokens and a mask embedding paradigm that enhance the spatial understanding of complex scenes;(2) a reasoning segmentation pipeline designed to optimize the generation of editing masks based on natural language instructions;and (3) a hypergraph-augmented inpainting module that ensures the preservation of both structural integrity and semantic coherence during complex edits, overcoming the limitations of local-based image generation. Extensive experiments on the Reason-Edit benchmark demonstrate that SmartFreeEdit surpasses current state-of-the-art methods across multiple evaluation metrics, including segmentation accuracy, instruction adherence, and visual quality preservation, while addressing the issue of local information focus and improving global consistency in the edited image. Our project will be available at https://github.com/smileformylove/SmartFreeEdit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。