用多模态大模型实现复杂指令图像编辑,提升准确率与背景一致性。
MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance
- 通过空间感知注意力模块对齐指令与图像区域,增强指令遵循能力。
- 在基准测试中指令遵循率提升23.96%,优于现有方法。
- 适合需要精准复杂编辑的视觉生成、设计自动化场景。
基于指令的图像编辑近期取得显著进展,但现有方法仍局限于简单操作,难以满足需复杂组合指令的实际应用需求。本文从架构设计、数据构建和评估协议三方面提出改进。针对当前模型存在指令遵循度低与背景不一致的问题,提出MCIE-E1方法,集成空间感知交叉注意力模块与背景一致性交叉注意力模块:前者在去噪过程中利用空间引导显式对齐语义指令与图像区域,提升指令遵循能力;后者保留未编辑区域特征,维持背景一致性。为支持有效训练,构建专用数据流水线,结合强大多模态大模型的细粒度自动筛选与严格人工验证以缓解复杂指令图像编辑数据稀缺问题。最后,提出CIE-Bench新基准,包含两项新评估指标。在该基准上的实验表明,MCIE-E1在定量与定性评估中均持续超越先前最优方法,指令遵循率提升23.96%。
原文摘要 · Abstract (English)
Recent advances in instruction-based image editing have shown remarkable progress. However, existing methods remain limited to relatively simple editing operations, hindering real-world applications that require complex and compositional instructions. In this work, we address these limitations from the perspectives of architectural design, data, and evaluation protocols. Specifically, we identify two key challenges in current models: insufficient instruction compliance and background inconsistency. To this end, we propose MCIE-E1, a Multimodal Large Language Model-Driven Complex Instruction Image Editing method that integrates two key modules: a spatial-aware cross-attention module and a background-consistent cross-attention module. The former enhances instruction-following capability by explicitly aligning semantic instructions with spatial regions through spatial guidance during the denoising process, while the latter preserves features in unedited regions to maintain background consistency. To enable effective training, we construct a dedicated data pipeline to mitigate the scarcity of complex instruction-based image editing datasets, combining fine-grained automatic filtering via a powerful MLLM with rigorous human validation. Finally, to comprehensively evaluate complex instruction-based image editing, we introduce CIE-Bench, a new benchmark with two new evaluation metrics. Experimental results on CIE-Bench demonstrate that MCIE-E1 consistently outperforms previous state-of-the-art methods in both quantitative and qualitative assessments, achieving a 23.96% improvement in instruction compliance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。