让图像编辑更懂上下文,自动过滤不合理指令。
CAMILA: Context-Aware Masking for Image Editing with Language Alignment
- 根据指令与图像的上下文一致性,智能判断是否执行修改
- 在复杂或矛盾指令下仍保持图像真实性和语义一致
- 适合需要精准控制图像生成的AI创作和设计场景
文本引导的图像编辑允许用户通过自然语言指令变换和合成图像,具有高度灵活性。然而,现有模型往往盲目遵循所有用户指令,即使指令本身不可行或相互矛盾,常导致不合理输出。为此,我们提出一种名为CAMILA(上下文感知掩码的图像编辑与语言对齐)的方法,旨在验证指令与图像之间的上下文一致性,仅对相关区域执行有效编辑,忽略无法实现的指令。为全面评估该方法,我们构建了包含不可行请求的单指令和多指令图像编辑数据集。实验表明,相比现有最优模型,CAMILA在性能和语义对齐度上均表现更优,有效应对复杂指令挑战的同时保持图像完整性。
原文摘要 · Abstract (English)
Text-guided image editing has been allowing users to transform and synthesize images through natural language instructions, offering considerable flexibility. However, most existing image editing models naively attempt to follow all user instructions, even if those instructions are inherently infeasible or contradictory, often resulting in nonsensical output. To address these challenges, we propose a context-aware method for image editing named as CAMILA (Context-Aware Masking for Image Editing with Language Alignment). CAMILA is designed to validate the contextual coherence between instructions and the image, ensuring that only relevant edits are applied to the designated regions while ignoring non-executable instructions. For comprehensive evaluation of this new method, we constructed datasets for both single- and multi-instruction image editing, incorporating the presence of infeasible requests. Our method achieves better performance and higher semantic alignment than state-of-the-art models, demonstrating its effectiveness in handling complex instruction challenges while preserving image integrity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。