无需训练即可精准编辑图像,三者兼顾:文字匹配、背景保真、衔接自然。
CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing
- 仅对未编辑区域使用Canny结构引导,保留原图细节。
- 双提示机制:局部提示控修改,全局提示保整体协调。
- 支持粗略遮罩或点提示,适配复杂指令编辑任务。
近期文本到图像模型的发展使无需训练的区域图像编辑成为可能,依赖基础模型的生成先验。然而,现有方法难以平衡编辑区域的文字契合度、未编辑区域的上下文保真度以及编辑融合的自然性。本文提出CannyEdit,一种新型无训练框架,通过两项关键创新解决这一三重困境。首先,选择性Canny控制仅在未编辑区域应用结构引导,保留原图细节的同时,实现指定可编辑区域的精准文本驱动修改。其次,双提示引导结合局部提示(用于具体编辑)与全局提示(用于整体场景一致性)。该协同机制实现了对象添加、替换和移除的可控局部编辑,在文字契合度、上下文保真度和编辑无缝性之间达到优于当前基于区域方法的权衡。此外,CannyEdit具备优异灵活性:可在粗糙掩码甚至单点提示下完成添加任务;且能无训练地与视觉-语言模型集成,支持需规划与推理的复杂指令编辑。大量评估显示,其在复杂对象添加场景中表现超越主流指令式编辑器。
原文摘要 · Abstract (English)
Recent advances in text-to-image (T2I) models have enabled training-free regional image editing by leveraging the generative priors of foundation models. However, existing methods struggle to balance text adherence in edited regions, context fidelity in unedited areas, and seamless integration of edits. We introduce CannyEdit, a novel training-free framework that addresses this trilemma through two key innovations. First, Selective Canny Control applies structural guidance from a Canny ControlNet only to the unedited regions, preserving the original image's details while allowing for precise, text-driven changes in the specified editable area. Second, Dual-Prompt Guidance utilizes both a local prompt for the specific edit and a global prompt for overall scene coherence. Through this synergistic approach, these components enable controllable local editing for object addition, replacement, and removal, achieving a superior trade-off among text adherence, context fidelity, and editing seamlessness compared to current region-based methods. Beyond this, CannyEdit offers exceptional flexibility: it operates effectively with rough masks or even single-point hints in addition tasks. Furthermore, the framework can seamlessly integrate with vision-language models in a training-free manner for complex instruction-based editing that requires planning and reasoning. Our extensive evaluations demonstrate CannyEdit's strong performance against leading instruction-based editors in complex object addition scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。