单次操作实现多指令图像编辑,精准控制每处修改。
PromptArtisan: Multi-instruction Image Editing in Single Pass with Complete Attention Control
- 用新注意力机制一次性完成多个编辑指令。
- 支持掩码重叠,可精确控制修改区域。
- 无需训练,适合初学者与专业人士使用。
我们提出 PromptArtisan,一种突破性的单次通过多指令图像编辑方法,显著提升效率并消除传统迭代修正的耗时问题。用户可为图像中不同区域提供多个编辑指令,每个指令对应一个特定掩码,支持掩码交集或重叠,实现复杂精细的图像变换。该方法结合预训练 InstructPix2Pix 模型与新颖的完整注意力控制机制(CACM),确保对用户指令的精准遵循,实现细粒度编辑控制。此外,本方法为零样本,无需额外训练,且处理复杂度优于传统迭代方法。通过无缝集成多指令能力、单次执行效率与完整注意力控制,PromptArtisan 开启了创意高效图像编辑的新可能,适用于各类用户。
原文摘要 · Abstract (English)
We present PromptArtisan, a groundbreaking approach to multi-instruction image editing that achieves remarkable results in a single pass, eliminating the need for time-consuming iterative refinement. Our method empowers users to provide multiple editing instructions, each associated with a specific mask within the image. This flexibility allows for complex edits involving mask intersections or overlaps, enabling the realization of intricate and nuanced image transformations. PromptArtisan leverages a pre-trained InstructPix2Pix model in conjunction with a novel Complete Attention Control Mechanism (CACM). This mechanism ensures precise adherence to user instructions, granting fine-grained control over the editing process. Furthermore, our approach is zero-shot, requiring no additional training, and boasts improved processing complexity compared to traditional iterative methods. By seamlessly integrating multi-instruction capabilities, single-pass efficiency, and complete attention control, PromptArtisan unlocks new possibilities for creative and efficient image editing workflows, catering to both novice and expert users alike.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。