用分层视觉提示实现精准图像编辑,让生成更可控。
MagicQuillV2: Precise and Interactive Image Editing with Layered Visual Cues
- 将创作意图拆分为内容、位置、结构、颜色四层控制
- 支持精确局部编辑,包括物体移除与位置调整
- 适合需要精细调控的设计师或内容创作者
我们提出 MagicQuill V2,一种基于分层视觉提示的生成式图像编辑系统,弥合了扩散模型的语义能力与传统图形软件的精细控制之间的差距。扩散变换器虽擅长整体生成,但单一整体提示难以分离用户对内容、位置、外观的不同意图。为此,我们的方法将创作意图分解为可操控的四层:内容层(创什么)、空间层(放哪里)、结构层(怎么形)和颜色层(调什么色)。技术上包含上下文感知的内容生成流水线、统一控制模块处理所有提示,以及微调的空间分支实现精确局部编辑(含物体移除)。大量实验表明,该分层方法有效解决了用户意图分离问题,赋予创作者对生成过程的直接、直观控制。
原文摘要 · Abstract (English)
We propose MagicQuill V2, a novel system that introduces a \textbf{layered composition} paradigm to generative image editing, bridging the gap between the semantic power of diffusion models and the granular control of traditional graphics software. While diffusion transformers excel at holistic generation, their use of singular, monolithic prompts fails to disentangle distinct user intentions for content, position, and appearance. To overcome this, our method deconstructs creative intent into a stack of controllable visual cues: a content layer for what to create, a spatial layer for where to place it, a structural layer for how it is shaped, and a color layer for its palette. Our technical contributions include a specialized data generation pipeline for context-aware content integration, a unified control module to process all visual cues, and a fine-tuned spatial branch for precise local editing, including object removal. Extensive experiments validate that this layered approach effectively resolves the user intention gap, granting creators direct, intuitive control over the generative process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。