让扩散模型更易编辑,通过掩码训练和推理时扩展提示。
MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing
- 用双重扰动训练法增强模型对局部结构的感知能力。
- 推理时插入暂停标记可动态提升生成算力,支持复杂编辑。
- 适合需要精细图像编辑与可控生成的研究者与开发者。
尽管扩散模型在文本到图像生成中表现卓越,但在基于场景的视觉编辑和组合控制方面仍面临挑战。我们提出一种名为MADI的框架,通过两项核心创新显著提升扩散模型的可编辑性、组合性和可控性。首先引入掩码增强高斯扩散(MAgD),采用双重扰动训练策略,结合标准去噪评分匹配与掩码重建,使模型学习更具区分性和组合性的视觉表征,实现局部化、结构感知的编辑。其次提出基于暂停标记(Pause Tokens)的推理时容量扩展机制,这些特殊占位符插入提示词中,可在推理阶段动态增加计算容量。研究发现,训练时使用丰富密集的提示能进一步提升性能,尤其对MAgD有效。MADI显著增强了扩散模型的可编辑性,为构建通用、上下文感知的生成架构铺平道路。
原文摘要 · Abstract (English)
Despite the remarkable success of diffusion models in text-to-image generation, their effectiveness in grounded visual editing and compositional control remains challenging. Motivated by advances in self-supervised learning and in-context generative modeling, we propose a series of simple yet powerful design choices that significantly enhance diffusion model capacity for structured, controllable generation and editing. We introduce Masking-Augmented Diffusion with Inference-Time Scaling (MADI), a framework that improves the editability, compositionality and controllability of diffusion models through two core innovations. First, we introduce Masking-Augmented gaussian Diffusion (MAgD), a novel training strategy with dual corruption process which combines standard denoising score matching and masked reconstruction by masking noisy input from forward process. MAgD encourages the model to learn discriminative and compositional visual representations, thus enabling localized and structure-aware editing. Second, we introduce an inference-time capacity scaling mechanism based on Pause Tokens, which act as special placeholders inserted into the prompt for increasing computational capacity at inference time. Our findings show that adopting expressive and dense prompts during training further enhances performance, particularly for MAgD. Together, these contributions in MADI substantially enhance the editability of diffusion models, paving the way toward their integration into more general-purpose, in-context generative diffusion architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。