让扩散模型精准控制图像风格,不改内容只换风格。
Stylistic Attribute Control in Latent Diffusion Models

- 从合成数据学习独立风格编辑方向,实现细粒度控制。
- 相比文本编辑,风格调整更精确且可连续调节。
- 适合需要稳定风格迁移的创意设计与图像编辑场景。
文本到图像的扩散模型已革新图像生成与编辑,但对风格属性的精确控制仍具挑战,常导致内容意外改动。本文提出一种在潜在扩散模型中实现细粒度参数化风格属性控制的方法,通过从合成数据集学习解耦的编辑方向。利用指导组合缩小风格微调模型与基础模型间的域差距,在保留原始图像语义的同时实现风格调整。为确保编辑一致性,引入训练正则化损失,并通过优化空条件嵌入增强DDIM反演以支持真实图像编辑。我们在包含轮廓、局部对比度、水彩效果和几何图案等多种风格属性的合成数据集上进行验证,结果表明,相比现有基于文本的编辑技术,该方法能实现更融合、更精确且连续可调的风格修改。
原文摘要 · Abstract (English)
Text-to-image diffusion models have revolutionized image synthesis and editing, but precise control over stylistic attributes remains a challenge, often causing unintended content modifications. We propose an approach for fine-grained parametric control of stylistic attributes in latent diffusion models by learning disentangled editing directions from synthetic datasets. We use guidance composition to close the domain gap between stylistically finetuned and foundation models, preserving the original image semantics while applying stylistic adjustments. To ensure consistent edits, we introduce a training regularization loss and enhance DDIM inversion with optimized null-conditional embeddings for real image editing. We validate our approach by learning from stylistically filtered synthetic datasets varying a range of stylistic attributes, including outlines, local contrast, watercolorization effects, and geometric patterns. Our evaluations demonstrate that compared to current text-based editing techniques, our method offers well-integrated, more precise and continuously adjustable stylistic modifications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。