无需微调即可精确控制扩散模型中的颜色,实现自由配色。
Leveraging Semantic Attribute Binding for Free-Lunch Color Control in Diffusion Models
- 通过重连文本与图像特征的语义绑定实现无训练颜色控制。
- 在多种物体类别上实现高精度颜色匹配,优于现有方法。
- 适合需要快速、灵活配色的生成式设计场景。
近期文本到图像(T2I)扩散模型在属性控制方面取得显著进展,但精确的颜色指定仍是根本挑战。现有方法如ColorPeel依赖模型个性化,需额外优化且灵活性受限。本文提出ColorWave,一种无需训练的新方法,可在不微调的前提下实现扩散模型中精确的RGB级颜色控制。通过对IP-Adapter中交叉注意力机制的系统分析,我们发现文本颜色描述与参考图像特征之间存在隐式绑定。基于此洞察,方法通过重连这些绑定,强制实现精确颜色指派,同时保留预训练模型的生成能力。实验表明,该方法在生成质量与多样性上保持优异,在多种物体类别上均优于先前方法,准确性和适用性更佳。全面评估验证了ColorWave为结构化、色彩一致的扩散图像合成树立了新范式。
原文摘要 · Abstract (English)
Recent advances in text-to-image (T2I) diffusion models have enabled remarkable control over various attributes, yet precise color specification remains a fundamental challenge. Existing approaches, such as ColorPeel, rely on model personalization, requiring additional optimization and limiting flexibility in specifying arbitrary colors. In this work, we introduce ColorWave, a novel training-free approach that achieves exact RGB-level color control in diffusion models without fine-tuning. By systematically analyzing the cross-attention mechanisms within IP-Adapter, we uncover an implicit binding between textual color descriptors and reference image features. Leveraging this insight, our method rewires these bindings to enforce precise color attribution while preserving the generative capabilities of pretrained models. Our approach maintains generation quality and diversity, outperforming prior methods in accuracy and applicability across diverse object categories. Through extensive evaluations, we demonstrate that ColorWave establishes a new paradigm for structured, color-consistent diffusion-based image synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。