arXiv:2509.02000cs.CV2025-09

让AI画画时精准用指定配色,还能自由控制颜色变化程度。

Palette Aligned Image Diffusion

  • 将配色视为稀疏直方图,用两个参数控制配色忠实度和色彩多样性。
  • 引入负直方图机制,可屏蔽不想要的颜色,提升配色准确性。
  • 在涵盖冷门与常见色的训练数据上优化,适配各种复杂配色需求。

我们提出Palette-Adapter,一种基于用户指定配色生成图像的新方法。尽管配色是创作中常用且直观的工具,但直接用于图像生成时存在显著歧义与不稳定性。本方法将配色视为稀疏直方图,并引入两个标量控制参数:直方图熵与配色到直方图的距离,以灵活调节配色遵循程度与颜色变化范围。此外,我们设计了负直方图机制,可在标准无分类器引导下抑制特定不期望的色调,从而增强对目标配色的忠实度。为确保在全色域上的良好泛化性,我们在精心构建的数据集上进行训练,该数据集对罕见与常见颜色均有均衡覆盖。实验表明,该方法在多种配色和提示下均能实现稳定、语义一致的图像生成,且在定性、定量评估及用户研究中持续优于现有方法,在强配色一致性与高图像质量之间取得更好平衡。

原文摘要 · Abstract (English)

We introduce the Palette-Adapter, a novel method for conditioning text-to-image diffusion models on a user-specified color palette. While palettes are a compact and intuitive tool widely used in creative workflows, they introduce significant ambiguity and instability when used for conditioning image generation. Our approach addresses this challenge by interpreting palettes as sparse histograms and introducing two scalar control parameters: histogram entropy and palette-to-histogram distance, which allow flexible control over the degree of palette adherence and color variation. We further introduce a negative histogram mechanism that allows users to suppress specific undesired hues, improving adherence to the intended palette under the standard classifier-free guidance mechanism. To ensure broad generalization across the color space, we train on a carefully curated dataset with balanced coverage of rare and common colors. Our method enables stable, semantically coherent generation across a wide range of palettes and prompts. We evaluate our method qualitatively, quantitatively, and through a user study, and show that it consistently outperforms existing approaches in achieving both strong palette adherence and high image quality.

图像生成配色控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。