用大模型精准解析模糊颜色词,提升文本生成图像的配色准确性。
Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion Generation
- 用大模型解析模糊颜色词,再在文本嵌入空间调整颜色混合
- 无需训练或参考图,颜色对齐准确率显著提升
- 适合时尚、产品设计等需精确配色的场景
文本到图像生成中准确的色彩对齐对于时尚、产品可视化和室内设计等应用至关重要,但现有扩散模型在处理如‘蒂芙尼蓝’‘青柠绿’‘亮粉红’等复杂颜色术语时仍存在偏差,常与人类意图不符。现有方法依赖交叉注意力调整、参考图像或微调,却无法系统解决颜色描述歧义问题。为此,我们提出一种无需训练的框架,通过大语言模型(LLM)解析文本提示中的模糊颜色词,并基于其在CIELAB色彩空间中的空间关系,直接在文本嵌入空间优化颜色融合策略。该方法无需额外训练或外部参考图像,在保持图像质量的同时显著提升色彩对齐精度,有效弥合了文本语义与视觉生成之间的差距。
原文摘要 · Abstract (English)
Accurate color alignment in text-to-image (T2I) generation is critical for applications such as fashion, product visualization, and interior design, yet current diffusion models struggle with nuanced and compound color terms (e.g., Tiffany blue, lime green, hot pink), often producing images that are misaligned with human intent. Existing approaches rely on cross-attention manipulation, reference images, or fine-tuning but fail to systematically resolve ambiguous color descriptions. To precisely render colors under prompt ambiguity, we propose a training-free framework that enhances color fidelity by leveraging a large language model (LLM) to disambiguate color-related prompts and guiding color blending operations directly in the text embedding space. Our method first employs a large language model (LLM) to resolve ambiguous color terms in the text prompt, and then refines the text embeddings based on the spatial relationships of the resulting color terms in the CIELAB color space. Unlike prior methods, our approach improves color accuracy without requiring additional training or external reference images. Experimental results demonstrate that our framework improves color alignment without compromising image quality, bridging the gap between text semantics and visual generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。