arXiv:2508.19791cs.CV2025-08被引 2

研究文本生成图像中颜色语义对齐问题,提出新编辑方法提升多色提示生成质量。

Color Bind: Exploring Color Perception in Text-to-Image Models

  • 针对多对象颜色提示的语义错位,设计专用图像编辑技术。
  • 在多种扩散模型上验证,显著提升多色生成准确率。
  • 适合关注生成细节一致性与可控性的研究人员。

文本到图像生成近年来取得显著进展,使用户可通过文本创建高质量图像。然而,当前方法在处理复杂多对象提示时难以准确捕捉语义,导致生成结果与提示不符。现有工作多通过推理阶段修改去噪网络的注意力层来缓解此问题,但评估常依赖粗略指标(如文本与图像CLIP嵌入的余弦相似度)或人工评价,难以大规模应用。本文以颜色这一基础属性为切入点,构建严谨评估体系。分析发现,预训练模型在生成包含多个颜色属性的图像时表现远差于单色提示,且现有推理阶段技术与编辑方法均无法可靠解决该问题。为此,我们提出一种专门针对多对象语义对齐的图像编辑方法。实验表明,该方法在多种基于扩散模型的生成系统上,显著优于现有基准,全面提升了多色提示下的生成质量。

原文摘要 · Abstract (English)

Text-to-image generation has recently seen remarkable success, granting users with the ability to create high-quality images through the use of text. However, contemporary methods face challenges in capturing the precise semantics conveyed by complex multi-object prompts. Consequently, many works have sought to mitigate such semantic misalignments, typically via inference-time schemes that modify the attention layers of the denoising networks. However, prior work has mostly utilized coarse metrics, such as the cosine similarity between text and image CLIP embeddings, or human evaluations, which are challenging to conduct on a larger-scale. In this work, we perform a case study on colors -- a fundamental attribute commonly associated with objects in text prompts, which offer a rich test bed for rigorous evaluation. Our analysis reveals that pretrained models struggle to generate images that faithfully reflect multiple color attributes-far more so than with single-color prompts-and that neither inference-time techniques nor existing editing methods reliably resolve these semantic misalignments. Accordingly, we introduce a dedicated image editing technique, mitigating the issue of multi-object semantic alignment for prompts containing multiple colors. We demonstrate that our approach significantly boosts performance over a wide range of metrics, considering images generated by various text-to-image diffusion-based techniques.

文本生成图像编辑语义对齐颜色感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。