测试文生图模型对隐含色彩概念的理解能力,发现其表现远未达人类水平。
ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models
- 构建包含1281个隐含色彩概念的评估基准,基于6584条人工标注。
- 9个主流文生图模型在抽象概念上生成色彩分布差异大,平均准确率不足50%。
- 即使使用无分类器引导,模型仍难以理解情绪等抽象语义中的色彩含义。
文生图(T2I)模型在从文本生成高质量图像方面取得了显著进展,但其对颜色与概念关联的能力仍局限于显式颜色名称或代码,而对情感、视觉状态等隐含概念的处理能力尚未充分探索。为弥补这一空白,我们提出了ColorConceptBench,一个由专家标注的基准,通过概率性颜色分布系统评估颜色-概念关联。该基准考察了1281个隐含颜色概念,基于6584条人类标注。对九个领先文生图模型的评估显示,性能在不同语义类别间差异显著,且模型对抽象语义的敏感度明显不足。这些局限即使在推理时采用无分类器引导缩放也依然存在,表明实现类人色彩理解需要从根本上改变模型对隐含语义的学习与表征方式。
原文摘要 · Abstract (English)
Text-to-image (T2I) models have advanced considerably in generating high-quality images from textual descriptions. However, their ability to associate colors with concepts remains largely constrained to explicit color names or codes, while their capacity to handle \emph{implicit concepts}, such as emotions and visual states, remains underexplored. To address this gap, we introduce ColorConceptBench, an expert-annotated benchmark that systematically evaluates color-concept associations through probabilistic color distributions. ColorConceptBench moves beyond explicit color specifications by examining how models interpret 1,281 implicit color concepts, grounded in 6,584 human annotations. Our evaluation of nine leading T2I models reveals that performance varies substantially across semantic categories, and models exhibit a significant lack of sensitivity to abstract semantics. These limitations persist even when applying classifier-free guidance scaling at inference time, suggesting that achieving human-like color understanding demands a shift in how models learn and represent implicit semantic meaning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。