测试文字生成图像模型在分类体系概念上的表现,发现现有模型效果差异大。
Do I look like a `cat.n.01` to you? A Taxonomy Image Generation Benchmark
- 构建分类体系图像生成评测基准,涵盖常识与随机词表概念。
- Playground-v2和FLUX在多指标中表现最佳,检索方法效果差。
- 首次用GPT-4进行成对评估,适合结构化数据自动化构建研究者。
本文探索了在零样本设置下使用文本到图像模型生成分类体系概念图像的可行性。尽管基于文本的分类体系扩展方法已成熟,但视觉维度潜力尚未被挖掘。为此,我们提出了一个全面的分类体系图像生成评测基准,用于评估模型理解分类概念并生成相关高质量图像的能力。该基准包含常识性和随机采样的WordNet概念,以及大语言模型生成的预测。对12个模型使用9种新颖的分类体系相关文本到图像指标和人工反馈进行评估。此外,我们首次采用基于GPT-4的成对评估方法进行图像生成评价。实验结果表明,模型排名与标准文本到图像任务显著不同。Playground-v2和FLUX在各项指标和子集上均表现最优,而基于检索的方法表现较差。这些发现凸显了自动化构建结构化数据资源的潜力。
原文摘要 · Abstract (English)
This paper explores the feasibility of using text-to-image models in a zero-shot setup to generate images for taxonomy concepts. While text-based methods for taxonomy enrichment are well-established, the potential of the visual dimension remains unexplored. To address this, we propose a comprehensive benchmark for Taxonomy Image Generation that assesses models' abilities to understand taxonomy concepts and generate relevant, high-quality images. The benchmark includes common-sense and randomly sampled WordNet concepts, alongside the LLM generated predictions. The 12 models are evaluated using 9 novel taxonomy-related text-to-image metrics and human feedback. Moreover, we pioneer the use of pairwise evaluation with GPT-4 feedback for image generation. Experimental results show that the ranking of models differs significantly from standard T2I tasks. Playground-v2 and FLUX consistently outperform across metrics and subsets and the retrieval-based approach performs poorly. These findings highlight the potential for automating the curation of structured data resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。