arXiv:2410.11824cs.CV2024-10被引 7

评测图像生成模型对真实物体的准确表现能力,发现主流模型仍存在细节失真问题。

KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities

  • 构建KITTEN基准,聚焦真实世界实体的视觉生成质量
  • 先进模型在地标、动物等实体上仍难还原准确细节
  • 检索增强模型依赖参考图,难生成创意新构型

文本到图像生成技术虽显著提升图像质量,但现有评估多关注美学或与文本提示的匹配度,难以判断模型是否能准确呈现多样化真实视觉实体。为此,我们提出KITTEN——一个针对真实世界实体的视觉生成知识密集型评测基准。通过该基准,我们系统评估了最新的文本到图像模型及检索增强模型在生成地标、动物等实体方面的能力。结合精心设计的人工评价、自动指标和多模态大模型(MLLM)评估,结果表明:即使最先进的文本到图像模型,在实体视觉细节上也存在明显偏差。检索增强模型虽通过引入参考图像提升了实体保真度,但过度依赖参考图,在创造性文本提示下难以生成新颖的实体组合。

原文摘要 · Abstract (English)

Recent advances in text-to-image generation have improved the quality of synthesized images, but evaluations mainly focus on aesthetics or alignment with text prompts. Thus, it remains unclear whether these models can accurately represent a wide variety of realistic visual entities. To bridge this gap, we propose KITTEN, a benchmark for Knowledge-InTensive image generaTion on real-world ENtities. Using KITTEN, we conduct a systematic study of the latest text-to-image models and retrieval-augmented models, focusing on their ability to generate real-world visual entities, such as landmarks and animals. Analysis using carefully designed human evaluations, automatic metrics, and MLLM evaluations show that even advanced text-to-image models fail to generate accurate visual details of entities. While retrieval-augmented models improve entity fidelity by incorporating reference images, they tend to over-rely on them and struggle to create novel configurations of the entity in creative text prompts.

图像生成视觉评测实体保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。