arXiv:2411.10183cs.CV2024-11中稿 · ISCAS2024被引 1

用视觉问答评估图文生成中每个物体的对齐精度

Visual question answering based evaluation metrics for text-to-image generation

  • 用ChatGPT生成针对图像的提问,实现细粒度图文对齐评估
  • 在COCO、MS-COCO数据集上验证,能同时评估对齐与图像质量
  • 支持调节对齐与质量权重,适合模型优化与对比实验

文本到图像生成与文本引导的图像编辑受到广泛关注。然而,现有主流评估方法难以判断输入文本中的所有信息是否准确反映在生成图像中,且多聚焦于整体图文对齐。本文提出新评估指标,针对每个独立物体评估输入文本与生成图像的对齐程度。首先利用ChatGPT根据输入文本生成对应问题;随后采用视觉问答(VQA)评估图像对问题的回答准确性,实现更细致的对齐分析。此外,结合无参考图像质量评估(NR-IQA),同时衡量图文对齐与生成图像质量。实验表明,该方法在COCO和MS-COCO数据集上优于现有指标,可灵活调整对齐与质量权重,实现更全面的评估。

原文摘要 · Abstract (English)

Text-to-image generation and text-guided image manipulation have received considerable attention in the field of image generation tasks. However, the mainstream evaluation methods for these tasks have difficulty in evaluating whether all the information from the input text is accurately reflected in the generated images, and they mainly focus on evaluating the overall alignment between the input text and the generated images. This paper proposes new evaluation metrics that assess the alignment between input text and generated images for every individual object. Firstly, according to the input text, chatGPT is utilized to produce questions for the generated images. After that, we use Visual Question Answering(VQA) to measure the relevance of the generated images to the input text, which allows for a more detailed evaluation of the alignment compared to existing methods. In addition, we use Non-Reference Image Quality Assessment(NR-IQA) to evaluate not only the text-image alignment but also the quality of the generated images. Experimental results show that our proposed evaluation approach is the superior metric that can simultaneously assess finer text-image alignment and image quality while allowing for the adjustment of these ratios.

图文对齐VQA生成评估图像质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。