arXiv:2412.05722cs.CV2024-12被引 9

用场景图问答评估文生图模型幻觉,更准更可解释。

Evaluating Hallucination in Text-to-Image Diffusion Models with Scene-Graph based Question-Answering Agent

  • 基于场景图和大模型问答,自动检测图文不一致
  • 在12000张图上验证,与人工评分高度一致
  • 适合研究文生图质量评估的开发者和研究人员

当前文生图模型多依赖主观的人工评估来判断图像与文本提示的一致性,缺乏可复现的量化工具。本文提出一种基于大语言模型的评估方法,通过提取场景图并生成问答任务,识别图像与文本之间的不一致(即‘幻觉’问题),记录幻觉类型与频率,并提供接近人类标准的综合评分。研究构建了包含12,000张合成图像的数据集,基于1,000个复合提示,使用三种先进文生图模型生成,所有图像均经人工评分验证。实验表明,该方法在对齐人类评分模式方面优于现有评估指标,且更具可控性和可解释性。

原文摘要 · Abstract (English)

Contemporary Text-to-Image (T2I) models frequently depend on qualitative human evaluations to assess the consistency between synthesized images and the text prompts. There is a demand for quantitative and automatic evaluation tools, given that human evaluation lacks reproducibility. We believe that an effective T2I evaluation metric should accomplish the following: detect instances where the generated images do not align with the textual prompts, a discrepancy we define as the `hallucination problem' in T2I tasks; record the types and frequency of hallucination issues, aiding users in understanding the causes of errors; and provide a comprehensive and intuitive scoring that close to human standard. To achieve these objectives, we propose a method based on large language models (LLMs) for conducting question-answering with an extracted scene-graph and created a dataset with human-rated scores for generated images. From the methodology perspective, we combine knowledge-enhanced question-answering tasks with image evaluation tasks, making the evaluation metrics more controllable and easier to interpret. For the contribution on the dataset side, we generated 12,000 synthesized images based on 1,000 composited prompts using three advanced T2I models. Subsequently, we conduct human scoring on all synthesized images and prompt pairs to validate the accuracy and effectiveness of our method as an evaluation metric. All generated images and the human-labeled scores will be made publicly available in the future to facilitate ongoing research on this crucial issue. Extensive experiments show that our method aligns more closely with human scoring patterns than other evaluation metrics.

文生图幻觉评估场景图LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。