用问答评估图文生成中的事实幻觉,发现主流模型常编造事实。
Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering
- 通过视觉问答自动检测图像是否忠实于文本描述
- 在1200组图像-文本对上验证,模型答对率普遍低于60%
- 数据集含1000个精校问题,适合研究真实性和可靠性
尽管文本到图像(TTI)生成模型取得了显著进展,现有研究忽略了这些模型是否准确传递事实信息的问题。本文聚焦图像幻觉问题,即生成图像未能忠实地呈现事实内容。为此,我们提出 I-HallA(基于问答的图像幻觉评估),一种通过视觉问答(VQA)衡量生成图像事实性的自动化评估指标。同时,我们构建了 I-HallA v1.0 基准数据集。该数据集由多代理 GPT-4 Omni 系统生成高质量问答对,并经人工判断确保准确性,共涵盖 1.2 千组跨九类的图像-文本对,包含 1000 个精心设计的问题,覆盖多种组合挑战。我们使用 I-HallA 评估五种主流 TTI 模型,发现其在事实表达上普遍存在缺陷。此外,我们通过强斯皮尔曼相关性(ρ=0.95)验证了该指标与人类判断的一致性。我们认为,该数据集和评估方法可为发展事实准确的 TTI 模型奠定基础。
原文摘要 · Abstract (English)
Despite the impressive success of text-to-image (TTI) generation models, existing studies overlook the issue of whether these models accurately convey factual information. In this paper, we focus on the problem of image hallucination, where images created by generation models fail to faithfully depict factual content. To address this, we introduce I-HallA (Image Hallucination evaluation with Question Answering), a novel automated evaluation metric that measures the factuality of generated images through visual question answering (VQA). We also introduce I-HallA v1.0, a curated benchmark dataset for this purpose. As part of this process, we develop a pipeline that generates high-quality question-answer pairs using multiple GPT-4 Omni-based agents, with human judgments to ensure accuracy. Our evaluation protocols measure image hallucination by testing if images from existing TTI models can correctly respond to these questions. The I-HallA v1.0 dataset comprises 1.2K diverse image-text pairs across nine categories with 1,000 rigorously curated questions covering various compositional challenges. We evaluate five TTI models using I-HallA and reveal that these state-of-the-art models often fail to accurately convey factual information. Moreover, we validate the reliability of our metric by demonstrating a strong Spearman correlation ($ρ$=0.95) with human judgments. We believe our benchmark dataset and metric can serve as a foundation for developing factually accurate TTI generation models. Additional resources can be found on our project page: https://sgt-lim.github.io/I-HallA/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。