arXiv:2509.21257cs.CVcs.CL2025-09中稿 · NeurIPS

提出用幻觉上限评估文生图模型,揭示隐藏偏见

Hallucination as an Upper Bound: A New Perspective on Text-to-Image Evaluation

  • 将幻觉定义为基于偏见的生成偏差,分三类:属性、关系、对象
  • 发现现有评估只看提示对齐,忽略额外生成内容
  • 为文生图模型提供更全面的评估框架,适合研究者参考

在语言和视觉-语言模型中,幻觉通常指模型基于自身先验知识或偏见生成的内容,而非来自输入。尽管该现象在这些领域已有研究,但尚未明确应用于文生图(T2I)生成模型。现有评估主要关注对齐性,即检查提示中指定元素是否出现,却忽略了模型在提示之外生成的内容。本文主张将文生图中的幻觉定义为由偏见驱动的偏离,并提出一个包含三类的分类体系:属性、关系和对象幻觉。这一框架为评估设定了上限,揭示了隐藏偏见,为文生图模型的更丰富评估提供了基础。

原文摘要 · Abstract (English)

In language and vision-language models, hallucination is broadly understood as content generated from a model's prior knowledge or biases rather than from the given input. While this phenomenon has been studied in those domains, it has not been clearly framed for text-to-image (T2I) generative models. Existing evaluations mainly focus on alignment, checking whether prompt-specified elements appear, but overlook what the model generates beyond the prompt. We argue for defining hallucination in T2I as bias-driven deviations and propose a taxonomy with three categories: attribute, relation, and object hallucinations. This framing introduces an upper bound for evaluation and surfaces hidden biases, providing a foundation for richer assessment of T2I models.

文生图幻觉评估模型偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。