arXiv:2505.18730cs.CV2025-05被引 3

评测文生图模型对提示外真实世界知识的对齐能力,发现顶尖模型仍有明显不足。

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation

  • 构建覆盖六类场景的2000+提示数据集,评估图像与真实世界知识的隐含对齐
  • 提出ABPScore指标,用多模态大模型打分,与人类判断高度一致
  • 提出推理时知识注入策略,使模型在关键样本上性能提升约43%

近期文本到图像(T2I)生成模型取得显著进展,可从文本提示生成高保真图像。然而,现有评测基准主要关注图像与提示的显式对齐,忽视了提示之外的真实世界知识对齐。为填补这一空白,我们提出Align Beyond Prompts(ABP),一个综合性评测基准,用于衡量生成图像与超出提示范围的真实世界知识的对齐程度。ABP包含超过2000个精心设计的提示,覆盖六大现实场景。我们进一步引入ABPScore,利用现有多模态大语言模型(MLLMs)评估图像与提示外世界知识的对齐度,该指标与人类判断具有强相关性。通过对8个主流T2I模型的全面评估,发现即使最先进的模型如GPT-4o,在将简单真实世界知识融入图像方面仍存在明显局限。为此,我们在ABP中提出一种无需训练的策略——推理时知识注入(ITKI)。在200个挑战性样本上应用该策略后,ABPScore提升约43%。数据集与代码已公开于https://github.com/smile365317/ABP。

原文摘要 · Abstract (English)

Recent text-to-image (T2I) generation models have advanced significantly, enabling the creation of high-fidelity images from textual prompts. However, existing evaluation benchmarks primarily focus on the explicit alignment between generated images and prompts, neglecting the alignment with real-world knowledge beyond prompts. To address this gap, we introduce Align Beyond Prompts (ABP), a comprehensive benchmark designed to measure the alignment of generated images with real-world knowledge that extends beyond the explicit user prompts. ABP comprises over 2,000 meticulously crafted prompts, covering real-world knowledge across six distinct scenarios. We further introduce ABPScore, a metric that utilizes existing Multimodal Large Language Models (MLLMs) to assess the alignment between generated images and world knowledge beyond prompts, which demonstrates strong correlations with human judgments. Through a comprehensive evaluation of 8 popular T2I models using ABP, we find that even state-of-the-art models, such as GPT-4o, face limitations in integrating simple real-world knowledge into generated images. To mitigate this issue, we introduce a training-free strategy within ABP, named Inference-Time Knowledge Injection (ITKI). By applying this strategy to optimize 200 challenging samples, we achieved an improvement of approximately 43% in ABPScore. The dataset and code are available in https://github.com/smile365317/ABP.

文生图知识对齐评测基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。