FAGER让文生图模型更懂事实,自动检测并修正图像中的事实错误。
FAGER: Factually Grounded Evaluation and Refinement of Text-to-Image Models

- 用大模型生成事实规则,结合视觉验证构建可评分的评估框架。
- 在五个跨领域数据集上,比现有方法更准确识别真实图像。
- 无需训练即可改进生成图像,适合科研、历史等需要精准表达的场景。
现有文生图评估指标主要关注生成图像是否符合提示中明确陈述的信息,但难以捕捉隐含、外部验证或身份定义的事实要求。因此,在涉及科学知识、历史事实、产品信息或文化特定概念的提示中,其对事实正确性的评估能力不足。我们提出FAGER(Factually Grounded Evaluation and Refinement),一个基于智能体的框架,用于评估生成图像是否准确反映提示中可视觉验证的事实,同时提供可操作的改进建议。FAGER首先通过大模型生成事实并结合参考图像进行视觉事实提取与验证,构建结构化事实评分标准,再转化为问答对供视觉语言模型评估。为验证其作为事实性度量的有效性,我们引入事实性A/B测试,衡量指标是否偏好真实参考图像而非生成图像。在涵盖科学、历史、产品、文化和知识密集型概念的五个数据集上,FAGER均显著优于现有方法。此外,我们证明FAGER可在完全无训练的情况下用于优化文生图输出,实现多个数据集上的事实性显著提升。
原文摘要 · Abstract (English)
Existing text-to-image (T2I) evaluation metrics mainly assess whether generated images align with information explicitly stated in the prompt, but often fail to capture factual requirements that are implicit, externally grounded, or identity-defining. As a result, they are not well suited for evaluating factual correctness in prompts involving scientific knowledge, historical facts, products, or culture-specific concepts. We propose FActually Grounded Evaluation and Refinement (FAGER), an agentic framework that evaluates whether generated images correctly reflect visually verifiable facts grounded in or implied by the prompt, while also providing actionable feedback for improvement. FAGER first constructs a structured factual rubric by combining LLM-based fact proposal with reference-guided visual fact extraction and verification, then converts the rubric into question-answer pairs for VLM-based evaluation. To validate FAGER as a factuality metric, we introduce a Factual A/B test, which measures whether a metric prefers factual reference images over corresponding generated images. Across five datasets spanning science, history, products, culture, and knowledge-intensive concepts, FAGER consistently outperforms prior metrics on this test. We further show that FAGER can be used to refine T2I outputs in a fully training-free manner, yielding substantial factuality gains across datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。