构建细粒度评估图像真实与合理性的数据集,助力生成模型优化
Q-REAL: Towards Realism and Plausibility Evaluation for AI-Generated Content
- 构建包含3088张图像的Q-Real数据集,标注实体位置并设计多维度判断题
- 在多模态大模型上验证,可精准识别生成图像的真实感与合理性缺陷
- 适合研究生成模型评估、多模态理解及模型优化的学者使用
AI生成内容的质量评估对模型能力衡量和优化至关重要。现有评估数据集与模型多仅提供单一质量评分,难以提供针对性改进指导。在当前文本到图像生成应用中,真实感与合理性是两个关键维度,随着统一生成-理解模型的发展,沿这两个维度进行细粒度评估更有效提升生成性能。为此,我们提出Q-Real,一个用于评估生成图像真实感与合理性的新数据集。Q-Real包含3,088张由主流文本到图像模型生成的图像,每张图像均标注主要实体位置,并提供针对这些实体在真实感与合理性维度上的判断问题与归因描述。考虑到多模态大模型(MLLM)近期进展使其具备细粒度评估能力,我们构建了Q-Real Bench,用于评估其在判断与基于推理的定位任务上的表现。为增强MLLM能力,我们设计了微调框架,并在多个MLLM上使用该数据集进行实验。实验结果表明,本数据集质量高、意义显著,基准测试全面。数据集与代码将在发表后公开。
原文摘要 · Abstract (English)
Quality assessment of AI-generated content is crucial for evaluating model capability and guiding model optimization. However, most existing quality assessment datasets and models provide only a single quality score, which is too coarse to offer targeted guidance for improving generative models. In current applications of AI-generated images, realism and plausibility are two critical dimensions, and with the emergence of unified generation-understanding models, fine-grained evaluation along these dimensions becomes especially effective for improving generative performance. Therefore, we introduce Q-Real, a novel dataset for fine-grained evaluation of realism and plausibility in AI-generated images. Q-Real consists of 3,088 images generated by popular text-to-image models. For each image, we annotate the locations of major entities and provide a set of judgment questions and attribution descriptions for these along the dimensions of realism and plausibility. Considering that recent advances in multi-modal large language models (MLLMs) enable fine-grained evaluation of AI-generated images, we construct Q-Real Bench to evaluate them on two tasks: judgment and grounding with reasoning. Finally, to enhance MLLM capabilities, we design a fine-tuning framework and conduct experiments on multiple MLLMs using our dataset. Experimental results demonstrate the high quality and significance of our dataset and the comprehensiveness of the benchmark. Dataset and code will be released upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。