arXiv:2504.09291cs.CVcs.MM2025-04被引 2

首个可解释的局部AI生成图像质量评估数据集与模型。

Towards Explainable Partial-AIGC Image Quality Assessment

  • 构建15000张局部AI编辑图像数据集,支持多维人类评分。
  • 提出三阶段训练框架,实现区域定位、质量打分与解释生成。
  • 适合研究AI图像质量评估、可解释性生成的学者和工程师。

AI视觉生成技术的快速发展推动了自然场景图像(NSIs)中局部逼真编辑效果的突破。尽管已有大量关于全量AI生成图像(AGIs)质量评估(IQA)的研究,但针对局部AI生成图像(PAIs)——即包含局部AI编辑的图像——的质量评估仍几乎空白。为此,我们构建了首个面向可解释局部AI生成图像质量评估(EPAIQA)的大规模数据集EPAIQA-15K,包含15,000张带有不同区域局部AI操作的图像及超过30万条多维度人类评分。基于此,我们利用大型多模态模型(LMMs),提出一种三阶段训练范式:逐步训练模型完成编辑区域定位、定量质量评分与质量解释生成。最终,我们开发出EPAIQA系列模型,具备可解释的质量反馈能力。本工作在局部AI生成图像的感知质量评估领域具有开创意义。

原文摘要 · Abstract (English)

The rapid advancement of AI-driven visual generation technologies has catalyzed significant breakthroughs in image manipulation, particularly in achieving photorealistic localized editing effects on natural scene images (NSIs). Despite extensive research on image quality assessment (IQA) for AI-generated images (AGIs), most studies focus on fully AI-generated outputs (e.g., text-to-image generation), leaving the quality assessment of partial-AIGC images (PAIs)-images with localized AI-driven edits an almost unprecedented field. Motivated by this gap, we construct the first large-scale PAI dataset towards explainable partial-AIGC image quality assessment (EPAIQA), the EPAIQA-15K, which includes 15K images with localized AI manipulation in different regions and over 300K multi-dimensional human ratings. Based on this, we leverage large multi-modal models (LMMs) and propose a three-stage model training paradigm. This paradigm progressively trains the LMM for editing region grounding, quantitative quality scoring, and quality explanation. Finally, we develop the EPAIQA series models, which possess explainable quality feedback capabilities. Our work represents a pioneering effort in the perceptual IQA field for comprehensive PAI quality assessment.

图像质量评估可解释性局部生成多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。