用多轮多模态模型评估AI生成图像的画质、文案匹配度和真伪。
M3-AGIQA: Multimodal, Multi-Round, Multi-Aspect AI-Generated Image Quality Assessment
- 用大语言模型多轮分析图像与文本,逐轮生成描述提升评估深度。
- 在多个数据集上表现优于现有方法,跨数据集泛化能力强。
- 适合需要全面评估AI图像质量的研究者与内容审核团队。
AI生成图像(AIGI)的快速发展带来了新的评估挑战,尤其体现在感知质量、提示词对应性和真实性三个方面。为此,我们提出M3-AGIQA框架,利用多模态大语言模型(MLLMs)实现对视觉与文本域更贴近人类判断的综合评估。该框架采用结构化的多轮评估流程,通过生成并分析中间图像描述,深入揭示上述三方面特性。相比现有方法,M3-AGIQA显著提升了评估结果的鲁棒性与可解释性。在多个基准测试上的实验证明,该方法在所测数据集与评估维度上均达到领先性能,并在多数跨数据集场景中展现出强大泛化能力。代码已开源:https://github.com/strawhatboy/M3-AGIQA。
原文摘要 · Abstract (English)
The rapid advancement of AI-generated image (AIGI) models presents new challenges for evaluating image quality, particularly across three aspects: perceptual quality, prompt correspondence, and authenticity. To address these challenges, we introduce M3-AGIQA, a comprehensive framework that leverages Multimodal Large Language Models (MLLMs) to enable more human-aligned, holistic evaluation of AI-generated images across both visual and textual domains. Besides, our framework features a structured multi-round evaluation process, generating and analyzing intermediate image descriptions to provide deeper insight into these three aspects. By aligning model outputs more closely with human judgment, M3-AGIQA delivers robust and interpretable quality scores. Extensive experiments on multiple benchmarks demonstrate that our method achieves state-of-the-art performance on tested datasets and aspects, and exhibits strong generalizability in most cross-dataset settings. Code is available at https://github.com/strawhatboy/M3-AGIQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。