用小模型复刻大模型的图像评估能力,成本更低效果更好。
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
- 将评估任务拆解为简单子任务,降低学习难度
- 7B小模型在多项指标上超越大模型基线4.6%以上
- 提供带推理过程的评测基准,适合模型对比与优化
受扩散模型进展推动,文本到图像生成已取得显著突破,对生成图像自动质量评估的需求日益迫切。现有最先进自动评估方法高度依赖多模态大语言模型(MLLM),尤其是性能强大的商业模型如GPT-4o。尽管这些模型效果优异,但高昂成本限制了大规模评估的可扩展性。采用开源MLLM是替代方案,但其处理多模态数据的能力远逊于商业模型。为此,我们首先基于GPT-4o提出一种任务分解评估框架,自动构建新训练数据集,将复杂评估任务解耦为更简单的子任务,有效降低学习复杂度。基于该数据集,设计创新训练策略,成功将GPT-4o的评估能力蒸馏至一个7B规模的开源MLLM MiniCPM-V-2.6。此外,为可靠全面评估现有工作及本模型,我们人工标注了一个包含链式思维解释和质量评分的元评估基准。实验结果表明,所提出的蒸馏模型在与人类判断的斯皮尔曼和肯德尔相关性上,显著优于当前最先进基线VIEScore,提升超过4.6%。
原文摘要 · Abstract (English)
Driven by the remarkable progress in diffusion models, text-to-image generation has made significant strides, creating a pressing demand for automatic quality evaluation of generated images. Current state-of-the-art automatic evaluation methods heavily rely on Multi-modal Large Language Models (MLLMs), particularly powerful commercial models like GPT-4o. While these models are highly effective, their substantial costs limit scalability in large-scale evaluations. Adopting open-source MLLMs is an alternative; however, their performance falls short due to significant limitations in processing multi-modal data compared to commercial MLLMs. To tackle these problems, we first propose a task decomposition evaluation framework based on GPT-4o to automatically construct a new training dataset, where the complex evaluation task is decoupled into simpler sub-tasks, effectively reducing the learning complexity. Based on this dataset, we design innovative training strategies to effectively distill GPT-4o's evaluation capabilities into a 7B open-source MLLM, MiniCPM-V-2.6. Furthermore, to reliably and comprehensively assess prior works and our proposed model, we manually annotate a meta-evaluation benchmark that includes chain-of-thought explanations alongside quality scores for generated images. Experimental results demonstrate that our distilled open-source MLLM significantly outperforms the current state-of-the-art GPT-4o-base baseline, VIEScore, with over 4.6\% improvement in Spearman and Kendall correlations with human judgments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。