用强化学习让开源模型自动评图,省去人工标注却更懂人类审美。
T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation
- 用粗粒度评分训练模型,不依赖高质量人工点评。
- 在三个评测集上接近人类判断,且解释更清晰可信。
- 适合想低成本评估生成图像质量的研究者和开发者。
基于扩散模型的文本到图像生成技术快速发展,亟需可解释的自动化评估方法以减少人工标注负担。现有监督微调方法依赖高质量评论数据集,这些数据要么由商业大模型生成(存在偏见与不一致),要么由人工标注(成本高昂),限制了可扩展性和泛化能力。为此,我们提出 T2I-Eval-R1,一种仅使用粗粒度质量评分的强化学习框架,训练开源多模态大语言模型作为图像评估器。该方法将分组相对策略优化(GRPO)融入指令微调过程,使模型仅凭易获取的评分或偏好即可生成标量分数与可解释推理链。同时引入连续奖励机制,促进评分多样性并提供稳定优化信号,提升评估鲁棒性与区分度。在三个主流文本到图像元评估基准上的实验表明,T2I-Eval-R1 在与人类评估对齐度上显著优于强基线方法,并能提供更准确、可解释的评分理由。
原文摘要 · Abstract (English)
The rapid progress in diffusion-based text-to-image (T2I) generation has created an urgent need for interpretable automatic evaluation methods that can assess the quality of generated images, therefore reducing the human annotation burden. To reduce the prohibitive cost of relying on commercial models for large-scale evaluation, and to improve the reasoning capabilities of open-source models, recent research has explored supervised fine-tuning (SFT) of multimodal large language models (MLLMs) as dedicated T2I evaluators. However, SFT approaches typically rely on high-quality critique datasets, which are either generated by proprietary LLMs-with potential issues of bias and inconsistency-or annotated by humans at high cost, limiting their scalability and generalization. To address these limitations, we propose T2I-Eval-R1, a novel reinforcement learning framework that trains open-source MLLMs using only coarse-grained quality scores, thereby avoiding the need for annotating high-quality interpretable evaluation rationale. Our approach integrates Group Relative Policy Optimization (GRPO) into the instruction-tuning process, enabling models to generate both scalar scores and interpretable reasoning chains with only easy accessible annotated judgment scores or preferences. Furthermore, we introduce a continuous reward formulation that encourages score diversity and provides stable optimization signals, leading to more robust and discriminative evaluation behavior. Experimental results on three established T2I meta-evaluation benchmarks demonstrate that T2I-Eval-R1 achieves significantly higher alignment with human assessments and offers more accurate interpretable score rationales compared to strong baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。