arXiv:2510.01010cs.CV2025-10被引 4

ImageDoctor通过多维度诊断,精准定位文生图缺陷并提升生成质量。

ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning

  • 采用'看-想-判'框架,分步诊断图像各维度问题
  • 生成热力图标记错误区域,支持像素级反馈
  • 比单一数值评估提升10%生成质量,适合模型调优

文本到图像(T2I)模型的快速发展带来了对可靠人类偏好建模的需求,尤其在强化学习用于偏好对齐的背景下更为突出。然而,现有方法通常用单一标量衡量生成图像质量,难以提供全面且可解释的反馈。为此,我们提出ImageDoctor,一个统一的多维度T2I模型评估框架,从合理性、语义对齐性、美学和整体质量四个互补维度评估图像质量。该框架还生成像素级缺陷热力图,标记错位或不合理的区域,并可用作密集奖励信号以优化T2I模型。受医疗诊断启发,我们引入“看-想-判”范式,先定位潜在缺陷,再生成推理过程,最后输出量化评分。基于视觉语言模型,结合监督微调与强化学习训练,ImageDoctor在多个数据集上展现出与人类偏好高度一致的表现,验证了其作为评估指标的有效性。当作为偏好对齐的奖励模型使用时,相比标量奖励模型,生成质量提升10%。

原文摘要 · Abstract (English)

The rapid advancement of text-to-image (T2I) models has increased the need for reliable human preference modeling, a demand further amplified by recent progress in reinforcement learning for preference alignment. However, existing approaches typically quantify the quality of a generated image using a single scalar, limiting their ability to provide comprehensive and interpretable feedback on image quality. To address this, we introduce ImageDoctor, a unified multi-aspect T2I model evaluation framework that assesses image quality across four complementary dimensions: plausibility, semantic alignment, aesthetics, and overall quality. ImageDoctor also provides pixel-level flaw indicators in the form of heatmaps, which highlight misaligned or implausible regions, and can be used as a dense reward for T2I model preference alignment. Inspired by the diagnostic process, we improve the detail sensitivity and reasoning capability of ImageDoctor by introducing a "look-think-predict" paradigm, where the model first localizes potential flaws, then generates reasoning, and finally concludes the evaluation with quantitative scores. Built on top of a vision-language model and trained through a combination of supervised fine-tuning and reinforcement learning, ImageDoctor demonstrates strong alignment with human preference across multiple datasets, establishing its effectiveness as an evaluation metric. Furthermore, when used as a reward model for preference tuning, ImageDoctor significantly improves generation quality -- achieving an improvement of 10% over scalar-based reward models.

文生图模型评估视觉推理强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。