改进图像生成评估,让模型更懂人类审美偏好。
Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment
- 提出ICT评分机制,衡量图像对文本内容的体现程度。
- 新模型在评分准确率上提升超10%,优化细节与美感。
- 适合关注生成质量与人类审美的研究者与开发者。
当前图像生成系统已超越基础图文对齐,达到高保真与优秀美学水平,但评估框架未能同步演进。本研究发现,基于CLIP和BLIP架构微调的人类偏好奖励模型存在固有缺陷:对细节丰富、美学价值高的图像错误地赋予低分,与真实人类审美严重偏离。为此,我们设计新型评估指标ICT(Image-Contained-Text)分数,通过衡量图像对文本内容的体现程度,实现并超越图文对齐目标。在此基础上,我们仅使用图像模态训练出HP(High-Preference)分数模型,显著提升图像美学与细节质量,同时保持图文对齐。实验表明,所提评估模型相比现有方法评分准确率提升超过10%,并在优化先进文生图模型方面取得显著成效。该研究为图像生成技术向更高阶人类审美偏好演进提供了理论与实证支持。代码见https://github.com/BarretBa/ICTHP。
原文摘要 · Abstract (English)
Contemporary image generation systems have achieved high fidelity and superior aesthetic quality beyond basic text-image alignment. However, existing evaluation frameworks have failed to evolve in parallel. This study reveals that human preference reward models fine-tuned based on CLIP and BLIP architectures have inherent flaws: they inappropriately assign low scores to images with rich details and high aesthetic value, creating a significant discrepancy with actual human aesthetic preferences. To address this issue, we design a novel evaluation score, ICT (Image-Contained-Text) score, that achieves and surpasses the objectives of text-image alignment by assessing the degree to which images represent textual content. Building upon this foundation, we further train an HP (High-Preference) score model using solely the image modality to enhance image aesthetics and detail quality while maintaining text-image alignment. Experiments demonstrate that the proposed evaluation model improves scoring accuracy by over 10\% compared to existing methods, and achieves significant results in optimizing state-of-the-art text-to-image models. This research provides theoretical and empirical support for evolving image generation technology toward higher-order human aesthetic preferences. Code is available at https://github.com/BarretBa/ICTHP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。