HPSv3++让奖励模型适应不同图像生成能力与迭代阶段,提升人类偏好预测效果。
HPSv3++: Scaling Reward Models Across the Full Spectrum of Diffusion Model Capabilities

- 构建21.2万条双维度偏好数据集HPDv3++,结合高能力模型与人工标注。
- 两阶段训练:保留原知识并融合多样审美,用无标签数据增强跨能力-迭代泛化。
- 在多个评测中优于HPSv3,适合用于不同阶段的文本到图像强化学习。
奖励模型引导文本到图像(T2I)系统生成符合人类偏好的结果。然而,现有奖励模型如HPSv3基于早期T2I模型的预标注数据训练,未考虑模型能力进化与强化学习迭代带来的质量判别变化,限制了其广泛适用性。本文提出HPSv3++,一个可覆盖全能力-迭代谱的奖励模型框架。首先,构建212K条双维度偏好数据集HPDv3++,使用高能力模型Qwen-Image结合人工监督标注文本保真度与美学质量。其次,提出两阶段训练:第一阶段采用数据感知正交梯度投影,融合HPDv3++中的多样化审美感知,同时保留HPSv3中有效的人类偏好知识;第二阶段利用涵盖不同能力水平与强化学习迭代阶段的T2I模型无标签数据,引入联合能力-迭代条件信号及标准差驱动的无监督引导机制,增强奖励模型在全谱系上的表现。HPSv3++在HPDv3上比HPSv3提升9.8%,在GenAI-Bench上提升5.5%,在自建的HPDv3++上达到79.1%/88.1%。用于T2I强化学习时,持续提升GenEval得分,验证其广泛适用性。代码已开源。
原文摘要 · Abstract (English)
Reward models guide text-to-image (T2I) systems toward outputs aligned with human preferences. However, typical reward models such as HPSv3 are trained on pre-annotated data from earlier T2I models, without accounting for quality discriminative shifts arising from evolving model capabilities and reinforcement learning (RL) iterations, limiting their broader applicability. In this work, we propose HPSv3++, a reward model framework that elevates the HPSv3 model for varying T2I model capabilities and their RL iteration changes across the full capability-iteration spectrum. Specifically, we first introduce HPDv3++, a 212K dual-dimension preference dataset annotated for text fidelity and aesthetic quality using a recent high-capability (Qwen-Image) model with human supervision. We then propose a two-stage training framework. Stage 1 employs data-aware orthogonal gradient projection to incorporate diverse aesthetic perception from HPDv3++ while preserving the original effective human preference knowledge in HPSv3. Stage 2 further leverages unlabeled data from T2I models spanning different capability levels and RL iterations, and introduces a joint capability-iterations conditioned signal for the reward model together with a standard deviation-driven unsupervised guidance mechanism, strengthening reward model across the capability-iteration spectrum. HPSv3++ achieves state-of-the-art preference prediction, outperforming HPSv3 9.8% on HPDv3, 5.5% on GenAI-Bench, while achieving 79.1%/88.1% on our proposed HPDv3++. When used for T2I RL training, it consistently improves GenEval scores across diverse T2I models, demonstrating its wide-range capabilities. The code is available at https://github.com/PlantPotatoOnMoon/HPSv3-PlusPlus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。