用生成模型的中间表示做图像评分,效果优于现有方法。
DiT-Reward: Generative Representations for Text-to-Image Reward Modeling

- 将预训练扩散Transformer转为奖励模型,利用多层文本条件表示
- 在四个基准上表现超HPSv3,最高达85.6%准确率,且支持冻结主干
- 适用于图像生成优化,推理速度提升1.65倍,适合追求真实感的场景
我们探讨了图像生成模型学习到的表示是否可用于生成图像的评估。为此,提出DiT-Reward,通过处理接近干净的图像潜变量,并聚合跨变压器层的文本条件表示,将预训练的文生图扩散Transformer转化为奖励模型。在与HPSv3相同的训练数据混合下,DiT-Reward在所有四个评估偏好基准上均表现更优,于HPDv2达到85.6%,于HPDv3达到77.6%。当生成主干冻结时,轻量级可学习头仍能从其表示中提取有意义的偏好预测。深度探测显示,中后层表示对下游奖励性能贡献最大,且多阶段融合进一步提升效果。同时观察到奖励性能随生成主干容量持续正向增长。当用于优化Stable Diffusion 3.5 Large(Flow-GRPO)时,DiT-Reward在相同训练轨迹下超越HPSv3,尤其在真实感方面提升显著。直接潜空间评分实现1.65倍推理加速,峰值显存相当。结果表明,预训练的生成型DiT可提供可迁移的表示用于奖励建模与策略优化。
原文摘要 · Abstract (English)
Can representations learned for image generation also support the evaluation of generated images? We study text-to-image reward prediction as a downstream task of generative representation learning. To this end, we introduce DiT-Reward, which converts a pretrained text-to-image Diffusion Transformer into a reward model by processing near-clean image latents and aggregating text-conditioned image representations across transformer layers. Under the same training data mixture as HPSv3, DiT-Reward outperforms HPSv3 on all four evaluated preference benchmarks, reaching 85.6% on HPDv2 and 77.6% on HPDv3. When the generative backbone is frozen, a lightweight learned head can still extract meaningful preference predictions from its representations. Probing across depth further reveals that downstream reward performance is strongest in the middle-to-late layers and benefits from combining representations across different stages. We also observe consistent positive scaling with generative backbone capacity. Finally, when used to optimize Stable Diffusion 3.5 Large with Flow-GRPO, DiT-Reward outperforms HPSv3 along the matched training trajectory, with particularly clear gains in realism. Direct latent scoring also achieves a 1.65x inference speedup over HPSv3 with comparable peak memory. These results show that pretrained generative DiTs provide transferable representations for reward modeling and policy optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。