让图像生成更懂用户需求,动态评估生成质量。
DyCoRM: Dynamic Criterion-Aware Reward Modeling for Text-to-Image Generation

- 根据用户具体要求动态调整评估标准
- 构建2万条带细粒度标注的数据集
- 适合需要精准控制生成效果的研究者
随着文本到图像(T2I)生成技术的不断进步,高质量图像的生成已变得越来越容易;因此,用户需求正转向更符合特定要求的图像。由于奖励模型在评估生成图像是否符合用户偏好方面的作用日益重要,这一趋势给奖励建模带来了新挑战:奖励模型不应仅依赖静态通用评价维度,而应考虑任务相关且细粒度的评估标准。为此,我们提出DyCoRM,一种动态、准则感知的奖励模型,能够基于任务相关准则进行评估并执行准则感知的偏好比较。为支持该设定,我们构建了DyCoDataset-20K,包含动态准则及准则级标注,并进一步衍生出DyCoBench-1K,用于系统评估奖励模型在动态准则下的表现。我们还引入DyCoPick,将准则感知奖励建模应用于T2I图像选择。本研究首次建立了面向动态与细粒度评估的奖励建模范式及其在T2I生成中的实际应用。
原文摘要 · Abstract (English)
With the continued advancement of text-to-image (T2I) generation, producing high-quality images is becoming increasingly attainable; consequently, user demands are shifting toward images that better satisfy their specific requirements. As reward models play an increasingly important role in assessing whether generated images align with user preference, this trend introduces an important challenge for reward modeling: rather than relying solely on static and general evaluation dimensions, reward models should account for the task-relevant and fine-grained criteria through which users assess whether generated images meet their specific requirements. To address this challenge, we propose DyCoRM, a dynamic, criterion-aware reward model that grounds task-relevant criteria and performs criterion-aware preference comparison. To support this setting, we construct DyCoDataset-20K, which provides dynamic criteria together with criterion-level annotations, and further derive DyCoBench-1K, a benchmark for systematically evaluating reward models under dynamic criteria. We further introduce DyCoPick, which applies criterion-aware reward modeling to selecting T2I images. Our contributions establish the first reward modeling framework for dynamic and fine-grained evaluation and practical application in T2I generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。