多奖励预训练提升文生图质量与效率
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
- 训练时同时纳入多个奖励信号,让模型直接学习用户偏好
- 在GenEval基准上达到最先进水平,生成图像更符合用户喜好
- 减少无效数据浪费,训练更快,兼顾多样性与语义准确
主流文生图模型通常采用后训练阶段筛选图像,并用单一奖励模型(如用户偏好)进行微调,导致信息丢失且仅优化单一目标,影响多样性、语义保真度和训练效率。为此,本文提出MIRO方法,在预训练阶段让模型同时接受多重奖励条件,使模型直接学习用户偏好。该方法不仅提升了生成图像的视觉质量,还显著加快了训练速度,在GenEval组合性生成基准以及用户偏好评分(PickAScore、ImageReward、HPSv2)上均达到当前最优表现。
原文摘要 · Abstract (English)
The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typically user preference. This discards informative data as well as optimizes only for a single reward, hence harming diversity, semantic fidelity and efficiency. Instead, we propose MIRO, a method that conditions the model on multiple rewards during training, thus letting the model learn user preferences directly. MIRO pre-training both improves the visual quality of the generated images and speeds up the training, achieving state of the art on the GenEval compositional benchmark and user-preference scores (PickAScore, ImageReward, HPSv2).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。