arXiv:2504.10329cs.CV2025-04被引 1

用自动化数据生成与交叉验证,让文生图模型更懂人类指令。

InstructEngine: Instruction-driven Text-to-Image Alignment

  • 基于文本生成任务分类体系,自动生成2.5万组图文偏好数据。
  • 在DrawBench上使SD v1.5和SDXL性能分别提升10.53%和5.30%。
  • 适合关注文生图对齐、低成本训练的开发者与研究者。

强化学习从人类/人工智能反馈(RLHF/RLAIF)已被广泛用于文生图模型的偏好对齐。现有方法在数据和算法层面均存在局限:数据方面,多数依赖人工标注的偏好数据,成本高昂且奖励模型增加计算开销,准确性难保障;算法层面,多数方法仅利用图像反馈作为比较信号,忽视文本信息,导致效率低、信号稀疏。为此,我们提出InstructEngine框架。在数据方面,首先构建文生图生成任务分类体系,再基于大模型与人工规则设计自动化数据生成流程,生成25,000组文本-图像偏好对;最后引入交叉验证对齐方法,通过组织语义相似样本形成可比对对,提升数据效率。在DrawBench上的评估表明,InstructEngine使SD v1.5和SDXL性能分别提升10.53%和5.30%,优于当前最优基线。消融实验证明其各组件有效性。人工评测胜率超50%,证明其更符合人类偏好。

原文摘要 · Abstract (English)

Reinforcement Learning from Human/AI Feedback (RLHF/RLAIF) has been extensively utilized for preference alignment of text-to-image models. Existing methods face certain limitations in terms of both data and algorithm. For training data, most approaches rely on manual annotated preference data, either by directly fine-tuning the generators or by training reward models to provide training signals. However, the high annotation cost makes them difficult to scale up, the reward model consumes extra computation and cannot guarantee accuracy. From an algorithmic perspective, most methods neglect the value of text and only take the image feedback as a comparative signal, which is inefficient and sparse. To alleviate these drawbacks, we propose the InstructEngine framework. Regarding annotation cost, we first construct a taxonomy for text-to-image generation, then develop an automated data construction pipeline based on it. Leveraging advanced large multimodal models and human-defined rules, we generate 25K text-image preference pairs. Finally, we introduce cross-validation alignment method, which refines data efficiency by organizing semantically analogous samples into mutually comparable pairs. Evaluations on DrawBench demonstrate that InstructEngine improves SD v1.5 and SDXL's performance by 10.53% and 5.30%, outperforming state-of-the-art baselines, with ablation study confirming the benefits of InstructEngine's all components. A win rate of over 50% in human reviews also proves that InstructEngine better aligns with human preferences.

文生图偏好对齐自动化数据RLHF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。