通过多轮对话优化扩散模型,让生成图像更贴合用户偏好。
Optimizing Multi-Round Enhanced Training in Diffusion Models for Improved Preference Understanding
- 引入人类反馈的视觉协同优化框架,结合多个奖励函数。
- 在多轮对话数据上微调扩散模型,提升图像一致性与意图对齐度。
- 适合需要精细交互的图像生成场景,如设计、创作类应用。
生成式AI已推动文本驱动图像生成变革,但在高分辨率输出与细粒度用户偏好对齐方面仍存挑战。为此,多轮交互成为必要手段。以往方法通过奖励反馈增强提示,但未在多轮对话数据集上进行优化。本文提出视觉协同适应(VCA)框架,融合人工反馈,利用与人类偏好对齐的预训练奖励模型。基于多样化的多轮对话数据集,该框架采用多样性、一致性和偏好反馈等多重奖励函数,通过LoRA对扩散模型进行微调,从而依据用户输入优化图像生成。同时,构建了与用户意图对齐的多轮对话提示-图像配对数据集。实验表明,本方法显著优于当前最优基线,在图像一致性和用户意图对齐方面均有提升,尤其在多轮对话场景中用户满意度更高。
原文摘要 · Abstract (English)
Generative AI has significantly changed industries by enabling text-driven image generation, yet challenges remain in achieving high-resolution outputs that align with fine-grained user preferences. Consequently, multi-round interactions are necessary to ensure the generated images meet expectations. Previous methods enhanced prompts via reward feedback but did not optimize over a multi-round dialogue dataset. In this work, we present a Visual Co-Adaptation (VCA) framework incorporating human-in-the-loop feedback, leveraging a well-trained reward model aligned with human preferences. Using a diverse multi-turn dialogue dataset, our framework applies multiple reward functions, such as diversity, consistency, and preference feedback, while fine-tuning the diffusion model through LoRA, thus optimizing image generation based on user input. We also construct multi-round dialogue datasets of prompts and image pairs aligned with user intent. Experiments demonstrate that our method outperforms state-of-the-art baselines, significantly improving image consistency and alignment with user intent. Our approach consistently surpasses competing models in user satisfaction, especially in multi-turn dialogue scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。