通过多轮反馈优化扩散模型,让AI更懂用户意图。
OMR-Diffusion:Optimizing Multi-Round Enhanced Training in Diffusion Models for Improved Intent Understanding
- 引入人类反馈与多奖励机制,用LoRA微调扩散模型
- 人评胜出508次,对话效率达3.4轮(DALL-E 3为13.7轮)
- 适合需要精准理解用户需求的对话式图像生成场景
生成式AI在文本驱动的图像生成上已取得显著进展,但在多轮对话中仍难以持续对齐用户不断变化的偏好与意图。本文提出视觉协同适应(VCA)框架,结合人工反馈,利用专门设计的奖励模型紧密贴近人类偏好。基于多样化的多轮对话数据集,该框架采用多样性、一致性及偏好反馈等多重奖励函数,通过LoRA对扩散模型进行优化,有效提升图像生成与用户输入的一致性。研究还构建了包含提示与图像对的多轮对话数据集,以更好匹配用户意图。实验表明,该模型在人类评估中获得508胜,优于DALL-E 3的463胜;对话效率达3.4轮(对比DALL-E 3的13.7轮),并在LPIPS(0.15)和BLIP(0.59)等指标上表现优异。多项实验验证了该方法在图像一致性和意图对齐上的显著优势。
原文摘要 · Abstract (English)
Generative AI has significantly advanced text-driven image generation, but it still faces challenges in producing outputs that consistently align with evolving user preferences and intents, particularly in multi-turn dialogue scenarios. In this research, We present a Visual Co-Adaptation (VCA) framework that incorporates human-in-the-loop feedback, utilizing a well-trained reward model specifically designed to closely align with human preferences. Using a diverse multi-turn dialogue dataset, the framework applies multiple reward functions (such as diversity, consistency, and preference feedback) to refine the diffusion model through LoRA, effectively optimizing image generation based on user input. We also constructed multi-round dialogue datasets with prompts and image pairs that well-fit user intent. Experiments show the model achieves 508 wins in human evaluation, outperforming DALL-E 3 (463 wins) and others. It also achieves 3.4 rounds in dialogue efficiency (vs. 13.7 for DALL-E 3) and excels in metrics like LPIPS (0.15) and BLIP (0.59). Various experiments demonstrate the effectiveness of the proposed method over state-of-the-art baselines, with significant improvements in image consistency and alignment with user intent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。