用视觉语言模型生成详细反馈,优化图像生成模型的微调数据。
RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences
- 通过VLM生成图像批评和可执行修改指令,构建高质量偏好对。
- 在扩散模型上微调后,生成图像质量显著提升,细节更准确。
- 适合研究可控图像生成、偏好学习与多模态优化的学者。
传统大模型/视觉生成模型的偏好微调方法通常仅依赖奖励模型标注,存在决策过程不透明、难以理解偏好原因,且易出现奖励黑客或过拟合等问题。本文提出一种新型管道——丰富偏好优化(RPO),利用视觉语言模型(VLM)提供的丰富反馈信号,改进文本到图像扩散模型等视觉生成模型的微调数据构建。流程首先通过提示VLM对合成图像生成详细批评,再进一步提示VLM从中提取可靠且可操作的图像编辑指令。依据这些指令对图像进行优化,形成具有信息量的合成偏好对,作为增强的微调数据集。实验表明该方法在提升主流扩散模型性能方面效果显著。
原文摘要 · Abstract (English)
Traditional preference tuning methods for LLMs/Visual Generative Models often rely solely on reward model labeling, which can be opaque, offer limited insights into the rationale behind preferences, and are prone to issues such as reward hacking or overfitting. We introduce Rich Preference Optimization (RPO), a novel pipeline that leverages rich feedback signals from Vision Language Models (VLMs) to improve the curation of preference pairs for fine-tuning visual generative models like text-to-image diffusion models. Our approach begins with prompting VLMs to generate detailed critiques of synthesized images, from which we further prompt VLMs to extract reliable and actionable image editing instructions. By implementing these instructions, we create refined images, resulting in synthetic, informative preference pairs that serve as enhanced tuning datasets. We demonstrate the effectiveness of our pipeline and the resulting datasets in fine-tuning state-of-the-art diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。