用AI自动生成更优图像提示,无需人工标注。
Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
- 用大视觉语言模型自动重写提示词
- 自动生成反馈评分,实现自我优化
- 适合想提升图像生成质量的用户
文本到图像模型能根据文本提示生成高质量图像,但编写有效提示需专业词汇。现有方法依赖大量人工标注数据和训练好的美学评估模型,存在数据依赖和偏差问题。本文提出一种新型提示优化框架,利用大视觉语言模型(LVLM)作为解题器重写用户简单提示,并同时作为奖励模型,评估生成图像的美学与对齐度。通过利用LVLM的先验知识提供AI反馈,替代人工标注,实现端到端自反馈。解题器与奖励模型统一于同一模型中,通过强化学习迭代优化,实现自我改进。在两个主流数据集上的实验表明,该方法优于多个强基线模型。
原文摘要 · Abstract (English)
Text-to-image models are powerful for producing high-quality images based on given text prompts, but crafting these prompts often requires specialized vocabulary. To address this, existing methods train rewriting models with supervision from large amounts of manually annotated data and trained aesthetic assessment models. To alleviate the dependence on data scale for model training and the biases introduced by trained models, we propose a novel prompt optimization framework, designed to rephrase a simple user prompt into a sophisticated prompt to a text-to-image model. Specifically, we employ the large vision language models (LVLMs) as the solver to rewrite the user prompt, and concurrently, employ LVLMs as a reward model to score the aesthetics and alignment of the images generated by the optimized prompt. Instead of laborious human feedback, we exploit the prior knowledge of the LVLM to provide rewards, i.e., AI feedback. Simultaneously, the solver and the reward model are unified into one model and iterated in reinforcement learning to achieve self-improvement by giving a solution and judging itself. Results on two popular datasets demonstrate that our method outperforms other strong competitors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。