用用户偏好反馈优化生成,让图像更贴近心中所想。
Personalized Image Generation via Human-in-the-loop Bayesian Optimization
- 通过多轮选择式偏好反馈,引导扩散模型生成更符合用户想象的图像。
- 仅需30位用户参与,定量与定性评估均优于5个基线方法。
- 适合需要精细个性化图像生成的用户,尤其在语言描述失效时有效。
设想用户艾丽斯脑海中有一幅具体图像 $x^\ast$,例如她童年居住街道的景象。她通过多次语言提示引导生成模型,得到近似图像 $x^{p*}$,但难以进一步接近目标。本文发现:即使语言提示失效,用户仍能判断新生成的图像 $x^+$ 是否比 $x^{p*}$ 更接近 $x^\ast$。基于此,提出 MultiBO(多选偏好贝叶斯优化)方法,以 $x^{p*}$ 为起点,每次生成 $K$ 张新图像,获取用户偏好反馈,利用反馈优化扩散模型,再生成下一组 $K$ 张图像。实验表明,在 $B$ 轮用户反馈后,可显著逼近 $x^\ast$,且生成模型完全不知晓 $x^\ast$ 的信息。30位用户的定性评分结合5种基线的定量指标对比,验证了该方法的有效性,证明多选偏好反馈可用于高效个性化图像生成。
原文摘要 · Abstract (English)
Imagine Alice has a specific image $x^\ast$ in her mind, say, the view of the street in which she grew up during her childhood. To generate that exact image, she guides a generative model with multiple rounds of prompting and arrives at an image $x^{p*}$. Although $x^{p*}$ is reasonably close to $x^\ast$, Alice finds it difficult to close that gap using language prompts. This paper aims to narrow this gap by observing that even after language has reached its limits, humans can still tell when a new image $x^+$ is closer to $x^\ast$ than $x^{p*}$. Leveraging this observation, we develop MultiBO (Multi-Choice Preferential Bayesian Optimization) that carefully generates $K$ new images as a function of $x^{p*}$, gets preferential feedback from the user, uses the feedback to guide the diffusion model, and ultimately generates a new set of $K$ images. We show that within $B$ rounds of user feedback, it is possible to arrive much closer to $x^\ast$, even though the generative model has no information about $x^\ast$. Qualitative scores from $30$ users, combined with quantitative metrics compared across $5$ baselines, show promising results, suggesting that multi-choice feedback from humans can be effectively harnessed for personalized image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。