用扩散模型生成更时尚的穿搭图像,自动优化不依赖人工提示。
Fashionability-Enhancing Outfit Image Editing with Conditional Diffusion Models
- 基于扩散模型,自动提升图像时尚度并保留原身材特征。
- 相比基线Fashion++,生成图像时尚评分显著更高。
- 通过专家评分与多维度对比实现无提示的时尚优化。
时尚图像生成研究多关注保持身体特征或遵循输入提示,却较少关注提升输出图像的内在时尚度。本文提出一种基于扩散模型的新方法,在保持关键属性可控的前提下,提升生成图像的时尚性。方法包含三方面:1)时尚度增强,确保生成图像比输入更时尚;2)身体特征保留,维持原图的体型与比例;3)自动时尚优化,无需人工输入或外部提示。我们采用两种数据收集方法:利用多个时尚专家通过OpenSkill框架和五项关键维度的成对比较对穿搭图像进行时尚度评分,为生成与评估提供互补视角。实验结果表明,该方法在时尚度上优于基线模型Fashion++,能有效生成更具吸引力与时尚感的图像。
原文摘要 · Abstract (English)
Image generation in the fashion domain has predominantly focused on preserving body characteristics or following input prompts, but little attention has been paid to improving the inherent fashionability of the output images. This paper presents a novel diffusion model-based approach that generates fashion images with improved fashionability while maintaining control over key attributes. Key components of our method include: 1) fashionability enhancement, which ensures that the generated images are more fashionable than the input; 2) preservation of body characteristics, encouraging the generated images to maintain the original shape and proportions of the input; and 3) automatic fashion optimization, which does not rely on manual input or external prompts. We also employ two methods to collect training data for guidance while generating and evaluating the images. In particular, we rate outfit images using fashionability scores annotated by multiple fashion experts through OpenSkill-based and five critical aspect-based pairwise comparisons. These methods provide complementary perspectives for assessing and improving the fashionability of the generated images. The experimental results show that our approach outperforms the baseline Fashion++ in generating images with superior fashionability, demonstrating its effectiveness in producing more stylish and appealing fashion images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。