arXiv:2507.22627cs.CVcs.AI2025-07ICCV被引 10

用草图+文字对生成时尚图像,支持精细设计定制。

LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing

  • 通过草图与文字配对编码,保留局部特征并融合全局信息
  • 在扩散模型中分步融合条件,生成效果优于现有方法
  • 新数据集支持多对一配对,适合个性化时尚生成研究

时尚设计融合视觉与文本表达:草图定义结构与元素,文字描述材质、质感与风格。本文提出基于草图-文字配对的时尚图像生成方法LOTS,通过全局描述与局部草图+文字对联合建模,引入分步融合策略改进扩散模型。首先,模块化配对中心表示将草图与文字映射至共享隐空间并保持局部特征独立;随后,在扩散模型多步去噪过程中,通过注意力机制实现局部与全局条件融合。为验证方法,基于Fashionpedia构建了首个每张图像提供多个文本-草图对的数据集Sketchy。定量结果显示LOTS在全局与局部指标上均达领先水平,定性分析与人工评估也证实其具备前所未有的设计定制能力。

原文摘要 · Abstract (English)

Fashion design is a complex creative process that blends visual and textual expressions. Designers convey ideas through sketches, which define spatial structure and design elements, and textual descriptions, capturing material, texture, and stylistic details. In this paper, we present LOcalized Text and Sketch for fashion image generation (LOTS), an approach for compositional sketch-text based generation of complete fashion outlooks. LOTS leverages a global description with paired localized sketch + text information for conditioning and introduces a novel step-based merging strategy for diffusion adaptation. First, a Modularized Pair-Centric representation encodes sketches and text into a shared latent space while preserving independent localized features; then, a Diffusion Pair Guidance phase integrates both local and global conditioning via attention-based guidance within the diffusion model's multi-step denoising process. To validate our method, we build on Fashionpedia to release Sketchy, the first fashion dataset where multiple text-sketch pairs are provided per image. Quantitative results show LOTS achieves state-of-the-art image generation performance on both global and localized metrics, while qualitative examples and a human evaluation study highlight its unprecedented level of design customization.

时尚生成草图生成文本控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。