用草图+文字配对生成更精准的时装图像,结构更稳、细节更丰富。
Multi-Level Conditioning by Pairing Localized Text and Sketch for Fashion Image Generation
- 将局部草图与文字描述配对,分层编码并融合到扩散模型中
- 在多个测试集上提升结构准确性与细节丰富度,优于当前最佳方法
- 适合时尚设计、跨模态生成领域的研究者与开发者使用
草图为设计师提供了简洁而富有表现力的早期时装构思方式,可明确结构、轮廓与空间关系;文本则补充材质、颜色和风格等细节。有效融合图文模态需在利用文本局部属性指导时保持草图的视觉结构。本文提出LOTS框架,通过全局草图引导与多组局部草图-文本配对实现多层次条件控制。其多级条件阶段在共享潜在空间中独立编码局部特征,同时保持全局结构协调;扩散配对引导阶段在扩散模型的多步去噪过程中,通过注意力机制整合局部与全局条件。为验证方法,我们构建了Sketchy数据集——首个每张图像提供多个文本-草图配对的时装数据集,包含高质量、专业级且结构一致的草图。为进一步评估鲁棒性,还引入“野外”子集,包含非专业人士绘制的草图,具有更高变异性与瑕疵。实验表明,该方法在保持全局结构一致性的同时,增强了局部语义引导能力,显著优于现有最先进方法。相关数据集、平台与代码已公开。
原文摘要 · Abstract (English)
Sketches offer designers a concise yet expressive medium for early-stage fashion ideation by specifying structure, silhouette, and spatial relationships, while textual descriptions complement sketches to convey material, color, and stylistic details. Effectively combining textual and visual modalities requires adherence to the sketch visual structure when leveraging the guidance of localized attributes from text. We present LOcalized Text and Sketch with multi-level guidance (LOTS), a framework that enhances fashion image generation by combining global sketch guidance with multiple localized sketch-text pairs. LOTS employs a Multi-level Conditioning Stage to independently encode local features within a shared latent space while maintaining global structural coordination. Then, the Diffusion Pair Guidance stage integrates both local and global conditioning via attention-based guidance within the diffusion model's multi-step denoising process. To validate our method, we develop Sketchy, the first fashion dataset where multiple text-sketch pairs are provided per image. Sketchy provides high-quality, clean sketches with a professional look and consistent structure. To assess robustness beyond this setting, we also include an "in the wild" split with non-expert sketches, featuring higher variability and imperfections. Experiments demonstrate that our method strengthens global structural adherence while leveraging richer localized semantic guidance, achieving improvement over state-of-the-art. The dataset, platform, and code are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。