arXiv:2411.17017cs.CV2024-11被引 9

用Transformer扩散模型提升虚拟试穿,让衣服文字不扭曲、细节更真实。

TED-VITON: Transformer-Empowered Diffusion Models for Virtual Try-On

  • 引入服装语义适配器增强衣物特征,提升生成精度。
  • 提出文本保真损失,实现无扭曲的服装文字渲染。
  • 结合大模型优化提示词,适合追求高真实感的视觉生成研究者。

虚拟试穿(VTO)近年在生成逼真图像和保留服装细节方面取得显著进展,主要得益于文本到图像(T2I)扩散模型的强大生成能力。然而,支撑这些方法的T2I模型已趋于陈旧,限制了进一步提升空间。现有方法在准确呈现服装文字(如品牌标识)时仍存在失真问题,且难以保持纹理与材质的精细度。基于扩散变压器(DiT)的新型T2I模型展现出卓越性能,为推动VTO发展带来契机。但直接将现有VTO技术应用于此类模型无效,因其架构差异阻碍了对高级生成能力的充分挖掘。为此,我们提出TED-VITON框架,包含:1)服装语义适配器(GS Adapter),用于增强特定衣物特征;2)文本保真损失(Text Preservation Loss),确保文字渲染准确无畸变;3)基于大语言模型(LLM)的约束机制,优化生成提示词。该框架在视觉质量与文字保真度上达到当前最优水平,为VTO任务设立了新基准。

原文摘要 · Abstract (English)

Recent advancements in Virtual Try-On (VTO) have demonstrated exceptional efficacy in generating realistic images and preserving garment details, largely attributed to the robust generative capabilities of text-to-image (T2I) diffusion backbones. However, the T2I models that underpin these methods have become outdated, thereby limiting the potential for further improvement in VTO. Additionally, current methods face notable challenges in accurately rendering text on garments without distortion and preserving fine-grained details, such as textures and material fidelity. The emergence of Diffusion Transformer (DiT) based T2I models has showcased impressive performance and offers a promising opportunity for advancing VTO. Directly applying existing VTO techniques to transformer-based T2I models is ineffective due to substantial architectural differences, which hinder their ability to fully leverage the models' advanced capabilities for improved text generation. To address these challenges and unlock the full potential of DiT-based T2I models for VTO, we propose TED-VITON, a novel framework that integrates a Garment Semantic (GS) Adapter for enhancing garment-specific features, a Text Preservation Loss to ensure accurate and distortion-free text rendering, and a constraint mechanism to generate prompts by optimizing Large Language Model (LLM). These innovations enable state-of-the-art (SOTA) performance in visual quality and text fidelity, establishing a new benchmark for VTO task. Project page: https://zhenchenwan.github.io/TED-VITON/

虚拟试穿扩散模型Transformer文本保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。