用统一模型同时实现虚拟试穿与脱衣,提升细节还原与推理效率。
OmniDiT: Extending Diffusion Transformer to Omni-VTON Framework
- 基于扩散Transformer构建统一框架,整合试穿与脱衣任务。
- 在380k数据集上实现最佳无模型试穿与脱衣性能。
- 引入滑动窗口注意力,线性复杂度加速推理,适合真实场景应用。
尽管虚拟试穿(VTON)和虚拟脱衣(VTOFF)技术快速进展,现有方法仍面临细粒度细节保留差、复杂场景泛化弱、流程繁琐及推理效率低等问题。为此,本文提出OmniDiT,一种基于扩散Transformer的全场景虚拟试穿框架,将试穿与脱衣任务统一建模。首先,构建自演化数据构建流程,生成包含超过380k多样高质量服装-模特-试穿图像对及详细文本提示的大规模数据集Omni-TryOn。其次,采用标记拼接与自适应位置编码,有效融合多参考条件;为缓解长序列计算瓶颈,首次将滑动窗口注意力引入扩散模型,实现线性复杂度。为进一步缓解局部窗口注意力导致的性能下降,采用多时间步预测与对齐损失提升生成保真度。实验表明,在多种复杂场景下,该方法在无模型VTON与VTOFF任务中表现最优,且在基于模型的VTON任务中达到当前最先进水平。
原文摘要 · Abstract (English)
Despite the rapid advancement of Virtual Try-On (VTON) and Try-Off (VTOFF) technologies, existing VTON methods face challenges with fine-grained detail preservation, generalization to complex scenes, complicated pipeline, and efficient inference. To tackle these problems, we propose OmniDiT, an omni Virtual Try-On framework based on the Diffusion Transformer, which combines try-on and try-off tasks into one unified model. Specifically, we first establish a self-evolving data curation pipeline to continuously produce data, and construct a large VTON dataset Omni-TryOn, which contains over 380k diverse and high-quality garment-model-tryon image pairs and detailed text prompts. Then, we employ the token concatenation and design an adaptive position encoding to effectively incorporate multiple reference conditions. To relieve the bottleneck of long sequence computation, we are the first to introduce Shifted Window Attention into the diffusion model, thus achieving a linear complexity. To remedy the performance degradation caused by local window attention, we utilize multiple timestep prediction and an alignment loss to improve generation fidelity. Experiments reveal that, under various complex scenes, our method achieves the best performance in both the model-free VTON and VTOFF tasks and a performance comparable to current SOTA methods in the model-based VTON task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。