arXiv:2501.05369cs.CV2025-01被引 2

提出单网络方法MNVTON,让虚拟试衣更高效清晰

1-2-1: Renaissance of Single-Network Paradigm for Virtual Try-On

  • 用模态归一化让文本图像视频共享注意力层
  • 图像视频试衣质量优于传统双网络方法
  • 适合需要高分辨率和长时序试衣的电商场景

虚拟试衣(VTON)在电商中日益重要,可真实模拟服装穿在人物上的效果,同时保持原貌与姿态。早期方法依赖单一生成网络,但因特征提取与融合能力有限,难以保留精细服装细节。近期方法改用双网络结构,引入辅助的“ReferenceNet”提升特征处理,虽有效但带来显著计算开销,限制了高分辨率及长时序图像/视频应用的扩展性。本文挑战双网络范式,提出新型单网络VTON方法MNVTON,通过模态特定归一化策略,分别处理文本、图像与视频输入,使三者能在同一注意力层中共享。大量实验表明,该方法在图像与视频试衣任务中持续产出更高质量、更精细的结果。结果表明,单网络范式可媲美双网络性能,为高质量、可扩展的VTON应用提供更高效方案。

原文摘要 · Abstract (English)

Virtual Try-On (VTON) has become a crucial tool in ecommerce, enabling the realistic simulation of garments on individuals while preserving their original appearance and pose. Early VTON methods relied on single generative networks, but challenges remain in preserving fine-grained garment details due to limitations in feature extraction and fusion. To address these issues, recent approaches have adopted a dual-network paradigm, incorporating a complementary "ReferenceNet" to enhance garment feature extraction and fusion. While effective, this dual-network approach introduces significant computational overhead, limiting its scalability for high-resolution and long-duration image/video VTON applications. In this paper, we challenge the dual-network paradigm by proposing a novel single-network VTON method that overcomes the limitations of existing techniques. Our method, namely MNVTON, introduces a Modality-specific Normalization strategy that separately processes text, image and video inputs, enabling them to share the same attention layers in a VTON network. Extensive experimental results demonstrate the effectiveness of our approach, showing that it consistently achieves higher-quality, more detailed results for both image and video VTON tasks. Our results suggest that the single-network paradigm can rival the performance of dualnetwork approaches, offering a more efficient alternative for high-quality, scalable VTON applications.

虚拟试衣单网络生成模型电商应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。