arXiv:2511.18957cs.CV2025-11被引 6

打造高分辨率视频试穿数据集,解决细节还原难题。

Eevee: Towards Close-up High-resolution Video-based Virtual Try-on

  • 构建包含近景与全景的高清试穿视频数据集
  • 提出新评估指标VGID,精准衡量服装纹理结构一致性
  • 验证现有模型在细节还原上的不足,适合时尚电商研究者

视频虚拟试穿技术为时尚电商提供低成本制作营销视频的方案。然而其实际应用受限于两大问题:一是现有数据集仅依赖单张服装图像,难以准确捕捉真实纹理细节;二是多数方法仅生成全身视频,无法满足商业对近景细节展示的需求。为此,我们提出一个高分辨率视频虚拟试穿数据集,具备两大特性:一是提供高保真服装图像(含细节近景)和文本描述;二是首次包含真人模特的全身与近景试穿视频。此外,近景视频对服装一致性的要求更高,需精细保留纹理与结构。为此,我们提出新的评估指标VGID(Video Garment Inception Distance),量化纹理与结构的保持程度。实验表明,利用本数据集中的细节图像,现有生成模型可有效提取并融合纹理特征,显著提升试穿结果的真实感与细节保真度。我们还对近期模型进行了全面基准测试,有效识别出当前方法在纹理与结构保留方面的缺陷。

原文摘要 · Abstract (English)

Video virtual try-on technology provides a cost-effective solution for creating marketing videos in fashion e-commerce. However, its practical adoption is hindered by two critical limitations. First, the reliance on a single garment image as input in current virtual try-on datasets limits the accurate capture of realistic texture details. Second, most existing methods focus solely on generating full-shot virtual try-on videos, neglecting the business's demand for videos that also provide detailed close-ups. To address these challenges, we introduce a high-resolution dataset for video-based virtual try-on. This dataset offers two key features. First, it provides more detailed information on the garments, which includes high-fidelity images with detailed close-ups and textual descriptions; Second, it uniquely includes full-shot and close-up try-on videos of real human models. Furthermore, accurately assessing consistency becomes significantly more critical for the close-up videos, which demand high-fidelity preservation of garment details. To facilitate such fine-grained evaluation, we propose a new garment consistency metric VGID (Video Garment Inception Distance) that quantifies the preservation of both texture and structure. Our experiments validate these contributions. We demonstrate that by utilizing the detailed images from our dataset, existing video generation models can extract and incorporate texture features, significantly enhancing the realism and detail fidelity of virtual try-on results. Furthermore, we conduct a comprehensive benchmark of recent models. The benchmark effectively identifies the texture and structural preservation problems among current methods.

视频试穿高分辨率服装一致性数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。