arXiv:2602.16502cs.CV2026-02被引 2

从一张街拍图直接生成可模拟的服装裁剪图,无需多视角或迭代优化。

DressWild: Feed-Forward Pose-Agnostic Garment Sewing Pattern Generation from In-the-Wild Images

  • 用视觉语言模型统一图像姿态,提取3D感知的服装特征。
  • 通过Transformer融合特征,直接预测可物理模拟的裁剪参数。
  • 适合需要快速生成可编辑服装模型的虚拟试穿与动画场景。

近期服装裁剪图生成技术取得进展,但现有前馈方法在不同姿态和视角下表现不佳,而基于优化的方法计算成本高且难以扩展。本文聚焦于需可编辑、可分离、可模拟的服装建模与制作应用,提出DressWild——一种新颖的前馈管道,仅需单张街拍图像即可重建符合物理规律的2D裁剪图及对应3D服装。输入图像后,该方法利用视觉语言模型(VLMs)在图像层面归一化姿态差异,提取具有姿态感知与3D信息的服装特征,经由Transformer编码器融合后,直接预测裁剪图参数,可直接用于物理模拟、纹理生成与多层虚拟试穿。大量实验表明,本方法无需多视角输入或迭代优化,即可鲁棒地从真实场景图像中恢复多样化的裁剪图与3D服装,为真实服装模拟与动画提供高效可扩展的解决方案。

原文摘要 · Abstract (English)

Recent advances in garment pattern generation have shown promising progress. However, existing feed-forward methods struggle with diverse poses and viewpoints, while optimization-based approaches are computationally expensive and difficult to scale. This paper focuses on sewing pattern generation for garment modeling and fabrication applications that demand editable, separable, and simulation-ready garments. We propose DressWild, a novel feed-forward pipeline that reconstructs physics-consistent 2D sewing patterns and the corresponding 3D garments from a single in-the-wild image. Given an input image, our method leverages vision-language models (VLMs) to normalize pose variations at the image level, then extract pose-aware, 3D-informed garment features. These features are fused through a transformer-based encoder and subsequently used to predict sewing pattern parameters, which can be directly applied to physical simulation, texture synthesis, and multi-layer virtual try-on. Extensive experiments demonstrate that our approach robustly recovers diverse sewing patterns and the corresponding 3D garments from in-the-wild images without requiring multi-view inputs or iterative optimization, offering an efficient and scalable solution for realistic garment simulation and animation.

服装生成视觉语言模型3D建模前馈网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。