用自然语言描述服装,零训练生成可缝制的裁剪图
NGL: Natural Garment Language for Training-Free Sewing Pattern Estimation
- 用自然语言构建服装描述框架,直接映射为裁剪图
- 在3个数据集上超越现有方法,多层服装也能准确还原
- 无需训练,适合真实场景中复杂服装的快速建模
从图像估计裁剪图是生成高质量3D服装的有效方法,但受限于真实图像与裁剪图配对数据稀缺。现有方法通过训练视觉-语言模型(VLM)从参数化服装模型生成的合成数据中学习低级裁剪表示,但在真实图像上泛化能力差,难以捕捉服装部件间真实关联,且仅限单层服饰。我们发现VLM擅长用自然语言描述服装,但将描述转化为有效裁剪图仍具挑战。为此,提出NGL(Natural Garment Language),一种与VLM自然描述能力对齐的领域专用语言,实现完全免训练的流水线:通过查询大VLM提取结构化服装规格,并确定性转换为有效裁剪图。在Dress4D、CloSe及新收集的253张真实时尚图像数据集上评估,本方法在标准几何指标上达到最先进水平,且在人类和GPT-based感知评估中更受青睐。此外,NGL可恢复多层服饰,而现有方法多局限于单层,证明其对真实图像(包括遮挡部分)具有强泛化能力。结果表明,高效服装表示对基于VLM的裁剪图估计至关重要。代码与数据将公开用于研究。
原文摘要 · Abstract (English)
Estimating sewing patterns from images is a practical approach for creating high-quality 3D garments, but it remains challenging due to the scarcity of paired real-world image and sewing-pattern data. Existing methods address this limitation by training vision-language models (VLMs) to learn low-level sewing-pattern representations from synthetic garments sampled from parametric garment models. However, they often struggle to generalize to in-the-wild images, fail to capture real-world correlations between garment parts, and are restricted to single-layer outfits. In contrast, we observe that VLMs are effective at describing garments in natural language, but mapping these descriptions into valid sewing patterns remains difficult. To bridge this gap, we propose NGL (Natural Garment Language), a novel domain-specific language that represents garments in terms aligned with VLMs' natural descriptive abilities. Leveraging NGL, we introduce a fully training-free pipeline that queries large VLMs to extract structured garment specifications and deterministically converts them into valid sewing patterns. We evaluate our method on the Dress4D, CloSe and a newly collected dataset of 253 in-the-wild fashion images. Our approach achieves state-of-the-art performance on standard geometry metrics and is preferred in both human and GPT-based perceptual evaluations compared to existing baselines. Furthermore, NGL recovers multi-layer outfits whereas competing methods focus mostly on single-layer garments, highlighting its strong generalization to real-world images even with occluded parts. These results demonstrate that an efficient garment representation is critical for sewing pattern estimation with VLMs. Our code and data will be released for research use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。