arXiv:2603.19053cs.CVcs.GR2026-03

用几何图像统一生成服装,速度提升至30秒内完成。

SwiftTailor: Efficient 3D Garment Generation with Geometry Image Representation

  • 将服装设计与3D建模融合为两阶段流程,用几何图像压缩表示。
  • 在多模态数据集上实现顶尖视觉质量,推理时间缩短至30秒以内。
  • 适合数字时尚、虚拟试衣等需要快速生成高保真3D服装的场景。

真实且高效的3D服装生成仍是计算机视觉与数字时尚领域的长期挑战。现有方法通常依赖大型视觉-语言模型生成2D缝制图的序列化表示,再通过如GarmentCode等服装建模框架转换为可仿真的3D网格,但推理时间常达30秒至1分钟。本文提出SwiftTailor,一种新型两阶段框架,通过紧凑的几何图像表示统一缝制图推理与基于几何的网格生成。该框架包含两个轻量模块:PatternMaker(高效视觉-语言模型,从多种输入模态预测缝制图)与GarmentSewer(高效密集预测变压器,将缝制图转为新型服装几何图像,以统一UV空间编码所有面板的3D表面)。最终3D网格通过高效逆映射过程重建,结合重网格化与动态缝合算法直接组装服装,从而摊销物理仿真开销。在Multimodal GarmentCodeData上的大量实验表明,SwiftTailor在准确率与视觉保真度上达到最先进水平,同时显著降低推理时间。本工作为下一代3D服装生成提供了可扩展、可解释且高性能的解决方案。

原文摘要 · Abstract (English)

Realistic and efficient 3D garment generation remains a longstanding challenge in computer vision and digital fashion. Existing methods typically rely on large vision- language models to produce serialized representations of 2D sewing patterns, which are then transformed into simulation-ready 3D meshes using garment modeling framework such as GarmentCode. Although these approaches yield high-quality results, they often suffer from slow inference times, ranging from 30 seconds to a minute. In this work, we introduce SwiftTailor, a novel two-stage framework that unifies sewing-pattern reasoning and geometry-based mesh synthesis through a compact geometry image representation. SwiftTailor comprises two lightweight modules: PatternMaker, an efficient vision-language model that predicts sewing patterns from diverse input modalities, and GarmentSewer, an efficient dense prediction transformer that converts these patterns into a novel Garment Geometry Image, encoding the 3D surface of all garment panels in a unified UV space. The final 3D mesh is reconstructed through an efficient inverse mapping process that incorporates remeshing and dynamic stitching algorithms to directly assemble the garment, thereby amortizing the cost of physical simulation. Extensive experiments on the Multimodal GarmentCodeData demonstrate that SwiftTailor achieves state-of-the-art accuracy and visual fidelity while significantly reducing inference time. This work offers a scalable, interpretable, and high-performance solution for next-generation 3D garment generation.

3D服装生成几何图像高效建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。