从单张图生成可仿真服装,无需多视角或迭代优化
Image2Garment: Simulation-ready Garment Generation from a Single Image
- 用视觉语言模型从图片推断面料成分和属性
- 仅需少量物理测量数据,精准预测仿真所需的材质参数
- 适合虚拟试衣、数字人服装模拟等场景
从单张图像生成物理上准确且可仿真的服装极具挑战,原因在于缺乏图像到物理的标注数据集,且问题本身病态。以往方法要么需要多视角拍摄与昂贵的可微分仿真,要么仅预测服装几何而忽略仿真所需的材质属性。我们提出一种前馈框架:首先微调视觉语言模型,从真实图像中推断材料组成与织物属性;随后训练轻量级预测器,利用小规模材料-物理测量数据集将这些属性映射为对应的物理参数。本工作构建了两个新数据集(FTAG 和 T2P),实现无需迭代优化的单图生成仿真可用服装。实验表明,该方法在材料组成估计和织物属性预测上优于现有技术,结合物理参数预测后,仿真精度显著超越当前最优图像到服装方法。
原文摘要 · Abstract (English)
Estimating physically accurate, simulation-ready garments from a single image is challenging due to the absence of image-to-physics datasets and the ill-posed nature of this problem. Prior methods either require multi-view capture and expensive differentiable simulation or predict only garment geometry without the material properties required for realistic simulation. We propose a feed-forward framework that sidesteps these limitations by first fine-tuning a vision-language model to infer material composition and fabric attributes from real images, and then training a lightweight predictor that maps these attributes to the corresponding physical fabric parameters using a small dataset of material-physics measurements. Our approach introduces two new datasets (FTAG and T2P) and delivers simulation-ready garments from a single image without iterative optimization. Experiments show that our estimator achieves superior accuracy in material composition estimation and fabric attribute prediction, and by passing them through our physics parameter estimator, we further achieve higher-fidelity simulations compared to state-of-the-art image-to-garment methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。