arXiv:2506.05210cs.CV2025-06

用文字和图片生成服装,让大模型懂时尚设计

Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation

  • 构建视觉-语言-服装三模态模型,从文本和图像生成服饰
  • 零样本测试显示能泛化到未见服装风格与提示
  • 适合研究时尚生成、跨模态理解的开发者与设计师

多模态基础模型展现出强大泛化能力,但其在服装生成等专业领域知识迁移的研究仍不充分。本文提出VLG模型,能够根据文本描述和视觉图像合成服装。通过零样本实验评估其将网络规模推理能力迁移到未见过的服装风格和提示的能力。初步结果表明该模型具备良好的迁移潜力,验证了多模态基础模型在时尚设计等专业化领域有效适配的可能性。

原文摘要 · Abstract (English)

Multimodal foundation models have demonstrated strong generalization, yet their ability to transfer knowledge to specialized domains such as garment generation remains underexplored. We introduce VLG, a vision-language-garment model that synthesizes garments from textual descriptions and visual imagery. Our experiments assess VLG's zero-shot generalization, investigating its ability to transfer web-scale reasoning to unseen garment styles and prompts. Preliminary results indicate promising transfer capabilities, highlighting the potential for multimodal foundation models to adapt effectively to specialized domains like fashion design.

服装生成多模态零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。