让服装生成模型理解自然语言指令,更真实地匹配用户需求。
Grounding Free-Form Instructions for Fashion Complementary Image Generation

- 用自然语言指令替代固定模板,提升生成语义匹配度。
- 在多维度评测中,生成服装与指令一致且风格统一。
- 结构简洁,推理成本低,适合实际应用落地。
时尚搭配图像生成(CIG)旨在根据种子服装和用户意图生成风格匹配的衣物,属于典型的多模态语义对齐问题。现有基准依赖固定模板提示(如“一张裙子的照片”),无法反映真实用户查询,且掩盖了模型在不同语言细节程度下的表现。本文提出自由形式指令下的时尚搭配图像生成任务,通过视觉-语言模型生成低、中、高三个层次具体性的自然语言指令,并经人工验证后扩充三个现有CIG基准数据集。我们基于StyleFlow——一种联合条件于种子图像与指令的单模态变压器架构的修正流匹配模型,实现该任务。在图像质量、商品库对齐分析、消融实验及人类评估中,StyleFlow均稳定生成符合指令且风格一致的服装,同时相比含额外模块的方法显著降低架构复杂度与推理开销。
原文摘要 · Abstract (English)
Fashion complementary image generation (CIG) aims to create garments that stylistically match a seed item based on user intent, making it a natural multimodal grounding problem where models must interpret language in visual context. Existing CIG benchmarks rely on rigid template prompts (e.g., "a photo of a skirt"), failing to reflect natural user queries and obscuring model behavior across levels of linguistic specificity. We introduce fashion complementary image generation with free-form instructions, a multimodal language-grounding setting where a model generates a compatible garment from a seed image and a natural-language instruction. To this end, we enrich three CIG benchmarks with low-, medium-, and high-specificity instructions generated by a vision-language model and validated by human annotators. We instantiate the task with StyleFlow, a Rectified Flow Matching model that jointly conditions on the seed image and instruction within a single multimodal transformer. Across image quality metrics, catalog-alignment analysis, ablations, and human evaluation, StyleFlow consistently produces instruction-aligned and stylistically coherent garments while reducing architectural complexity and inference cost relative to auxiliary-module approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。