arXiv:2504.17826cs.CVcs.AI2025-04被引 5

基于视觉语言模型的多轮时尚助手,支持个性化穿搭推荐与虚拟试穿。

FashionM3: Multimodal, Multitask, and Multiround Fashion Assistant based on Unified Vision-Language Model

  • 在33万条多模态对话数据上微调,实现多轮交互式推荐
  • 支持个性化推荐、替代方案生成和虚拟试穿,提升穿搭满意度
  • 适合电商、时尚导购场景,实测效果优于传统方法

时尚搭配与个性化推荐在现代零售中至关重要,为时尚产业带来巨大经济价值。随着视觉语言模型(VLM)的发展,通过自然语言与视觉交互提升零售体验成为可能。本文提出FashionM3,一个基于专用微调VLM的多模态、多任务、多轮次时尚助手,可帮助用户发现满意穿搭,具备个性化推荐、替代方案建议、产品图像生成和虚拟试穿模拟等能力。模型在包含331,124条多模态对话样本的新颖FashionRec数据集上训练,覆盖基础、个性化及替代推荐任务,通过多轮交互实现上下文感知的个性化建议迭代优化。定量分析、定性评估及用户研究均表明,FashionM3在推荐有效性和实际应用价值方面表现卓越。

原文摘要 · Abstract (English)

Fashion styling and personalized recommendations are pivotal in modern retail, contributing substantial economic value in the fashion industry. With the advent of vision-language models (VLM), new opportunities have emerged to enhance retailing through natural language and visual interactions. This work proposes FashionM3, a multimodal, multitask, and multiround fashion assistant, built upon a VLM fine-tuned for fashion-specific tasks. It helps users discover satisfying outfits by offering multiple capabilities including personalized recommendation, alternative suggestion, product image generation, and virtual try-on simulation. Fine-tuned on the novel FashionRec dataset, comprising 331,124 multimodal dialogue samples across basic, personalized, and alternative recommendation tasks, FashionM3 delivers contextually personalized suggestions with iterative refinement through multiround interactions. Quantitative and qualitative evaluations, alongside user studies, demonstrate FashionM3's superior performance in recommendation effectiveness and practical value as a fashion assistant.

时尚推荐多模态虚拟试穿对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。