无需训练,用大模型实现更懂审美的穿搭推荐
TATTOO: Training-free AesTheTic-aware Outfit recOmmendation
- 用多模态大模型生成服饰描述,再通过审美思维链提取风格特征
- 在Aesthetic-100数据集上超越已有训练方法,零样本检索能力更强
- 适合想快速部署、追求美学一致性的电商推荐场景
全球时尚电商市场严重依赖智能且符合审美的穿搭补全工具来提升销量。以往研究虽已探索穿搭补全与搭配商品检索,但大多需在大规模标注数据上进行昂贵的任务专用训练,且未显式引入人类审美。在多模态大语言模型(MLLMs)时代,我们证明传统训练范式可简化为无需训练的新模式,不仅推荐得分更高,审美感知也更优。提出TATTOO:首先用MLLM生成目标商品描述,再通过审美思维链将图像提炼为包含色彩、风格、场合、季节、材质和平衡性的结构化审美特征。结合视觉摘要、文本描述与审美向量,利用动态熵门机制融合,将候选商品映射至共享嵌入空间并排序。在真实评估集Aesthetic-100上的实验表明,TATTOO性能优于现有训练方法;另在标准Polyvore数据集上验证了其先进的零样本检索能力。
原文摘要 · Abstract (English)
The global fashion e-commerce market relies significantly on intelligent and aesthetic-aware outfit-completion tools to promote sales. While previous studies have approached the problem of fashion outfit-completion and compatible-item retrieval, most of them require expensive, task-specific training on large-scale labeled data, and no effort is made to guide outfit recommendation with explicit human aesthetics. In the era of Multimodal Large Language Models (MLLMs), we show that the conventional training-based pipeline could be streamlined to a training-free paradigm, with better recommendation scores and enhanced aesthetic awareness. We achieve this with TATTOO, a Training-free AesTheTic-aware Outfit recommendation approach. It first generates a target-item description using MLLMs, followed by an aesthetic chain-of-thought used to distill the images into a structured aesthetic profile including color, style, occasion, season, material, and balance. By fusing the visual summary of the outfit with the textual description and aesthetics vectors using a dynamic entropy-gated mechanism, candidate items can be represented in a shared embedding space and be ranked accordingly. Experiments on a real-world evaluation set Aesthetic-100 show that TATTOO achieves state-of-the-art performance compared with existing training-based methods. Another standard Polyvore dataset is also used to measure the advanced zero-shot retrieval capability of our training-free method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。