用直接反馈优化大模型,让穿搭推荐更懂潮流且个性化。
Decoding Style: Efficient Fine-Tuning of LLMs for Image-Guided Outfit Recommendation with Preference
- 用多模态大模型提取图像风格特征,打通图文信息断层。
- 在Polyvore数据集上微调,显著提升搭配一致性和趋势契合度。
- 通过负例反馈实现自我迭代,适合时尚电商和个性化推荐场景。
个性化穿搭推荐仍面临巨大挑战,需兼具时尚搭配理解与潮流感知能力。本文提出一种新框架,利用大语言模型(LLM)的表达力完成该任务,通过微调与直接反馈整合缓解其“黑箱”与静态问题。通过多模态大语言模型(MLLM)进行图像描述生成,弥合商品图文之间的语义鸿沟,使LLM能从人工精选的时尚图像中提取风格与色彩特征,作为个性化推荐的基础。在开源的Polyvore数据集上对LLM进行高效微调,优化其推荐时尚搭配的能力。采用基于负例的直接偏好机制增强决策过程,构建持续优化的AI反馈循环,使推荐结果与季节性潮流保持一致。框架在Polyvore数据集上评估,涵盖填空与互补商品检索两项任务,验证了其生成符合潮流、结构连贯的穿搭建议的能力,并通过直接反馈实现性能持续提升。实验表明,该框架显著优于基线LLM,生成搭配更具一致性,展现出提升购物体验的潜力。
原文摘要 · Abstract (English)
Personalized outfit recommendation remains a complex challenge, demanding both fashion compatibility understanding and trend awareness. This paper presents a novel framework that harnesses the expressive power of large language models (LLMs) for this task, mitigating their "black box" and static nature through fine-tuning and direct feedback integration. We bridge the item visual-textual gap in items descriptions by employing image captioning with a Multimodal Large Language Model (MLLM). This enables the LLM to extract style and color characteristics from human-curated fashion images, forming the basis for personalized recommendations. The LLM is efficiently fine-tuned on the open-source Polyvore dataset of curated fashion images, optimizing its ability to recommend stylish outfits. A direct preference mechanism using negative examples is employed to enhance the LLM's decision-making process. This creates a self-enhancing AI feedback loop that continuously refines recommendations in line with seasonal fashion trends. Our framework is evaluated on the Polyvore dataset, demonstrating its effectiveness in two key tasks: fill-in-the-blank, and complementary item retrieval. These evaluations underline the framework's ability to generate stylish, trend-aligned outfit suggestions, continuously improving through direct feedback. The evaluation results demonstrated that our proposed framework significantly outperforms the base LLM, creating more cohesive outfits. The improved performance in these tasks underscores the proposed framework's potential to enhance the shopping experience with accurate suggestions, proving its effectiveness over the vanilla LLM based outfit generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。