聚焦服装细节属性,提升视觉语言模型在时尚领域的理解能力
FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training
- 设计属性强调的文本预测任务,强化文本中细粒度特征学习
- 提出属性促进的图像重建任务,增强图像中细节特征表达
- 在检索和识别任务上显著优于现有方法,适合时尚跨模态应用
大规模视觉语言预训练在通用领域表现优异,但在时尚领域,商品差异依赖纹理、材质等细粒度属性,这对检索等任务至关重要。现有模型难以有效利用图文双模态中的细粒度属性。为此,我们提出时尚领域细粒度属性增强的视觉语言预训练方法 FashionFAE,聚焦时尚数据的细节特征。设计属性强调的文本预测任务,迫使模型关注文本中的关键属性;同时提出属性促进的图像重建任务,利用图像模态的代表性属性进一步提升模型对细粒度特征的建模能力。大量实验表明,FashionFAE 显著优于现有最先进方法,在子测试集和全测试集上的检索性能分别提升 2.9% 和 5.2%,在识别任务上平均提升 1.6%。
原文摘要 · Abstract (English)
Large-scale Vision-Language Pre-training (VLP) has demonstrated remarkable success in the general domain. However, in the fashion domain, items are distinguished by fine-grained attributes like texture and material, which are crucial for tasks such as retrieval. Existing models often fail to leverage these fine-grained attributes from both text and image modalities. To address the above issues, we propose a novel approach for the fashion domain, Fine-grained Attributes Enhanced VLP (FashionFAE), which focuses on the detailed characteristics of fashion data. An attribute-emphasized text prediction task is proposed to predict fine-grained attributes of the items. This forces the model to focus on the salient attributes from the text modality. Additionally, a novel attribute-promoted image reconstruction task is proposed, which further enhances the fine-grained ability of the model by leveraging the representative attributes from the image modality. Extensive experiments show that FashionFAE significantly outperforms State-Of-The-Art (SOTA) methods, achieving 2.9% and 5.2% improvements in retrieval on sub-test and full test sets, respectively, and a 1.6% average improvement in recognition tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。