arXiv:2605.22552cs.CVcs.MM2026-05被引 1

用统一框架实现多种时尚图像检索,支持真实场景下的多样化查询。

FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning

论文配图:FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning
图 1 · 摘自论文原文
  • 基于多模态大模型构建统一检索框架,动态适配不同任务目标。
  • 在U-FIRE数据集上达到顶尖性能,跨任务泛化能力强。
  • 适合电商、推荐系统等需要灵活图像检索的场景使用。

时尚图像检索是现代电商平台的核心功能。现有方法多局限于特定检索任务,难以应对多样化的实际需求。为此,本文提出统一框架FashionLens,支持多种查询格式与搜索意图。首先构建U-FIRE基准数据集,整合多个碎片化时尚数据集,并补充两个人工标注数据集以评估泛化能力。在此基础上,FashionLens采用多模态大模型架构,设计提案引导的球面查询校准器,通过自适应球面线性插值将查询表示动态映射到任务对齐的度量空间。同时,提出梯度引导的自适应采样策略,根据实时学习难度和数据规模先验自动调整任务权重,缓解优化失衡问题。在U-FIRE上的实验表明,FashionLens在多种检索场景中均达先进水平,且对未见任务具备强泛化能力。代码与数据已开源。

原文摘要 · Abstract (English)

Fashion image retrieval is a cornerstone of modern e-commerce systems. A unified framework that supports diverse query formats and search intentions is highly desired in practice. However, existing approaches focus on narrow retrieval tasks and do not fully capture such diversity. Therefore, in this work, we aim to develop a unified framework capable of handling diverse realistic fashion retrieval scenarios, achieving truly versatile fashion image retrieval. To establish a data foundation, we first introduce U-FIRE, a comprehensive benchmark that consolidates fragmented fashion datasets into a unified collection, supplemented by two manually curated datasets for testing generalization. Building upon this, we propose FashionLens, a unified framework based on Multimodal Large Language Models. To handle divergent matching objectives, we design a Proposal-Guided Spherical Query Calibrator that dynamically shifts query representations into task-aligned metric spaces via adaptive spherical linear interpolation. Additionally, to mitigate the optimization imbalance caused by varying task complexities and data scales, we develop a Gradient-Guided Adaptive Sampling strategy that automatically re-weights tasks based on realtime learning difficulty and the data scale prior. Experiments on U-FIRE show that FashionLens achieves state-of-the-art performance across diverse retrieval scenarios and generalizes robustly to unseen tasks. The data and code are publicly released at https://github.com/haokunwen/FashionLens.

时尚检索多模态自适应电商

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。