arXiv:2607.20291cs.CV2026-07

构建多样意图的时尚图像多轮检索数据集并提出直接对齐的视觉对话框架。

Diverse-Intent Multi-Turn Fashion Image Retrieval

论文配图:Diverse-Intent Multi-Turn Fashion Image Retrieval
图 1 · 摘自论文原文
  • 构建26000条多轮交互数据,涵盖7类任务与多种意图变化模式。
  • 无需文本中间表示,直接将多模态查询对齐到时尚图像嵌入空间。
  • 适合研究多轮时尚搜索、跨模态对齐与交互式推荐的学者与工程师。

真实世界的时尚搜索涉及多轮交互式检索。然而,现有方法受限于每轮都遵循相同属性编辑范式的假设,未探索意图的多样性变化;且通常依赖文本化过程连接多模态查询与视觉检索,可能丢失细粒度视觉线索。为此,我们提出了DIM-Fashion,一个包含26,000条多轮会话的基准数据集,源自7个任务下的13个时尚检索数据集,具有多样化的意图转换和回滚行为。同时,我们提出FashionAM,一种基于多模态大语言模型与视觉语言预训练的框架,可直接将多模态对话查询对齐至面向时尚的图像嵌入空间,避免中间文本化步骤。大量实验表明,FashionAM在多个指标上优于现有方法。数据集与代码将在论文被接受后公开。

原文摘要 · Abstract (English)

Real-world fashion search involves interactive retrieval across multiple turns. However, existing multi-turn retrieval methods are built on a restrictive assumption that every interaction follows the same attribute-editing paradigm, leaving heterogeneous intent transitions unexplored. Moreover, existing approaches often rely on textification to bridge multimodal queries and visual retrieval, which may lose fine-grained visual cues. To address these gaps, we introduce DIM-Fashion, a benchmark of 26K multi-turn sessions constructed from 13 fashion retrieval datasets across 7 tasks, featuring diverse intent transitions and rollback behaviors. We further propose FashionAM, an MLLM-VLP framework that directly aligns multimodal conversational queries with a fashion-oriented gallery embedding space, avoiding intermediate textification. Extensive experiments demonstrate the effectiveness of FashionAM over existing approaches. The dataset and code will be made publicly available upon acceptance.

多轮检索时尚搜索视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。