arXiv:2606.10697cs.IR2026-06

用超像素令牌提升服装属性检索精度,更准定位细节特征。

Beyond Patches: Superpixel Token-based Transformers for Attribute-Specific Fashion Retrieval

论文配图:Beyond Patches: Superpixel Token-based Transformers for Attribute-Specific Fashion Retrieval
图 1 · 摘自论文原文
  • 用语义引导的超像素分割生成紧凑语义块
  • 在三个数据集上相对领先模型提升9.35%检索准确率
  • 适合需要精细图像匹配的电商和时尚检索场景

属性特定服装检索(ASFR)旨在通过聚焦特定属性来提升细粒度图像检索性能。然而,现有基于补丁的注意力与Transformer方法常因无法对齐不规则属性区域且易受背景噪声干扰,难以捕捉细微的像素级微结构。为此,我们提出SuperFashion,首个采用超像素令牌的ASFR框架。该框架首先通过属性引导注意力机制提取属性相关特征,并据此裁剪出语义有意义的图像区域;随后在这些区域上进行超像素分割,生成紧凑且语义一致的超像素令牌。通过为属性与超像素令牌分别引入模态特异性嵌入,超像素令牌化的Transformer实现了自适应交互与融合,显著提升了属性定位与区分能力。在FashionAI、DARN和DeepFashion上的大量实验表明,相比先前最优方法,其整体MAP分别提升1.84%、9.27%和9.35%。SuperFashion为基于网络的图像检索提供了新思路。

原文摘要 · Abstract (English)

Attribute-Specific Fashion Retrieval (ASFR) aims to improve fine-grained image retrieval by focusing on specific attributes. However, existing patch-based attention and Transformer methods often misalign with irregular attribute regions and are prone to background noise, limiting their ability to capture subtle, pixel-level microstructures. To tackle these challenges, we propose SuperFashion, the first ASFR framework that adopts superpixel tokens within a Transformer architecture. SuperFashion initially employs an attribute-guided attention mechanism to extract attribute-related features, which in turn guide the cropping of semantically meaningful image regions. Superpixel segmentation is then leveraged on these regions to generate compact, semantically coherent superpixel tokens. By incorporating modality-specific embeddings for both attribute and superpixel tokens, the superpixel token-based Transformer facilitates adaptive interaction and fusion, thereby enhancing attribute localization and discrimination. Extensive experiments on FashionAI, DARN, and DeepFashion demonstrate relative overall MAP improvements of 1.84%, 9.27%, and 9.35% over prior SOTA. SuperFashion offers a new solution for web-based image retrieval.

服装检索超像素Transformer细粒度识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。