arXiv:2511.00925cs.CV2025-11

解决草图图像检索中跨模态不平衡问题,提升零样本检索效果。

Dynamic Multi-level Weighted Alignment Network for Zero-shot Sketch-based Image Retrieval

  • 动态加权对齐机制,分层次评估草图与图像匹配质量。
  • 在Sketchy、TU-Berlin、QuickDraw上均超越现有最优方法。
  • 适合关注跨模态对齐与零样本检索的研究者。

零样本草图图像检索(ZS-SBIR)因在电商等场景的广泛应用而受到越来越多关注。尽管该领域已有进展,但以往方法存在模态样本不平衡及训练过程中低质量信息不一致的问题,导致性能受限。为此,本文提出动态多层级加权对齐网络(Dynamic Multi-level Weighted Alignment Network)。该方法包含三部分:(i) 单模态特征提取模块,使用CLIP文本编码器和ViT分别提取文本与视觉标记;(ii) 跨模态多层级加权模块,通过局部与全局聚合块生成对齐权重列表,衡量草图与图像样本的对齐质量;(iii) 加权四元组损失模块,优化三元组损失中的域平衡性。在Sketchy、TU-Berlin和QuickDraw三个基准数据集上的实验表明,本方法显著优于当前最先进的ZS-SBIR方法。

原文摘要 · Abstract (English)

The problem of zero-shot sketch-based image retrieval (ZS-SBIR) has achieved increasing attention due to its wide applications, e.g. e-commerce. Despite progress made in this field, previous works suffer from using imbalanced samples of modalities and inconsistent low-quality information during training, resulting in sub-optimal performance. Therefore, in this paper, we introduce an approach called Dynamic Multi-level Weighted Alignment Network for ZS-SBIR. It consists of three components: (i) a Uni-modal Feature Extraction Module that includes a CLIP text encoder and a ViT for extracting textual and visual tokens, (ii) a Cross-modal Multi-level Weighting Module that produces an alignment weight list by the local and global aggregation blocks to measure the aligning quality of sketch and image samples, (iii) a Weighted Quadruplet Loss Module aiming to improve the balance of domains in the triplet loss. Experiments on three benchmark datasets, i.e., Sketchy, TU-Berlin, and QuickDraw, show our method delivers superior performances over the state-of-the-art ZS-SBIR methods.

跨模态对齐零样本检索草图识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。