针对特定领域文本图像检索,构建多匹配新基准并提出高效优化框架。
Rethinking Text-Based Image Retrieval in Specific Domain

- 设计多匹配数据引擎,构建5万张监控图像的专用检索基准
- 在新基准上,模型平均mAP@20提升7.8点,超越传统对比学习
- 适合关注安全监控、工业质检等实际场景的检索研究者
随着视觉语言表征学习的发展,基于文本的图像检索(TBIR)取得显著进展。然而,现有基准普遍基于单匹配假设,难以反映特定领域(如监控)的实际性能——同一查询常对应多个相关图像。为此,我们设计了领域特定多匹配文本图像检索(DSMM-TBIR)数据引擎,并构建了包含5万张监控图像与200个综合查询的Security Multi-Match TBIR(SecMM-TBIR)基准。实验发现,通用对比学习在特定领域中因严重误负样本导致语义相似对被错误分离,损害检索效果。为此提出语义感知微调(SAFT)框架,结合语义感知软标签监督(SASS)与模态内结构蒸馏(ISD),形成新范式。在多种CLIP类模型上,SAFT相较标准图像-文本对比微调(ITC),在SecMM-TBIR上实现平均mAP@20提升7.8点,同时提升通用领域表现。完整基准将公开以推动后续研究。
原文摘要 · Abstract (English)
Driven by the rapid advancement of vision-language representation learning, Text-based Image Retrieval (TBIR) has made notable progress. However, existing benchmarks are predominantly constructed on an exclusive single-match assumption between query and images. While effective in general scenarios, this assumption fails to reflect practical system performance in specific domains (e.g., surveillance), where a single query often corresponds to multiple relevant candidate images. To address this limitation, we design a Domain-Specific Multi-Match Text-based Image Retrieval (DSMM-TBIR) data engine. Leveraging this engine, we construct Security Multi-Match TBIR (SecMM-TBIR), a benchmark comprising 50k surveillance images with 200 comprehensive queries. Furthermore, we observe that vanilla contrastive learning in specific domains suffers from severe false negatives, forcing the model to push apart semantically similar pairs and thus degrading retrieval performance. We propose the Semantic-Aware Fine-Tuning (SAFT) framework to address semantic compression in specific domains, which incorporates Semantic-Aware Soft-Label Supervision (SASS) and Intra-modal Structural Distillation (ISD) to establish a promising paradigm for domain-specific TBIR tasks. Experiments across diverse CLIP-like models demonstrate that SAFT yields an average mAP@20 gain of 7.8 points on SecMM-TBIR over standard image-text contrastive (ITC) fine-tuning, while also improving general-domain performance. The entire benchmark will be released to facilitate further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。