arXiv:2506.06938cs.MMcs.CV2025-06被引 2

用局部图像区域提升模糊文本查询的检索效果

Experimental Evaluation of Static Image Sub-Region-Based Search Models Using CLIP

  • 将图像切分为子区域,配合文本描述进行检索
  • 3×3分割和5网格重叠使检索准确率显著提升
  • 适合水下等相似度高场景的图像搜索

多模态文本-图像模型在大规模图像集合中实现了有效的文本查询。尽管在日常场景中表现良好,但在高度同质化的专业领域仍具挑战性,主要原因是用户缺乏专业知识,难以对同类实体进行区分,只能提供模糊的文本描述。本文研究在模糊文本查询基础上加入位置提示是否能提升检索性能。为此,我们收集了741条人工标注数据,包含短/长文本描述及水下复杂场景中感兴趣的区域边界框。基于这些标注,评估了在图像不同静态子区域上使用CLIP进行查询的表现,相比全图查询。结果表明,简单的3×3划分与5网格重叠均能显著提升检索效果,并对标注框扰动具有鲁棒性。

原文摘要 · Abstract (English)

Advances in multimodal text-image models have enabled effective text-based querying in extensive image collections. While these models show convincing performance for everyday life scenes, querying in highly homogeneous, specialized domains remains challenging. The primary problem is that users can often provide only vague textual descriptions as they lack expert knowledge to discriminate between homogenous entities. This work investigates whether adding location-based prompts to complement these vague text queries can enhance retrieval performance. Specifically, we collected a dataset of 741 human annotations, each containing short and long textual descriptions and bounding boxes indicating regions of interest in challenging underwater scenes. Using these annotations, we evaluate the performance of CLIP when queried on various static sub-regions of images compared to the full image. Our results show that both a simple 3-by-3 partitioning and a 5-grid overlap significantly improve retrieval effectiveness and remain robust to perturbations of the annotation box.

图像检索CLIP子区域水下图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。