arXiv:2505.24441cs.CV2025-05被引 1

针对复杂场景中难发现的小物体,提出新基准与多嵌入检索方法。

SORCE: Small Object Retrieval in Complex Environments

  • 用区域提示词引导多模态大模型生成图像的多个特征嵌入。
  • 在SORCE-1K上显著超越现有方法,小物体召回率提升超过20%。
  • 适合关注细粒度视觉检索、真实场景应用的研究者。

文本到图像检索(T2IR)旨在将文本查询与图库中的图像匹配。现有基准主要关注整体语义或显著前景物体的描述,可能忽略复杂环境中的不显眼小物体。这类小物体检索在真实应用中至关重要,因目标未必显著。为此,我们提出SORCE(复杂环境中小物体检索),一个T2IR的新子领域。构建新基准SORCE-1K,包含复杂环境图像和描述不明显小物体的文本查询,且上下文线索极少。初步分析显示,现有T2IR方法难以捕捉小物体并编码全部语义至单一嵌入,导致在SORCE-1K上表现不佳。因此,我们提出以多个独特嵌入表示每张图像,利用多模态大模型(MLLM)通过一组区域提示词(ReP)提取。实验表明,该多嵌入方法通过MLLM与ReP显著优于现有T2IR方法,在SORCE-1K上取得更高检索性能。实验验证了SORCE-1K作为基准的有效性,凸显多嵌入表示与定制化文本特征对解决该任务的潜力。

原文摘要 · Abstract (English)

Text-to-Image Retrieval (T2IR) is a highly valuable task that aims to match a given textual query to images in a gallery. Existing benchmarks primarily focus on textual queries describing overall image semantics or foreground salient objects, possibly overlooking inconspicuous small objects, especially in complex environments. Such small object retrieval is crucial, as in real-world applications, the targets of interest are not always prominent in the image. Thus, we introduce SORCE (Small Object Retrieval in Complex Environments), a new subfield of T2IR, focusing on retrieving small objects in complex images with textual queries. We propose a new benchmark, SORCE-1K, consisting of images with complex environments and textual queries describing less conspicuous small objects with minimal contextual cues from other salient objects. Preliminary analysis on SORCE-1K finds that existing T2IR methods struggle to capture small objects and encode all the semantics into a single embedding, leading to poor retrieval performance on SORCE-1K. Therefore, we propose to represent each image with multiple distinctive embeddings. We leverage Multimodal Large Language Models (MLLMs) to extract multiple embeddings for each image instructed by a set of Regional Prompts (ReP). Experimental results show that our multi-embedding approach through MLLM and ReP significantly outperforms existing T2IR methods on SORCE-1K. Our experiments validate the effectiveness of SORCE-1K for benchmarking SORCE performances, highlighting the potential of multi-embedding representation and text-customized MLLM features for addressing this task.

小物体检索多嵌入文本图像匹配复杂环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。