arXiv:2503.07038cs.CV2025-03NeurIPS被引 1

通过多物体注意力优化,提升复杂场景中微小目标的图像检索精度。

Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization

  • 设计多物体预训练+注意力特征融合的检索框架
  • 在零样本和轻量微调下显著优于现有方法
  • 适合需要精准小目标检索的应用场景

我们针对小物体图像检索(SoIR)任务展开研究,目标是在杂乱场景中找到包含特定微小物体的图像。该任务的核心挑战在于构建一个能有效表征图像中所有物体的单一图像描述符,以支持高效、可扩展的搜索。本文首先分析了现有方法在此任务上的局限性,并提出了新的基准评估体系。随后,我们提出多物体注意力优化(MaO)框架,包含专门的多物体预训练阶段,以及利用物体掩码进行注意力特征提取的精炼过程,最终整合为统一的图像描述符。实验表明,MaO在零样本和轻量级多物体微调设置下均显著超越现有方法与强基线。我们希望本工作能为该高实用性任务的性能提升奠定基础并激发后续研究。代码与数据可在项目主页获取:https://pihash2k.github.io/findyourneedle.github.io。

原文摘要 · Abstract (English)

We address the challenge of Small Object Image Retrieval (SoIR), where the goal is to retrieve images containing a specific small object, in a cluttered scene. The key challenge in this setting is constructing a single image descriptor, for scalable and efficient search, that effectively represents all objects in the image. In this paper, we first analyze the limitations of existing methods on this challenging task and then introduce new benchmarks to support SoIR evaluation. Next, we introduce Multi-object Attention Optimization (MaO), a novel retrieval framework which incorporates a dedicated multi-object pre-training phase. This is followed by a refinement process that leverages attention-based feature extraction with object masks, integrating them into a single unified image descriptor. Our MaO approach significantly outperforms existing retrieval methods and strong baselines, achieving notable improvements in both zero-shot and lightweight multi-object fine-tuning. We hope this work will lay the groundwork and inspire further research to enhance retrieval performance for this highly practical task. Code and Data are available on our project page: $\href{https://pihash2k.github.io/findyourneedle.github.io}{https://pihash2k.github.io/findyourneedle.github.io}$.

图像检索小目标识别注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。