arXiv:2606.18885cs.CVcs.IR2026-06中稿 · ICML

让图像检索关注被忽略的细节区域,提升复杂场景下的匹配精度

LARE: Low-Attention Region Encoding for Text-Image Retrieval

论文配图:LARE: Low-Attention Region Encoding for Text-Image Retrieval
图 1 · 摘自论文原文
  • 并行编码图像整体与低关注度区域,生成更丰富的视觉特征
  • 在新构建的Dense-Set数据集上,检索准确率显著提升
  • 适合需要细粒度匹配的视觉搜索任务,如复杂场景识别

在密集场景中,传统视觉编码器因显著性偏差,往往聚焦于主导对象而忽略对细粒度检索至关重要的低关注度区域。本文提出LARE(Low-Attention Region Encoding)框架,显式建模这些被忽视的区域。LARE采用双编码策略,同时编码图像的低注意力区域与完整图像,生成更具多样性和信息量的图像嵌入。为评估在复杂密集场景中的检索性能,我们构建了Dense-Set数据集,该数据集源自COCO和Flickr30K,通过重新标注使图像描述更丰富地涵盖低关注度或此前被忽略的区域。该数据集揭示了现有检索模型的局限性,支持在高密度场景下更严格的评估。实验表明,所提框架通过在共享潜在空间中保留细微、非主导的视觉线索,显著提升了检索性能。

原文摘要 · Abstract (English)

Image retrieval in crowded scenes is particularly challenging due to the salience bias of conventional visual encoders, which tend to focus on dominant objects while neglecting low-attention regions that are often crucial for fine-grained retrieval. We propose LARE (Low-Attention Region Encoding), a framework that explicitly models these overlooked regions. LARE adopts a dual-encoding strategy that encodes low-attention regions of an image and the full image in parallel, leading to more diverse and informative image embeddings. To evaluate image retrieval performance in challenging crowded scenes, we introduce Dense-Set, a challenging subset derived from COCO and Flickr30K. In this subset, images are re-captioned to provide richer descriptions of low-attention or previously overlooked regions. This dataset highlights the limitations of existing retrieval models and enables a more rigorous evaluation under densely crowded scene conditions. Experimental results demonstrate that the proposed framework improves retrieval performance by preserving subtle, non-dominant visual cues within the shared latent space.

图像检索细粒度匹配视觉编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。