arXiv:2608.26658cs.CVcs.AI2026-08

让模型像人一样看图搜索,精准聚焦目标物体

PailitaoGR: Latent Think-with-Images for Generative Image Retrieval

论文配图:PailitaoGR: Latent Think-with-Images for Generative Image Retrieval
图 1 · 摘自论文原文
  • 通过视觉目标增强和注意力引导,让模型专注搜索目标区域
  • 在真实搜索日志数据上提升13.8%的检索准确率
  • 适合需要精准图像检索的应用场景

生成式检索通过直接生成产品语义标识(SIDs)展现出强大性能。将其扩展至图像搜索面临挑战,因真实查询图像包含目标、辅助证据和无关内容等多元信息,要求模型能识别并聚焦目标,同时选择性利用辅助证据。本文提出PailitaoGR,一种基于潜在空间‘思考-图像’的生成式图像检索方法,将目标聚焦感知与选择性辅助证据利用能力内嵌于生成式检索模型中,实现‘无裁剪缩放’与‘无需OCR阅读’。具体设计包括:目标聚焦感知机制,由目标增强器与基于策略内蒸馏及注意力引导损失的学习策略组成,使模型强化目标区域的视觉标记;选择性辅助证据利用机制,包含辅助增强器与容量增量对比蒸馏策略,有效挖掘辅助证据。训练与验证集基于真实在线图像搜索日志构建。实验表明,该方法相较现有基线平均提升13.8%,验证了其有效性。

原文摘要 · Abstract (English)

Generative retrieval has demonstrated strong performance by directly generating product semantic identifiers (SIDs). Extending this paradigm to image search, however, is nontrivial because real-world query images contain diverse information, including the search target, useful auxiliary evidence, and irrelevant visual content. This requires the model to identify and focus on the search target while selectively utilizing auxiliary evidence. In this paper, we propose \textbf{PailitaoGR}, a \emph{Latent Think-with-Images} method for generative image retrieval, which internalizes target-focused perception and selective auxiliary-evidence utilization into a the generative retrieval model, enabling \textit{Zooming without Cropping} and \textit{Reading without OCR}. Specifically, we design a target-focused perception mechanism that identifies and enhances visual tokens of the search target, consisting of a target Enhancer and a learning strategy based on on-policy distillation and attention guidance loss, enabling the model to focus on search-target regions. We also design a selective auxiliary-evidence utilization mechanism that identifies and enhances visual tokens of auxiliary evidence, including an auxiliary enhancer and an in-capacity incremental contrastive distillation strategy, enabling the model to exploit auxiliary evidence. We construct training and validation sets sampled from real-world online image-search logs. Experiments show that our method outperforms existing baselines by an average of 13.8\%, validating its effectiveness.

图像检索生成式模型视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。