arXiv:2511.22253cs.IRcs.CV2025-11中稿 · ICDM - MMSR Worksh…

用轻量级表示提升零样本图像检索效率,支持图文混合查询。

UNION: A Lightweight Target Representation for Efficient Zero-Shot Image-Guided Retrieval with Optional Textual Queries

  • 将图像嵌入与空文本提示融合,构建通用目标表征
  • 仅用5000样本训练,跨基准达38.5和82.7的mAP@50/200
  • 无需修改预训练模型,适合资源受限场景

图像引导检索(IGROT)是一种通用检索范式,查询可包含锚图及可选文本,旨在检索语义相关的目标图像。该设定统一了组合图像检索(CIR)与草图图像检索(SBIR)。本文提出UNION,一种轻量且可泛化的目标表征,通过融合图像嵌入与空文本提示,在低数据监督下实现高效检索。不同于依赖固定目标特征的传统方法,UNION增强与多模态查询的语义对齐,且无需修改预训练视觉语言模型架构。仅使用5000个训练样本(来自LlavaSCo用于CIR,Training-Sketchy用于SBIR),在多个基准上取得竞争力结果:CIRCO mAP@50为38.5,Sketchy mAP@200为82.7,超越许多强监督基线。这表明UNION在不同查询类型间桥接视觉与语言的鲁棒性与高效性。

原文摘要 · Abstract (English)

Image-Guided Retrieval with Optional Text (IGROT) is a general retrieval setting where a query consists of an anchor image, with or without accompanying text, aiming to retrieve semantically relevant target images. This formulation unifies two major tasks: Composed Image Retrieval (CIR) and Sketch-Based Image Retrieval (SBIR). In this work, we address IGROT under low-data supervision by introducing UNION, a lightweight and generalisable target representation that fuses the image embedding with a null-text prompt. Unlike traditional approaches that rely on fixed target features, UNION enhances semantic alignment with multimodal queries while requiring no architectural modifications to pretrained vision-language models. With only 5,000 training samples - from LlavaSCo for CIR and Training-Sketchy for SBIR - our method achieves competitive results across benchmarks, including CIRCO mAP@50 of 38.5 and Sketchy mAP@200 of 82.7, surpassing many heavily supervised baselines. This demonstrates the robustness and efficiency of UNION in bridging vision and language across diverse query types.

图像检索零样本多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。