arXiv:2502.11748cs.CV2025-02CVPR被引 25

构建大规模实例级图像检索数据集,评估模型识别具体物体能力。

ILIAS: Instance-Level Image retrieval At Scale

  • 设计包含1000个物体实例的测试集,覆盖多样场景与挑战条件。
  • 在1亿张干扰图像中检索,模型表现仍有巨大提升空间。
  • 适合研究视觉语言模型、跨域检索与细粒度识别的学者使用。

本文提出ILIAS,一个用于大规模实例级图像检索的新测试数据集。该数据集旨在评估当前及未来基础模型和检索技术对特定物体的识别能力。相较于现有数据集,ILIAS具有大规模、领域多样性、精确标注以及性能尚未饱和等优势。数据集包含1000个物体实例的查询图与正样本图,均人工收集以涵盖复杂条件和多样化领域。大规模检索任务基于来自YFCC100M的1亿张干扰图像进行。为避免误判负例且无需额外标注,仅纳入2014年后出现的物体(即晚于YFCC100M编目时间)。广泛基准测试显示:i) 在特定领域(如地标或产品)微调的模型在该领域表现优异,但在ILIAS上表现不佳;ii) 通过多领域类别监督学习线性适配层可显著提升性能,尤其对视觉-语言模型有效;iii) 局部描述子在重排序阶段仍是关键,尤其在背景杂乱时;iv) 视觉-语言基础模型的文本到图像检索性能出人意料地接近图像到图像情况。

原文摘要 · Abstract (English)

This work introduces ILIAS, a new test dataset for Instance-Level Image retrieval At Scale. It is designed to evaluate the ability of current and future foundation models and retrieval techniques to recognize particular objects. The key benefits over existing datasets include large scale, domain diversity, accurate ground truth, and a performance that is far from saturated. ILIAS includes query and positive images for 1,000 object instances, manually collected to capture challenging conditions and diverse domains. Large-scale retrieval is conducted against 100 million distractor images from YFCC100M. To avoid false negatives without extra annotation effort, we include only query objects confirmed to have emerged after 2014, i.e. the compilation date of YFCC100M. An extensive benchmarking is performed with the following observations: i) models fine-tuned on specific domains, such as landmarks or products, excel in that domain but fail on ILIAS ii) learning a linear adaptation layer using multi-domain class supervision results in performance improvements, especially for vision-language models iii) local descriptors in retrieval re-ranking are still a key ingredient, especially in the presence of severe background clutter iv) the text-to-image performance of the vision-language foundation models is surprisingly close to the corresponding image-to-image case. website: https://vrg.fel.cvut.cz/ilias/

图像检索实例级大规模视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。