无需标注,用自监督模型发现图像中实例级特征区域。
Class Agnostic Instance-level Descriptor for Visual Instance Search
- 基于自监督ViT特征,分层检测紧凑特征子集以定位实例区域。
- 在三个基准上对已知与未知类别均保持高效,支持单/多实例搜索。
- 适合无标注数据下的图像检索与实例搜索场景。
尽管深度特征在基于内容的图像检索中取得巨大成功,但视觉实例搜索仍具挑战性,主要因缺乏有效的实例级特征表示。监督或弱监督目标检测方法不适用,因其在未知类别上表现不佳。本文基于自监督ViT输出的特征集,将实例级区域发现建模为分层方式检测紧凑特征子集。层级分解生成实例区域的层次结构:非叶节点与叶节点分别对应图像中不同粒度的实例区域。由此产生的特征长度统一,可覆盖主导区域、多个实例组合或单一实例。该特征集合将图像检索、多实例搜索与实例搜索统一于同一框架。三个基准上的实证研究显示,该实例级描述符在已知与未知类别上均有效,且在单实例、多实例搜索及图像检索任务中表现优异。
原文摘要 · Abstract (English)
Despite the great success of the deep features in content-based image retrieval, the visual instance search remains challenging due to the lack of effective instance-level feature representation. Supervised or weakly supervised object detection methods are not the appropriate solutions due to their poor performance on the unknown object categories. In this paper, based on the feature set output from self-supervised ViT, the instance-level region discovery is modeled as detecting the compact feature subsets in a hierarchical fashion. The hierarchical decomposition results in a hierarchy of instance regions. On the one hand, this kind of hierarchical decomposition well addresses the problem of object embedding and occlusions, which are widely observed in real scenarios. On the other hand, the non-leaf nodes and leaf nodes on the hierarchy correspond to the instance regions in different granularities within an image. Therefore, features in uniform length are produced for these instance regions, which may cover across a dominant image region, an integral of multiple instances, or various individual instances. Such a collection of features allows us to unify the image retrieval, multi-instance search, and instance search into one framework. The empirical studies on three benchmarks show that such an instance-level descriptor remains effective on both the known and unknown object categories. Moreover, the superior performance is achieved on single-instance and multi-instance search, as well as image retrieval tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。