arXiv:2411.17490cs.CV2024-11ICCV被引 11

在双曲空间中学习图像的多层语义层次结构,提升检索准确性。

Learning Visual Hierarchies in Hyperbolic Space for Image Retrieval

  • 用对比损失和蕴含度量显式建模图像间的层级关系
  • 在部件级图像检索任务中显著优于传统方法
  • 适合需要理解复杂语义结构的视觉检索场景

将潜在表示以层次化方式组织,有助于模型在多个抽象层次上学习模式。然而,当前主流图像理解模型主要关注视觉相似性,对视觉层次结构的学习仍不充分。本文首次提出一种学习范式,可在无需显式层次标签的情况下,将用户定义的多层级复杂视觉层次结构编码至双曲空间。具体而言,首先利用图像内及跨图像的对象级标注构建基于部件的图像层次结构;接着引入基于成对蕴含度量的对比损失来强制实现该层次结构;最后设计新的评估指标,有效衡量层次化图像检索性能。该编码方式使学习到的表示不仅包含视觉相似性,还捕捉了语义与结构信息。在基于部件的图像检索实验中,本方法在层次化检索任务上取得显著提升,验证了其在捕捉视觉层次结构方面的有效性。

原文摘要 · Abstract (English)

Structuring latent representations in a hierarchical manner enables models to learn patterns at multiple levels of abstraction. However, most prevalent image understanding models focus on visual similarity, and learning visual hierarchies is relatively unexplored. In this work, for the first time, we introduce a learning paradigm that can encode user-defined multi-level complex visual hierarchies in hyperbolic space without requiring explicit hierarchical labels. As a concrete example, first, we define a part-based image hierarchy using object-level annotations within and across images. Then, we introduce an approach to enforce the hierarchy using contrastive loss with pairwise entailment metrics. Finally, we discuss new evaluation metrics to effectively measure hierarchical image retrieval. Encoding these complex relationships ensures that the learned representations capture semantic and structural information that transcends mere visual similarity. Experiments in part-based image retrieval show significant improvements in hierarchical retrieval tasks, demonstrating the capability of our model in capturing visual hierarchies.

图像检索双曲空间层次结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。