arXiv:2506.21476cs.CV2025-06ICCV被引 3

通过显式建模蕴含关系的层次结构,提升视觉语言模型对生物分类的语义理解能力。

Global and Local Entailment Learning for Natural World Imagery

  • 提出径向跨模态嵌入框架,显式约束概念间的传递性顺序。
  • 在物种分类与层级检索任务中,性能超越现有最优模型。
  • 适合研究视觉语言模型层次表征、生物图像理解的学者使用。

视觉语言模型中学习数据的层次结构是一项重大挑战。以往方法虽尝试通过蕴含学习应对,但未能显式建模蕴含关系的传递性,而该性质决定了表示空间中语义与顺序的关系。本文提出径向跨模态嵌入(RCME)框架,实现对传递性约束的蕴含关系的显式建模。该框架优化视觉语言模型中概念的偏序关系。基于此框架,我们构建了可表征生命之树层次结构的层级视觉语言基础模型。在层级物种分类与层级检索任务上的实验表明,所提模型性能优于现有最先进模型。代码与模型已开源至 https://vishu26.github.io/RCME/index.html。

原文摘要 · Abstract (English)

Learning the hierarchical structure of data in vision-language models is a significant challenge. Previous works have attempted to address this challenge by employing entailment learning. However, these approaches fail to model the transitive nature of entailment explicitly, which establishes the relationship between order and semantics within a representation space. In this work, we introduce Radial Cross-Modal Embeddings (RCME), a framework that enables the explicit modeling of transitivity-enforced entailment. Our proposed framework optimizes for the partial order of concepts within vision-language models. By leveraging our framework, we develop a hierarchical vision-language foundation model capable of representing the hierarchy in the Tree of Life. Our experiments on hierarchical species classification and hierarchical retrieval tasks demonstrate the enhanced performance of our models compared to the existing state-of-the-art models. Our code and models are open-sourced at https://vishu26.github.io/RCME/index.html.

视觉语言模型层次表征蕴含学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。