提出新模型统一视觉语言中的层次与组合关系,提升语义表达能力。
PHyCLIP: $\ell_1$-Product of Hyperbolic Factors Unifies Hierarchy and Compositionality in Vision-Language Representation Learning
- 在双曲空间的笛卡尔积上使用ℓ₁-乘积度量,分别建模家族内层级与跨家族组合。
- 在零样本分类、检索和组合理解任务中均超越现有方法,最高提升12.3%。
- 结构更可解释,适合需要精准语义建模的应用场景。
视觉语言模型在大规模图像与文本配对数据上取得了显著进展,但仍难以同时表达两种不同类型的语义结构:概念族内的层次关系(如狗 ≤ 哺乳动物 ≤ 动物)和跨概念族的组合关系(如“车里的狗” ≤ 狗,车)。现有方法采用双曲空间以高效捕捉树状层次结构,但其对组合性的表达能力尚不明确。为此,我们提出PHyCLIP,通过在双曲因子的笛卡尔积上使用ℓ₁-乘积度量,在单个双曲因子内自然生成族内层次,并通过ℓ₁-乘积度量捕捉跨族组合,类似于布尔代数。在零样本分类、检索、层次分类和组合理解任务上的实验表明,PHyCLIP优于现有单空间方法,且嵌入空间结构更具可解释性。
原文摘要 · Abstract (English)
Vision-language models have achieved remarkable success in multi-modal representation learning from large-scale pairs of visual scenes and linguistic descriptions. However, they still struggle to simultaneously express two distinct types of semantic structures: the hierarchy within a concept family (e.g., dog $\preceq$ mammal $\preceq$ animal) and the compositionality across different concept families (e.g., "a dog in a car" $\preceq$ dog, car). Recent works have addressed this challenge by employing hyperbolic space, which efficiently captures tree-like hierarchy, yet its suitability for representing compositionality remains unclear. To resolve this dilemma, we propose PHyCLIP, which employs an $\ell_1$-Product metric on a Cartesian product of Hyperbolic factors. With our design, intra-family hierarchies emerge within individual hyperbolic factors, and cross-family composition is captured by the $\ell_1$-product metric, analogous to a Boolean algebra. Experiments on zero-shot classification, retrieval, hierarchical classification, and compositional understanding tasks demonstrate that PHyCLIP outperforms existing single-space approaches and offers more interpretable structures in the embedding space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。