提出新型超球面图文模型,解决层次结构嵌入的失真与评估难题。
ARGENT: Adaptive Hierarchical Image-Text Representations
- 设计自适应蕴含损失与范数正则化,防止锥形坍塌。
- 在图像分类、图文检索和层次评估上分别提升0.7、1.1、0.8分。
- 引入角度概率蕴含协议,实现更可靠层次理解评估。
大规模视觉-语言模型(如CLIP)虽具强大语义表征能力,但运行于欧氏空间,难以捕捉视觉与语言概念的固有层次结构。超球几何因指数级体积增长,可低失真地嵌入此类层次结构。然而,现有超球面VLM采用蕴含损失,易导致父节点嵌入向原点收缩,其蕴含锥体向半空间扩张,引发灾难性锥体坍塌,破坏层次结构。此外,现有层次评估方法多依赖检索或相关性指标,受分类体系影响大且负样本模糊。为此,本文提出自适应蕴含损失与范数正则化,有效防止锥体坍塌而无需启发式夹紧。进一步设计基于角度的概率蕴含协议(PEP),以AUC-ROC与平均精度评分。本文提出更强的超球面VLM基线模型ARGENT,其在图像分类、文本到图像检索及新提出的层次评估指标上分别取得0.7、1.1、0.8的绝对性能提升。
原文摘要 · Abstract (English)
Large-scale Vision-Language Models (VLMs) such as CLIP learn powerful semantic representations but operate in Euclidean space, which fails to capture the inherent hierarchical structure of visual and linguistic concepts. Hyperbolic geometry, with its exponential volume growth, offers a principled alternative for embedding such hierarchies with low distortion. However, existing hyperbolic VLMs use entailment losses that are unstable: as parent embeddings contract toward the origin, their entailment cones widen toward a half-space, causing catastrophic cone collapse that destroys the intended hierarchy. Additionally, hierarchical evaluation of these models remains unreliable, being largely retrieval-based and correlation-based metrics and prone to taxonomy dependence and ambiguous negatives. To address these limitations, we propose an adaptive entailment loss paired with a norm regularizer that prevents cone collapse without heuristic aperture clipping. We further introduce an angle-based probabilistic entailment protocol (PEP) for evaluating hierarchical understanding, scored with AUC-ROC and Average Precision. This paper introduces a stronger hyperbolic VLM baseline ARGENT, Adaptive hieRarchical imaGe-tExt represeNTation. ARGENT improves the SOTA hyperbolic VLM by 0.7, 1.1, and 0.8 absolute points on image classification, text-to-image retrieval, and proposed hierarchical metrics, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。