arXiv:2603.22042cs.CVcs.AI2026-03中稿 · CVPR被引 2

用不确定性建模图像部件对整体的语义代表性,提升视觉语言模型理解复杂场景能力。

Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models

  • 引入超球面不确定性表示部件对整体的语义重要性,越重要则不确定性越低。
  • 在零样本分类、检索和多标签分类上达到当前最佳性能。
  • 适合关注多物体组合理解与层次结构建模的研究者。

尽管视觉语言模型(VLMs)取得了显著进展,其欧氏空间嵌入仍难以捕捉如部件-整体或父子关系等层次结构,在多对象组合场景中表现不佳。超球面视觉语言模型通过蕴含关系更好地保留层次结构并建模部件-整体关系。然而现有方法未考虑各部件对整体具有不同语义代表性。本文提出不确定性引导的组合式超球面对齐(UNCHA),利用超球面不确定性建模部件对整体的语义代表性:越具代表性的部件不确定性越低,反之越高。该代表性被融入对比学习目标,以不确定性加权。最终通过熵正则化的蕴含损失校准不确定性。所提方法使超球面嵌入更准确地反映部件-整体顺序,捕捉图像内在组合结构,提升对复杂多对象场景的理解。UNCHA在零样本分类、检索和多标签分类基准上均达到最先进水平。代码与模型已开源:https://github.com/jeeit17/UNCHA.git。

原文摘要 · Abstract (English)

While Vision-Language Models (VLMs) have achieved remarkable performance, their Euclidean embeddings remain limited in capturing hierarchical relationships such as part-to-whole or parent-child structures, and often face challenges in multi-object compositional scenarios. Hyperbolic VLMs mitigate this issue by better preserving hierarchical structures and modeling part-whole relations (i.e., whole scene and its part images) through entailment. However, existing approaches do not model that each part has a different level of semantic representativeness to the whole. We propose UNcertainty-guided Compositional Hyperbolic Alignment (UNCHA) for enhancing hyperbolic VLMs. UNCHA models part-to-whole semantic representativeness with hyperbolic uncertainty, by assigning lower uncertainty to more representative parts and higher uncertainty to less representative ones for the whole scene. This representativeness is then incorporated into the contrastive objective with uncertainty-guided weights. Finally, the uncertainty is further calibrated with an entailment loss regularized by entropy-based term. With the proposed losses, UNCHA learns hyperbolic embeddings with more accurate part-whole ordering, capturing the underlying compositional structure in an image and improving its understanding of complex multi-object scenes. UNCHA achieves state-of-the-art performance on zero-shot classification, retrieval, and multi-label classification benchmarks. Our code and models are available at: https://github.com/jeeit17/UNCHA.git.

视觉语言模型超球面嵌入组合理解不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。