arXiv:2505.11029cs.LG2025-05NeurIPS被引 5

让视觉语言模型更懂不确定性的分布差异。

Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit Hypersphere

  • 在单位超球面上构建不对称不确定性分布,捕捉图文数据差异。
  • 在多个基准上验证概率嵌入有效提升模型鲁棒性。
  • 适合关注模型可信度与多模态不确定性建模的研究者。

视觉语言模型(VLMs)作为基础模型,在众多视觉与文本任务中显著提升了性能,无需为下游任务从头大规模训练。然而,这些确定性 VLMs 无法捕捉自然语言和视觉数据中的固有模糊性与不确定性。近期的后处理概率化方法通过将确定性嵌入映射到概率分布来缓解此问题;但现有方法未考虑模态间的不对称不确定性结构,且忽视有意义的确定性嵌入位于单位超球面上的约束,可能导致性能不佳。本文针对文本与视觉数据内在的不对称不确定性结构,提出 AsymVLM,从预训练的 VLMs 构建单位超球面上的概率嵌入,实现不确定性量化。我们在标准基准上验证了概率嵌入的有效性,并通过全面的消融实验揭示了文本与视觉数据不确定性结构中固有的不对称性。

原文摘要 · Abstract (English)

Vision-language models (VLMs) as foundation models have significantly enhanced performance across a wide range of visual and textual tasks, without requiring large-scale training from scratch for downstream tasks. However, these deterministic VLMs fail to capture the inherent ambiguity and uncertainty in natural language and visual data. Recent probabilistic post-hoc adaptation methods address this by mapping deterministic embeddings onto probability distributions; however, existing approaches do not account for the asymmetric uncertainty structure of the modalities, and the constraint that meaningful deterministic embeddings reside on a unit hypersphere, potentially leading to suboptimal performance. In this paper, we address the asymmetric uncertainty structure inherent in textual and visual data, and propose AsymVLM to build probabilistic embeddings from pre-trained VLMs on the unit hypersphere, enabling uncertainty quantification. We validate the effectiveness of the probabilistic embeddings on established benchmarks, and present comprehensive ablation studies demonstrating the inherent nature of asymmetry in the uncertainty structure of textual and visual data.

多模态不确定性视觉语言模型概率建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。