arXiv:2605.13352cs.LG2026-05

让视觉语言模型同时量化两种不确定性,提升检索与分类的可信度。

GeoFlowVLM: Geometry-Aware Joint Uncertainty for Frozen Vision-Language Embedding

  • 通过黎曼流匹配建模双编码器嵌入的联合分布,保留球面几何特性。
  • 推出条件检索熵和边缘典型性得分,分别衡量模态间模糊性和训练分布外风险。
  • 在三个检索任务和四个零样本分类任务中均实现理想校准,适合高可靠性场景。

标准双编码器视觉语言模型通过ℓ₂归一化将图像和文本映射到共享单位超球面上,通常不暴露统计不确定性(跨模态模糊)和认知不确定性(训练分布外)。现有后处理方法要么仅恢复其中一种不确定性,要么忽略嵌入的超球面几何结构。我们提出GeoFlowVLM,一种后处理适配器,通过单一掩码速度场,在乘积超球面ℝ^{d−1}×ℝ^{d−1}上学习双编码器嵌入的联合分布,采用黎曼流匹配。一致性结果表明:在总体极限下,训练网络能生成有效的黎曼流匹配速度场,覆盖联合分布及跨模态条件分布。由此导出两个量:基于费诺型界解释的条件检索熵,用于量化统计不确定性;以及由联合负对数似然精确链式分解支持的边缘典型性得分,用于评估认知不确定性。该分解分离出一个结构性区分项——点对点互信息,实证上为唯一始终无信息量的独立成分。实验表明,该熵在三个检索基准上近乎理想地单调校准召回率@1,边缘典型性总分在四个零样本分类基准上保持一致校准的可选择准确率。

原文摘要 · Abstract (English)

Standard dual-encoder vision-language models that map images and text to deterministic points on a shared unit hypersphere through $\ell_2$ normalization typically expose neither \emph{aleatoric} uncertainty (cross-modal ambiguity) nor \emph{epistemic} uncertainty (lack of training-distribution support). Existing post-hoc methods either recover at most one of the two uncertainty components, or ignore the hyperspherical geometry of these models' embeddings. We propose \textbf{GeoFlowVLM} as a post-hoc adapter that learns the joint distribution of paired $\ell_2$-normalised dual-encoder VLM embeddings on the product hypersphere $\mathbb{S}^{d-1} \times \mathbb{S}^{d-1}$ via Riemannian flow matching with a single masked velocity field. A consistency result shows that, in the population limit, the trained network exposes the joint flow and both cross-modal conditional flows as valid Riemannian flow-matching velocity fields on their respective domains. We derive two quantities from this single model: a conditional retrieval entropy that quantifies aleatoric ambiguity with a decision-theoretic interpretation via a Fano-type bound, and a marginal-typicality epistemic score justified by an exact chain-rule decomposition of the joint NLL. This decomposition isolates a cross-modal pointwise-mutual-information term that is structurally discriminative rather than epistemic, and is empirically the only consistently uninformative standalone component. Empirically, the entropy tracks Recall@1 with near-ideal monotonic calibration across three retrieval benchmarks in both directions, and the marginal-typicality sum yields consistently calibrated selective accuracy across four zero-shot classification benchmarks.

视觉语言不确定性流匹配检索校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。