arXiv:2601.21662cs.LG2026-01被引 2

通过黎曼流匹配量化视觉语言模型的不确定性,提升模型可信度判断能力。

Epistemic Uncertainty Quantification for Pre-trained VLMs via Riemannian Flow Matching

  • 基于嵌入空间密度构建不确定性度量,低密度区域表征模型无知
  • 与预测误差相关性接近完美,显著优于现有方法
  • 适用于分布外检测与数据清洗,适合高可靠性场景应用

视觉-语言模型(VLMs)通常为确定性模型,缺乏量化认知不确定性(epistemic uncertainty)的内在机制,而该不确定性反映了模型对其自身表征的知识缺失。本文从理论出发,将嵌入空间的负对数密度作为认知不确定性的代理指标,其中低密度区域代表模型无知。所提方法REPVLM利用黎曼流匹配,在VLM嵌入的超球面流形上计算概率密度。实验表明,REPVLM在不确定性与预测误差之间实现了近乎完美的相关性,显著优于现有基线方法。除分类任务外,该方法还提供了一种可扩展的分布外检测与自动化数据清理指标。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) are typically deterministic in nature and lack intrinsic mechanisms to quantify epistemic uncertainty, which reflects the model's lack of knowledge or ignorance of its own representations. We theoretically motivate negative log-density of an embedding as a proxy for the epistemic uncertainty, where low-density regions signify model ignorance. The proposed method REPVLM computes the probability density on the hyperspherical manifold of the VLM embeddings using Riemannian Flow Matching. We empirically demonstrate that REPVLM achieves near-perfect correlation between uncertainty and prediction error, significantly outperforming existing baselines. Beyond classification, we also demonstrate that the model also provides a scalable metric for out-of-distribution detection and automated data curation.

不确定性量化视觉语言模型黎曼几何数据清洗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。