arXiv:2606.14728cs.CV2026-06

提出融合两种不确定性的新方法,提升视觉语言模型输出可信度

FUSE: Quantifying Uncertainty in Vision-Language Models by Bayesian Fusing Epistemic and Aleatoric Uncertainty

论文配图:FUSE: Quantifying Uncertainty in Vision-Language Models by Bayesian Fusing Epistemic and Aleatoric Uncertainty
图 1 · 摘自论文原文
  • 通过贝叶斯融合整合输入模糊与模型响应多样性
  • 在多个数据集上实现最优不确定性校准性能
  • 适合对可靠性要求高的机器人等场景应用

视觉语言模型在多个领域日益重要。在机器人等应用中,量化模型输出的不确定性至关重要。我们提出FUSE,一种概率框架,捕捉视觉语言建模中的两种互补不确定性来源:(i) 源自输入数据视觉-语言模糊性的似然性嵌入级不确定性,(ii) 基于模型语义响应多样性的认知性模型级不确定性。该方法构建了贝叶斯融合机制,解析地结合两类不确定性,生成标量形式的不确定性度量。该度量可有效预测下游任务中模型输出的正确性。实验表明,本方法优于基线,在多个数据集上达到最先进不确定性校准效果。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are playing an increasingly important role across multiple domains. In many applications, such as robotics, it is crucial to quantify the uncertainty in the output of these models. } We develop FUSE, a probabilistic framework for capturing two complementary sources of uncertainty in vision-language modeling: (i) aleatoric embedding-level uncertainty derived from input data vision-language ambiguity, and (ii) epistemic model-level uncertainty estimated from the semantic response diversity of VLMs. Our approach formulates a Bayesian fusion mechanism that analytically combines these uncertainty sources to produce a scalar measure of uncertainty. This measure can be used to reliably predict the model's output correctness for downstream applications. We demonstrate that our method outperforms baselines and achieves SOTA uncertainty calibration.

视觉语言模型不确定性量化贝叶斯方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。