arXiv:2608.09011cs.LG2026-08

动态感知分布变化,提升视觉语言模型预测可靠性

Dynamic Distribution-Aware Uncertainty Tracking in Vision-Language Representation Learning

论文配图:Dynamic Distribution-Aware Uncertainty Tracking in Vision-Language Representation Learning
图 1 · 摘自论文原文
  • 用高斯混合模型建模嵌入空间,动态捕捉数据分布
  • 在测试分布变化时仍保持高不确定性估计精度
  • 适合安全关键场景下对模型可信度有要求的应用

不确定性量化(UQ)旨在衡量模型预测的可靠性,是将视觉语言模型(VLMs)部署于安全关键场景的重要保障。后处理方法因轻量而被广泛采用,通过可学习模块或归纳总结将VLM输出映射为不确定性度量。然而,这类方法仅能拟合源域的失效模式,忽视了测试分布的动态特性。为此,我们提出动态分布感知不确定性量化框架(DDA-UQ),将范式从静态映射转向动态分布感知过程。训练阶段,利用高斯混合模型建模VLM嵌入空间并提取分布证据,从而动态生成不确定性估计;推理阶段,系统能动态响应数据分布变化。大量实验表明,该方法显著优于现有最优方法。

原文摘要 · Abstract (English)

Uncertainty Quantification (UQ) aims to measure the reliability of model predictions, serving as a critical safeguard for deploying Vision-Language Models (VLMs) in safety-critical scenarios. Post-hoc approaches are widely adopted due to their lightweight nature, mapping the outputs of VLMs to uncertainty measures through learnable modules or inductive summarization. However, Post-hoc approaches remain inherently confined to fitting the failure patterns of the source domain, ignoring the dynamic nature of test distributions. To address this challenge, we propose a Dynamic Distribution-Aware Uncertainty Quantification framework (DDA-UQ) that shifts the paradigm from static mapping to a dynamic distribution-aware process. During training, we leverage a Gaussian Mixture Model to model the VVLMs'embedding space and extract distributional evidence, thereby dynamically deriving uncertainty estimates. During inference, the design dynamically responds to changes in the data distribution. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods.

不确定性量化视觉语言模型动态分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。