用几何方法量化视觉语言模型的语义不确定性,更稳定可靠。
Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes
- 通过分析答案嵌入的几何结构,避免脆弱的聚类方法。
- 在8个数据集上实现最优校准与区分度性能,ECE更低。
- 适用于多种模型和模态,能捕捉通用语义不确定性模式。
大型视觉语言模型(LVLMs)常产生看似合理但不可靠的输出,因此需要稳健的不确定性估计。现有语义不确定性方法依赖外部模型对多个采样回答进行聚类并测量语义一致性,但此类聚类方法易受细微表述变化影响,常错误分组或分离语义相近的答案,导致不确定性估计不可靠。本文提出语义高斯过程不确定性(SGPU),一种基于贝叶斯框架的方法,通过分析答案嵌入的几何结构来量化语义不确定性,避免了脆弱的聚类过程。SGPU将生成的答案映射到密集语义空间,计算其嵌入的格拉姆矩阵,并通过特征谱总结其语义配置。该谱表示被输入高斯过程分类器,学习将语义一致性模式映射为预测不确定性,可在黑盒与白盒场景中应用。在六种LLM与LVLM、八个涵盖视觉问答、图像分类和文本问答的数据集上,SGPU持续取得最佳校准(ECE)与判别性能(AUROC, AUARC)。此外,我们证明了SGPU在模型与模态间具有迁移能力,表明其谱表示捕捉到了语义不确定性的通用模式。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) often produce plausible but unreliable outputs, making robust uncertainty estimation essential. Recent work on semantic uncertainty estimates relies on external models to cluster multiple sampled responses and measure their semantic consistency. However, these clustering methods are often fragile, highly sensitive to minor phrasing variations, and can incorrectly group or separate semantically similar answers, leading to unreliable uncertainty estimates. We propose Semantic Gaussian Process Uncertainty (SGPU), a Bayesian framework that quantifies semantic uncertainty by analyzing the geometric structure of answer embeddings, avoiding brittle clustering. SGPU maps generated answers into a dense semantic space, computes the Gram matrix of their embeddings, and summarizes their semantic configuration via the eigenspectrum. This spectral representation is then fed into a Gaussian Process Classifier that learns to map patterns of semantic consistency to predictive uncertainty, and that can be applied in both black-box and white-box settings. Across six LLMs and LVLMs on eight datasets spanning VQA, image classification, and textual QA, SGPU consistently achieves state-of-the-art calibration (ECE) and discriminative (AUROC, AUARC) performance. We further show that SGPU transfers across models and modalities, indicating that its spectral representation captures general patterns of semantic uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。