通过语义扰动提升视觉语言模型的自信表达与实际正确性的匹配度。
Object-Level Verbalized Confidence Calibration in Vision-Language Models via Semantic Perturbation
- 用高斯噪声扰动关键物体区域,模拟不同置信度下的视觉模糊。
- 在多个基准上显著提升自信表达与回答正确率的一致性。
- 适合关注模型可信度与可解释性的研究人员和开发者。
视觉语言模型(VLMs)在多模态任务中表现优异,但常出现置信度与回答正确性不匹配的问题,影响用户信任,尤其当模型自信地生成错误或虚构信息时。本文提出一种基于语义扰动的置信度校准框架(CSP),针对以物体为中心的查询优化模型的自信表达。首先构建一个扰动数据集,对关键物体区域施加高斯噪声,以模拟不同置信水平下的视觉不确定性,建立视觉模糊与置信度之间的显式映射。随后采用两阶段训练:先在扰动数据集上进行监督微调,再通过偏好优化进一步提升校准效果。大量实验表明,该方法显著改善了自信表达与回答正确性的一致性,同时保持或提升了整体任务性能。结果表明,语义扰动是一种提升VLM可靠性与可解释性的有效工具。
原文摘要 · Abstract (English)
Vision-language models (VLMs) excel in various multimodal tasks but frequently suffer from poor calibration, resulting in misalignment between their verbalized confidence and response correctness. This miscalibration undermines user trust, especially when models confidently provide incorrect or fabricated information. In this work, we propose a novel Confidence Calibration through Semantic Perturbation (CSP) framework to improve the calibration of verbalized confidence for VLMs in response to object-centric queries. We first introduce a perturbed dataset where Gaussian noise is applied to the key object regions to simulate visual uncertainty at different confidence levels, establishing an explicit mapping between visual ambiguity and confidence levels. We further enhance calibration through a two-stage training process combining supervised fine-tuning on the perturbed dataset with subsequent preference optimization. Extensive experiments on popular benchmarks demonstrate that our method significantly improves the alignment between verbalized confidence and response correctness while maintaining or enhancing overall task performance. These results highlight the potential of semantic perturbation as a practical tool for improving the reliability and interpretability of VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。