arXiv:2509.25844cs.CLcs.HC2025-09ACL

为视觉语言模型解释质量设计评分,提升盲人用户判断准确性

Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations

  • 提出视觉保真度与对比性双指标评估解释质量
  • 在三组数据集上提升预测正确性判断准确率11.1%
  • 适合无障碍视觉系统与可解释AI研究者使用

当用户无法看到视觉上下文(如视障人士)时,仅靠视觉语言模型(VLM)的自然语言解释可能误导判断。现有方法显示,解释易使用户误信错误预测。为此,本文提出两个互补的质量评分函数:视觉保真度(Visual Fidelity),衡量解释与图像内容的一致性;对比性(Contrastiveness),衡量解释是否识别出区分预测与合理替代方案的关键视觉细节。在A-OKVQA、VizWiz和MMMU-Pro任务上,该评分方法比现有解释质量度量更准确地反映模型预测的正确性。用户研究发现,在不查看图像的情况下,展示质量评分可使参与者判断模型正确性的准确率提升11.1%,其中错误相信错误预测的比例下降15.4%。结果表明,解释质量评分有助于引导用户恰当地依赖VLM输出。

原文摘要 · Abstract (English)

When people query Vision-Language Models (VLMs) but cannot see the accompanying visual context (e.g. for blind and low-vision users), augmenting VLM predictions with natural language explanations can signal which model predictions are reliable. However, prior work has found that explanations can easily convince users that inaccurate VLM predictions are correct. To remedy undesirable overreliance on VLM predictions, we propose evaluating two complementary qualities of VLM-generated explanations via two quality scoring functions. We propose Visual Fidelity, which captures how faithful an explanation is to the visual context, and Contrastiveness, which captures how well the explanation identifies visual details that distinguish the model's prediction from plausible alternatives. On the A-OKVQA, VizWiz, and MMMU-Pro tasks, these quality scoring functions are better calibrated with model correctness than existing explanation qualities. We conduct a user study in which participants have to decide whether a VLM prediction is accurate without viewing its visual context. We observe that showing our quality scores alongside VLM explanations improves participants' accuracy at predicting VLM correctness by 11.1%, including a 15.4% reduction in the rate of falsely believing incorrect predictions. These findings highlight the utility of explanation quality scores in fostering appropriate reliance on VLM predictions.

可解释AI视觉语言模型无障碍设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。