arXiv:2503.06670cs.CVcs.CL2025-03被引 8

PixelSHAP让视觉语言模型的注意力可解释,看清它真正关注图像哪些区域。

Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On

  • 基于博弈论的像素级归因方法,通过扰动图像物体分析其对模型输出的影响。
  • 无需模型内部结构,仅用输入输出对即可工作,兼容各类开源与商业模型。
  • 适用于自动驾驶等高风险场景,帮助理解模型决策依据,提升可信度。

视觉语言模型(VLMs)的可解释性对于高风险应用中的信任、调试和决策至关重要。本文提出PixelSHAP,一种模型无关的框架,将基于Shapley值的分析扩展到结构化的视觉实体。与以往聚焦文本提示的方法不同,PixelSHAP通过系统性扰动图像中的物体,量化其对VLM响应的影响,实现基于视觉的推理解释。该方法不依赖模型内部结构,仅需输入-输出对,兼容开源与商用模型。支持多种基于嵌入的相似性度量,并采用受Shapley方法启发的优化技术实现高效计算。我们在自动驾驶场景中验证了PixelSHAP的有效性,展示了其增强可解释性的潜力。主要挑战包括分割敏感性和物体遮挡问题。我们已开源实现,以促进后续研究。

原文摘要 · Abstract (English)

Interpretability in Vision-Language Models (VLMs) is crucial for trust, debugging, and decision-making in high-stakes applications. We introduce PixelSHAP, a model-agnostic framework extending Shapley-based analysis to structured visual entities. Unlike previous methods focusing on text prompts, PixelSHAP applies to vision-based reasoning by systematically perturbing image objects and quantifying their influence on a VLM's response. PixelSHAP requires no model internals, operating solely on input-output pairs, making it compatible with open-source and commercial models. It supports diverse embedding-based similarity metrics and scales efficiently using optimization techniques inspired by Shapley-based methods. We validate PixelSHAP in autonomous driving, highlighting its ability to enhance interpretability. Key challenges include segmentation sensitivity and object occlusion. Our open-source implementation facilitates further research.

可解释性视觉语言模型像素归因

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。