提出可解释大模型视觉推理的新方法,精准定位图文关联证据
GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models
- 融合梯度与层传播,联合分析图像和文本的贡献
- 生成高保真跨模态热图,人类注意力匹配度达新高
- 适用于诊断幻觉、偏见等模型缺陷,提升透明性
近期的大规模视觉语言模型(LVLM)在视觉问答任务中表现卓越,但理解其视觉注意力聚焦位置仍具挑战,而这对于解析模型行为至关重要。我们提出 GLIMPSE(用于提示式视觉显著性解释的梯度-层重要性映射),一种轻量级、模型无关的框架,能够联合归因于支持开放式生成的最相关视觉证据与文本信号。GLIMPSE 通过融合梯度加权注意力、自适应层传播和相关性加权标记聚合,生成响应级别的整体热图,以解释跨模态推理,在忠实性上超越先前方法,并在人类注意力对齐方面达到最新水平。我们展示了一种分析方法,可揭示 LVLM 跨模态归因的细粒度洞察,追踪推理动态,分析系统性错位,诊断幻觉与偏见,确保模型透明性。
原文摘要 · Abstract (English)
Recent large vision-language models (LVLMs) have advanced capabilities in visual question answering (VQA). However, interpreting where LVLMs direct their visual attention remains a significant challenge, yet is essential for understanding model behavior. We introduce GLIMPSE (Gradient-Layer Importance Mapping for Prompted Visual Saliency Explanation), a lightweight, model-agnostic framework that jointly attributes LVLM outputs to the most relevant visual evidence and textual signals that support open-ended generation. GLIMPSE fuses gradient-weighted attention, adaptive layer propagation, and relevance-weighted token aggregation to produce holistic response-level heat maps for interpreting cross-modal reasoning, outperforming prior methods in faithfulness and pushing the state-of-the-art in human-attention alignment. We demonstrate an analytic approach to uncover fine-grained insights into LVLM cross-modal attribution, trace reasoning dynamics, analyze systematic misalignment, diagnose hallucination and bias, and ensure transparency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。