arXiv:2506.18985cs.CVcs.AI2025-06被引 1

提出可解释大模型视觉推理的新方法,精准定位图文关联证据

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models

  • 融合梯度与层传播,联合分析图像和文本的贡献
  • 生成高保真跨模态热图,人类注意力匹配度达新高
  • 适用于诊断幻觉、偏见等模型缺陷,提升透明性

近期的大规模视觉语言模型(LVLM)在视觉问答任务中表现卓越,但理解其视觉注意力聚焦位置仍具挑战,而这对于解析模型行为至关重要。我们提出 GLIMPSE(用于提示式视觉显著性解释的梯度-层重要性映射),一种轻量级、模型无关的框架,能够联合归因于支持开放式生成的最相关视觉证据与文本信号。GLIMPSE 通过融合梯度加权注意力、自适应层传播和相关性加权标记聚合,生成响应级别的整体热图,以解释跨模态推理,在忠实性上超越先前方法,并在人类注意力对齐方面达到最新水平。我们展示了一种分析方法,可揭示 LVLM 跨模态归因的细粒度洞察,追踪推理动态,分析系统性错位,诊断幻觉与偏见,确保模型透明性。

原文摘要 · Abstract (English)

Recent large vision-language models (LVLMs) have advanced capabilities in visual question answering (VQA). However, interpreting where LVLMs direct their visual attention remains a significant challenge, yet is essential for understanding model behavior. We introduce GLIMPSE (Gradient-Layer Importance Mapping for Prompted Visual Saliency Explanation), a lightweight, model-agnostic framework that jointly attributes LVLM outputs to the most relevant visual evidence and textual signals that support open-ended generation. GLIMPSE fuses gradient-weighted attention, adaptive layer propagation, and relevance-weighted token aggregation to produce holistic response-level heat maps for interpreting cross-modal reasoning, outperforming prior methods in faithfulness and pushing the state-of-the-art in human-attention alignment. We demonstrate an analytic approach to uncover fine-grained insights into LVLM cross-modal attribution, trace reasoning dynamics, analyze systematic misalignment, diagnose hallucination and bias, and ensure transparency.

可解释性视觉语言模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。