提升视觉语言模型定位可解释性,让模型决策更透明可信。
Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models
- 引入分层语义关系模块,多尺度融合图像区域语义
- 通过可学习权重优化注意力分配,提升解释精度
- 适用于医疗、自动驾驶等高可靠性场景
视觉语言模型在图像分析领域取得显著进展,但在安全关键应用中仍面临挑战,源于物体间复杂关系、细微视觉线索以及对透明度与可靠性的高要求。本文提出多模态可解释学习(MMEL)框架,在保持高性能的同时增强模型可解释性。基于梯度解释方法(Grad-eclip),MMEL引入新型分层语义关系模块,通过多尺度特征处理、自适应注意力加权和跨模态对齐,捕捉不同粒度下图像区域间的语义关联,并使用可学习的层特定权重平衡模型深度中各层贡献。实验表明,将语义关系信息融入梯度归因图,使可视化结果更聚焦且具上下文感知能力,更真实反映模型对复杂场景的处理方式。该框架具备跨领域泛化能力,为需要高可解释性与可靠性的应用提供决策洞察。
原文摘要 · Abstract (English)
Recent advances in vision-language models have significantly expanded the frontiers of automated image analysis. However, applying these models in safety-critical contexts remains challenging due to the complex relationships between objects, subtle visual cues, and the heightened demand for transparency and reliability. This paper presents the Multi-Modal Explainable Learning (MMEL) framework, designed to enhance the interpretability of vision-language models while maintaining high performance. Building upon prior work in gradient-based explanations for transformer architectures (Grad-eclip), MMEL introduces a novel Hierarchical Semantic Relationship Module that enhances model interpretability through multi-scale feature processing, adaptive attention weighting, and cross-modal alignment. Our approach processes features at multiple semantic levels to capture relationships between image regions at different granularities, applying learnable layer-specific weights to balance contributions across the model's depth. This results in more comprehensive visual explanations that highlight both primary objects and their contextual relationships with improved precision. Through extensive experiments on standard datasets, we demonstrate that by incorporating semantic relationship information into gradient-based attribution maps, MMEL produces more focused and contextually aware visualizations that better reflect how vision-language models process complex scenes. The MMEL framework generalizes across various domains, offering valuable insights into model decisions for applications requiring high interpretability and reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。