FastRM快速生成多模态模型可解释性热图,显著降低计算开销。
FastRM: An efficient and automatic explainability framework for multimodal generative models
- 基于高效算法预测多模态模型的可解释性相关图。
- 计算时间减少99.8%,内存占用降低44.4%。
- 适合需要实时验证输出可信度的应用场景。
大型视觉语言模型(LVLMs)在文本和视觉输入上展现出强大的推理能力,但容易产生无依据的虚假信息。识别并缓解不当响应对构建可信AI至关重要。传统基于梯度的可解释性方法虽能揭示决策过程,但计算成本高,难以用于实时输出验证。本文提出FastRM,一种高效预测LVLM可解释性相关图的方法,并提供模型置信度的定量与定性评估。实验表明,FastRM相比传统方法计算时间减少99.8%,内存占用降低44.4%。该方法使可解释AI更实用、可扩展,推动其在真实场景中的部署,帮助用户更有效评估模型输出可靠性。
原文摘要 · Abstract (English)
Large Vision Language Models (LVLMs) have demonstrated remarkable reasoning capabilities over textual and visual inputs. However, these models remain prone to generating misinformation. Identifying and mitigating ungrounded responses is crucial for developing trustworthy AI. Traditional explainability methods such as gradient-based relevancy maps, offer insight into the decision process of models, but are often computationally expensive and unsuitable for real-time output validation. In this work, we introduce FastRM, an efficient method for predicting explainable Relevancy Maps of LVLMs. Furthermore, FastRM provides both quantitative and qualitative assessment of model confidence. Experimental results demonstrate that FastRM achieves a 99.8% reduction in computation time and a 44.4% reduction in memory footprint compared to traditional relevancy map generation. FastRM allows explainable AI to be more practical and scalable, thereby promoting its deployment in real-world applications and enabling users to more effectively evaluate the reliability of model outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。