arXiv:2410.04609cs.CV2024-10被引 2

构建图文注意力数据集,揭示模型如何对齐视觉与语言信息

VISTA: A Visual and Textual Attention Dataset for Interpreting Multimodal Models

  • 人工标注图像区域与文本片段的对应关系,生成对齐数据集
  • 对比模型热力图与真实注意力分布,发现多数模型存在偏差
  • 适用于希望理解多模态模型决策机制的研究者

深度学习的发展推动了自然语言处理与计算机视觉的融合,催生了强大的视觉语言模型(VLMs)。尽管性能卓越,这些模型常被视为黑箱,引发关键问题:图像中的哪些区域对应文本中的特定部分?为解答此问题,我们构建了一个图像-文本对齐的人类视觉注意力数据集,精确映射图像区域与对应文本段落。随后,我们将VLMs生成的内部热力图与该数据集进行对比,系统分析模型的决策过程。研究聚焦于文本引导的视觉显著性检测,揭示不同模型在响应文本提示时对视觉元素的关注方式,深化对模型内部机制的理解,提升其可解释性与可信度。

原文摘要 · Abstract (English)

The recent developments in deep learning led to the integration of natural language processing (NLP) with computer vision, resulting in powerful integrated Vision and Language Models (VLMs). Despite their remarkable capabilities, these models are frequently regarded as black boxes within the machine learning research community. This raises a critical question: which parts of an image correspond to specific segments of text, and how can we decipher these associations? Understanding these connections is essential for enhancing model transparency, interpretability, and trustworthiness. To answer this question, we present an image-text aligned human visual attention dataset that maps specific associations between image regions and corresponding text segments. We then compare the internal heatmaps generated by VL models with this dataset, allowing us to analyze and better understand the model's decision-making process. This approach aims to enhance model transparency, interpretability, and trustworthiness by providing insights into how these models align visual and linguistic information. We conducted a comprehensive study on text-guided visual saliency detection in these VL models. This study aims to understand how different models prioritize and focus on specific visual elements in response to corresponding text segments, providing deeper insights into their internal mechanisms and improving our ability to interpret their outputs.

多模态可解释性视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。