arXiv:2510.02292cs.CLcs.CV2025-10EMNLP被引 11

通过提取中间层输出,解析视觉语言模型的内部表征差异。

From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens

  • 支持从任意层提取VLM中间输出,统一配置接口
  • 覆盖16个主流VLM及其30多个变体,可扩展性强
  • 适合研究者分析模型内部机制,加速理解与改进

我们提出VLM-Lens,一个用于系统化基准测试、分析和解释视觉语言模型(VLMs)的工具包。该工具包支持在前向传播过程中提取任何层的中间输出,提供统一的YAML可配置接口,抽象模型特有复杂性,实现跨多种VLM的友好操作。目前支持16个前沿基础VLM及其超过30个变体,并可通过不修改核心逻辑的方式扩展新模型。工具包可无缝集成多种可解释性与分析方法。我们通过两个简单实验展示了其应用效果,揭示了不同层及目标概念下VLM隐藏表征的系统性差异。VLM-Lens已开源,旨在加速社区对VLM的理解与优化。

原文摘要 · Abstract (English)

We introduce VLM-Lens, a toolkit designed to enable systematic benchmarking, analysis, and interpretation of vision-language models (VLMs) by supporting the extraction of intermediate outputs from any layer during the forward pass of open-source VLMs. VLM-Lens provides a unified, YAML-configurable interface that abstracts away model-specific complexities and supports user-friendly operation across diverse VLMs. It currently supports 16 state-of-the-art base VLMs and their over 30 variants, and is extensible to accommodate new models without changing the core logic. The toolkit integrates easily with various interpretability and analysis methods. We demonstrate its usage with two simple analytical experiments, revealing systematic differences in the hidden representations of VLMs across layers and target concepts. VLM-Lens is released as an open-sourced project to accelerate community efforts in understanding and improving VLMs.

视觉语言模型可解释性中间表征工具包

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。