arXiv:2508.07833cs.CV2025-08中稿 · CVPR被引 2

通过逆向生成图像解释视觉语言模型的内部概念。

MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization

  • 利用联合视觉语言模型反演与特征对齐,还原模型内部编码。
  • 在不同长度输出上成功逆向生成高质量且语义一致的图像。
  • 首次实现对视觉语言模型概念的可视化解释,适合模型可解释性研究者。

视觉语言模型(VLM)将多模态输入编码于庞大而复杂的架构中,限制了透明度与可信度。我们提出多模态反演框架MIMIC,用于逆向解析VLM的内部表征。MIMIC结合基于VLM的联合反演与特征对齐目标,以应对VLM的自回归处理机制,并引入三重正则化项:空间对齐、自然图像平滑性与语义真实性。我们在多种自由格式的VLM输出上评估MIMIC,涵盖不同长度的输出结果,采用标准视觉质量指标与基于文本的语义指标进行量化与定性分析。据我们所知,这是首个针对视觉语言模型概念进行视觉解释的模型反演方法。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) encode multimodal inputs over large, complex, and difficult-to-interpret architectures, which limit transparency and trust. We propose a Multimodal Inversion for Model Interpretation and Conceptualization (MIMIC) framework that inverts the internal encodings of VLMs. MIMIC uses a joint VLM-based inversion and a feature alignment objective to account for VLM's autoregressive processing. It additionally includes a triplet of regularizers for spatial alignment, natural image smoothness, and semantic realism. We evaluate MIMIC both quantitatively and qualitatively by inverting visual concepts across a range of free-form VLM outputs of varying length. Reported results include both standard visual quality metrics and semantic text-based metrics. To the best of our knowledge, this is the first model inversion approach addressing visual interpretations of VLM concepts.

模型解释视觉语言模型反演生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。