arXiv:2511.18359cs.CV2025-11中稿 · IEEE/CVF Conferenc…

用视频还原视觉模型的决策逻辑,让模型解释自己为何这样判断。

TRANSPORTER: Transferring Visual Semantics from VLM Manifolds

  • 通过最优传输映射将模型输出转换为可生成的视频方向
  • 生成的视频能清晰反映物体属性、动作和场景变化对预测的影响
  • 无需修改模型即可解释多种视觉语言模型的内部机制

当前视觉语言模型(VLM)虽能理解复杂场景中的多类对象、动作与动态变化,但其内部推理过程仍难以解析。受文本到视频生成模型启发,本文提出一种从逻辑值到视频(L2V)的新任务,并设计了不依赖具体模型的TRANSPORTER方法,用于生成体现VLM决策规则的视频。利用高保真文本到视频模型生成能力,TRANSPORTER学习将模型高语义嵌入空间与视频生成空间进行最优传输对齐,使模型输出的逻辑分数对应特定视频生成方向。该方法能生成随物体属性、动作副词及场景上下文变化而改变的视频内容。在多个VLM上的定量与定性评估表明,L2V为模型可解释性提供了前所未有的高保真新视角。

原文摘要 · Abstract (English)

How do video understanding models acquire their answers? Although current Vision Language Models (VLMs) reason over complex scenes with diverse objects, action performances, and scene dynamics, understanding and controlling their internal processes remains an open challenge. Motivated by recent advancements in text-to-video (T2V) generative models, this paper introduces a logits-to-video (L2V) task alongside a model-independent approach, TRANSPORTER, to generate videos that capture the underlying rules behind VLMs' predictions. Given the high-visual-fidelity produced by T2V models, TRANSPORTER learns an optimal transport coupling to VLM's high-semantic embedding spaces. In turn, logit scores define embedding directions for conditional video generation. TRANSPORTER generates videos that reflect caption changes over diverse object attributes, action adverbs, and scene context. Quantitative and qualitative evaluations across VLMs demonstrate that L2V can provide a fidelity-rich, novel direction for model interpretability that has not been previously explored.

模型解释视觉语言模型视频生成可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。