揭示视觉语言模型生成过程中各信息源的动态影响机制
Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework

- 构建因果时序框架,追踪视觉、问题和生成前缀的逐步影响
- 发现模型从早期依赖问题和视觉转向后期依赖生成内容
- 适合研究多模态生成机制或评估模型推理可信度的学者
视觉语言模型在复杂图像与视频理解任务中日益重要,但传统评估指标仅关注最终答案质量,难以揭示不同信息源如何影响生成过程。本文提出一种基于结构因果模型的因果时序评估框架,通过干预与后门调整,推导出三个分步索引的因果驱动度量:视觉因果驱动(VCD)、问题因果驱动(QCD)和前缀因果驱动(PCD),用于刻画源相关生成模式,无需参考答案。在Qwen3-VL-8B-Instruct上对MAVIS、LLaVA-Video-178K和MiraData的实验,以及在InternVL2-8B上的跨模型验证均显示,模型呈现从早期强问题与视觉引导向后期更依赖生成前缀的转变。随机干预验证表明,QCD与PCD相比观测PMI基线分别降低34.8%和47.1%的恢复误差。在VLMBias数据集上,前缀-视觉失衡分数达到0.767 AUROC和0.873 AUPRC,可有效区分先验驱动与视觉依存生成。结果表明,因果驱动轨迹为多模态生成提供了互补的源级诊断工具。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly evaluated on complex image and video understanding tasks, yet conventional metrics primarily assess final-answer quality and reveal little about how different information sources shape the generation process. We propose a causal and temporal evaluation framework that traces the evolving roles of visual input, question text, and generated prefixes during autoregressive decoding. Grounded in a Structural Causal Model, we use interventions and backdoor adjustment to derive three step-indexed causal-drive metrics---Visual Causal Drive (VCD), Question Causal Drive (QCD), and Prefix Causal Drive (PCD)---for characterizing source-specific generation patterns without requiring reference answers. Experiments on Qwen3-VL-8B-Instruct across MAVIS, LLaVA-Video-178K, and MiraData, together with cross-model validation on InternVL2-8B, reveal a consistent transition from stronger early question and visual guidance toward increasing reliance on generated prefixes. Randomized-intervention validation shows that QCD and PCD reduce recovery error over observational PMI baselines by 34.8\% and 47.1\%, respectively. On VLMBias, the prefix--visual imbalance score achieves 0.767 AUROC and 0.873 AUPRC for distinguishing prior-driven from visually grounded generations. These results show that causal-drive trajectories provide complementary source-level diagnostics for multimodal generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。