发现大模型推理时依赖文字而非视觉,提出新方法提升视觉感知能力。
Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs

- 通过控制视觉与文本变量,诊断模型因果推理失败机制。
- 17个主流模型均更依赖文本,性能越强越易受干扰。
- 新算法让模型主动依赖视觉证据,适合改进多模态推理任务。
尽管大型多模态模型(LMMs)在通用视频理解上表现优异,但其在因果发现中对文本先验的依赖已被视为关键缺陷。现有基准仅评估结果准确性,无法揭示缺陷来源与程度。本文提出ProCauEval,一种基于扰动的评估协议,从结果判断转向机制诊断,通过五种可控配置系统性地操纵视觉与文本模态,分解二者对模型行为的贡献并剖析失败模式。在17个主流LMM上评估发现:模型能准确感知视频内容,但在因果推理中系统性地忽视视觉信息。进一步观察到,更强的后训练反而加剧对文本先验的依赖,且基线性能越高,在扰动下越脆弱。为此,提出抗蒸馏策略优化(ADPO),一种基于负教师对齐的强化学习框架,通过显式将策略远离由视觉损坏诱导的仅依赖先验的反事实教师,最大化原始输入与视觉损坏输入下策略分布的差异,强制模型基于视觉证据进行推理。大量实验表明,ADPO在不牺牲基础理解的前提下提升了视觉参与度,为可靠因果发现迈出初步一步。
原文摘要 · Abstract (English)
Although Large Multimodal Models (LMMs) have achieved strong performance on general video understanding, their susceptibility to textual prior shortcuts during causal discovery has been recognized as a critical deficit. The underlying mechanisms of this phenomenon remain incompletely understood, as existing benchmarks only measure response accuracy without revealing the sources and extent of the deficit. We introduce ProCauEval, a perturbation-based evaluation protocol that shifts from outcome assessment to mechanism diagnosis, probing causal discovery through five controlled configurations that systematically manipulate visual and textual modalities to decompose their respective contributions to model behavior and dissect the failure modes. Evaluating 17 mainstream LMMs, we find that models faithfully perceive video content yet systematically underexploit it during causal reasoning. We further observe that stronger post-training amplifies rather than mitigates textual prior reliance, and that higher baseline performance correlates with greater fragility under perturbation. To address these, we propose Anti-Distillation Policy Optimization (ADPO), a reinforcement learning framework built on negative teacher alignment, which augments GRPO by explicitly pushing the policy away from a prior-only counterfactual teacher induced by visual corruption. Specifically, ADPO maximizes the divergence between the policy distributions conditioned on the original and visually corrupted inputs, thereby forcing the model to ground its reasoning in visual evidence rather than textual shortcuts. Extensive experiments show that ADPO improves visual engagement without sacrificing fundamental comprehension, thus offering a preliminary step toward reliable causal discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。