提出一种无需训练的诊断方法,可逐步分析模型依赖视觉区域的情况。
Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
- 通过对比掩码前后推理过程,逐步定位模型对视觉区域的依赖
- 发现部分模型缺失证据时会幻觉,另一些则对扰动极度敏感
- 适合关注多模态模型推理可信度的研究者
我们提出对比区域掩码(CRM),一种无需训练的诊断方法,可揭示多模态大语言模型(MLLM)在链式思维(CoT)推理每一步中对特定视觉区域的依赖。与以往仅关注最终答案或注意力图的方法不同,CRM通过系统性掩码标注区域,并对比结果推理轨迹与未掩码基线,提供因果性的、步骤级的归因。在VisArgs等数据集上的应用表明,某些模型虽保持推理结构,但证据缺失时会产生幻觉;另一些则紧密依赖视觉线索,但在扰动下迅速崩溃。该方法将评估重心从答案正确性转向推理忠实性,使视觉基准测试成为诊断工具,凸显了衡量多模态推理性能、鲁棒性和忠实性的必要性。
原文摘要 · Abstract (English)
We introduce Contrastive Region Masking (CRM), a training free diagnostic that reveals how multimodal large language models (MLLMs) depend on specific visual regions at each step of chain-of-thought (CoT) reasoning. Unlike prior approaches limited to final answers or attention maps, CRM provides causal, step-level attribution by systematically masking annotated regions and contrasting the resulting reasoning traces with unmasked baselines. Applied to datasets such as VisArgs, CRM reveals distinct failure modes: some models preserve reasoning structure, but hallucinate when evidence is missing, while others ground tightly to visual cues yet collapse under perturbations. By shifting the evaluation from correctness of answers to faithfulness of reasoning, CRM reframes visual benchmarks as diagnostic tools, highlighting the need for multimodal evaluation frameworks that measure not just performance, but also robustness and fidelity of reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。