arXiv:2509.23789cs.LGcs.CR2025-09被引 1

视觉思维链让视觉语言模型更聪明却更脆弱,提出增强鲁棒性的插件式方案

Visual CoT Makes VLMs Smarter but More Fragile

  • 将视觉编辑融入推理过程,提升多模态表现
  • 在12类图像噪声下,准确率下降更剧烈,显示更强敏感性
  • 通过定位模型注入可信视觉线索,稳定推理且无需修改架构

思维链(CoT)技术显著提升了视觉语言模型(VLMs)的推理能力。视觉思维链(Visual CoT)通过引入裁剪、标注等显式视觉编辑,进一步增强了多模态性能。然而,其对图像级噪声的鲁棒性尚未被研究。本文首次系统评估了视觉思维链在视觉扰动下的表现。我们在4个视觉问答(VQA)数据集上覆盖12种图像退化类型,全面比较使用与不使用视觉思维链的VLMs。结果表明,无论图像是否受噪声污染,视觉思维链均能提升绝对准确率;但同时也加剧了对输入扰动的敏感性,导致性能下降更剧烈。深入分析发现,视觉思维链中间环节的编辑图像块是脆弱性的主要来源。基于此,我们提出一种即插即用的鲁棒性增强方法,在视觉思维链流程中引入Grounding DINO模型,提供高置信度局部视觉提示以稳定推理。本工作揭示了视觉思维链的清晰脆弱模式,并提供了有效的、与架构无关的鲁棒性提升方案。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) techniques have significantly enhanced reasoning in Vision-Language Models (VLMs). Extending this paradigm, Visual CoT integrates explicit visual edits, such as cropping or annotating regions of interest, into the reasoning process, achieving superior multimodal performance. However, the robustness of Visual CoT-based VLMs against image-level noise remains unexplored. In this paper, we present the first systematic evaluation of Visual CoT robustness under visual perturbations. Our benchmark spans 12 image corruption types across 4 Visual Question Answering (VQA) datasets, enabling a comprehensive comparison between VLMs that use Visual CoT, and VLMs that do not. The results reveal that integrating Visual CoT consistently improves absolute accuracy regardless of whether the input images are clean or corrupted by noise; however, it also increases sensitivity to input perturbations, resulting in sharper performance degradation compared to standard VLMs. Through extensive analysis, we identify the intermediate reasoning components of Visual CoT, i.e., the edited image patches , as the primary source of fragility. Building on this analysis, we propose a plug-and-play robustness enhancement method that integrates Grounding DINO model into the Visual CoT pipeline, providing high-confidence local visual cues to stabilize reasoning. Our work reveals clear fragility patterns in Visual CoT and offers an effective, architecture-agnostic solution for enhancing visual robustness.

视觉推理鲁棒性VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。