arXiv:2603.07659cs.CV2026-03中稿 · CVPR被引 1

通过多轮反事实推理提升视觉语言模型的测试时鲁棒性

Scaling Test-Time Robustness of Vision-Language Models via Self-Critical Inference Framework

  • 引入自批判推理框架,结合文本与视觉扰动进行多轮反事实思考
  • 增加推理轮次可显著提升鲁棒性,优于单步反事实方法
  • 提出动态鲁棒性基准,针对不同模型评估语言偏见与敏感性

大型语言模型(LLMs)的兴起推动了多模态学习的快速发展,尤其促进了大型视觉语言模型(LVLMs)的发展。然而,现有LVLM训练范式过度依赖语言模型组件,导致语言偏见和语言敏感性两大关键鲁棒性问题。为同时解决这些问题,我们提出一种新型自批判推理(SCI)框架,通过在视觉对比解码基础上引入多轮反事实推理,结合文本与视觉扰动进行迭代思考。该过程进一步提出一种通过增加反事实推理轮次来提升鲁棒性的新策略。此外,我们观察到不同LVLM的失败模式存在显著差异,表明固定鲁棒性基准难以真实反映其可靠性。为此,我们提出动态鲁棒性基准(DRBench),一个针对语言偏见与敏感性问题的模型定制化评估框架。大量实验表明,SCI在DRBench上持续优于基线方法,且增加推理轮次能进一步提升鲁棒性,超越现有单步反事实推理方法。

原文摘要 · Abstract (English)

The emergence of Large Language Models (LLMs) has driven rapid progress in multi-modal learning, particularly in the development of Large Vision-Language Models (LVLMs). However, existing LVLM training paradigms place excessive reliance on the LLM component, giving rise to two critical robustness challenges: language bias and language sensitivity. To address both issues simultaneously, we propose a novel Self-Critical Inference (SCI) framework that extends Visual Contrastive Decoding by conducting multi-round counterfactual reasoning through both textual and visual perturbations. This process further introduces a new strategy for improving robustness by scaling the number of counterfactual rounds. Moreover, we also observe that failure cases of LVLMs differ significantly across models, indicating that fixed robustness benchmarks may not be able to capture the true reliability of LVLMs. To this end, we propose the Dynamic Robustness Benchmark (DRBench), a model-specific evaluation framework targeting both language bias and sensitivity issues. Extensive experiments show that SCI consistently outperforms baseline methods on DRBench, and that increasing the number of inference rounds further boosts robustness beyond existing single-step counterfactual reasoning methods.

视觉语言模型鲁棒性反事实推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。