arXiv:2412.02172cs.CVcs.AI2024-12CVPR被引 18

首个细粒度视觉推理纠错评估基准,揭示大模型自改进瓶颈

VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning

  • 设计细粒度步骤级批判与修正机制,要求逐步评估并解释
  • 人类批判使模型性能显著提升,但模型自动生成批判效果差
  • 发现三大批判失败模式,提出回溯图像验证策略提升13.5%

大型视觉语言模型(LVLMs)在推理过程中进行自我批判与修正的能力,是实现其自改进的关键。然而,对这一能力的系统性分析仍显不足。本文提出VISCO,首个全面评估LVLM细粒度批判与修正能力的基准。相较于以往仅用单一标量评价整个推理过程的方法,VISCO采用密集、细粒度的批判方式,要求模型评估思维链中每一步的正确性,并提供自然语言解释支持判断。对24个LVLM的广泛评估表明,人类编写的批判能显著提升修正后性能,展现出自改进策略潜力。但模型生成的批判帮助有限,甚至可能损害性能,说明批判是关键瓶颈。我们识别出三种常见批判失败模式:忽视视觉感知、不愿否定、过度假设错误传播。为此,提出有效回溯策略(LookBack),通过重新审视图像验证初始推理中的每一项信息,使批判与修正性能最高提升13.5%。

原文摘要 · Abstract (English)

The ability of large vision-language models (LVLMs) to critique and correct their reasoning is an essential building block towards their self-improvement. However, a systematic analysis of such capabilities in LVLMs is still lacking. We propose VISCO, the first benchmark to extensively analyze the fine-grained critique and correction capabilities of LVLMs. Compared to existing work that uses a single scalar value to critique the entire reasoning [4], VISCO features dense and fine-grained critique, requiring LVLMs to evaluate the correctness of each step in the chain-of-thought and provide natural language explanations to support their judgments. Extensive evaluation of 24 LVLMs demonstrates that human-written critiques significantly enhance the performance after correction, showcasing the potential of the self-improvement strategy. However, the model-generated critiques are less helpful and sometimes detrimental to the performance, suggesting that critique is the crucial bottleneck. We identified three common patterns in critique failures: failure to critique visual perception, reluctance to "say no", and exaggerated assumption of error propagation. To address these issues, we propose an effective LookBack strategy that revisits the image to verify each piece of information in the initial reasoning. LookBack significantly improves critique and correction performance by up to 13.5%.

视觉推理自改进批判修正基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。