arXiv:2410.08165cs.LGcs.CV2024-10被引 3

提出视觉推理新方法,让模型像解题一样分步思考。

Chain-of-Sketch: Enabling Global Visual Reasoning

  • 用分步视觉推理解题,模仿语言模型的思维链
  • 引入马尔可夫结构提升小模型在复杂任务上的表现
  • 适合需要全局推理的图像理解场景

现代视觉模型在依赖局部特征的任务中表现优异,但对需全局推理的任务仍存短板。本文扩展了包含图、字符串、迷宫和图像网格的全局视觉数据集,发现大视觉模型与主流多模态大模型在此类任务上均表现不佳。我们提出‘全局度’衡量指标解释学习效率低下的原因。为此,提出链式草图(Chain-of-Sketch, CoS)方法,将复杂任务分解为中间视觉步骤,类似语言模型的思维链。进一步发现,并非所有CoS策略效果相同;关键在于对CoS帧施加马尔可夫结构,提出归纳式CoS(inductive CoS),显著提升分布外泛化能力,且在小模型上表现更优。

原文摘要 · Abstract (English)

Modern vision models have achieved remarkable success in benchmarks where local features provide critical information about the target. There is now a growing interest in tackling tasks requiring more global reasoning, where local features do not provide significant information. Minsky and Papert put forward such tasks in 1969 with their connectivity study, exposing the limitations of the perceptron model. In this paper, we introduce an expanded set of global visual datasets involving graphs, strings, mazes, and image grids. We show that large vision models still struggle to learn these tasks efficiently. Similarly, state-of-the-art multi-modal LLMs perform poorly on these datasets. We explain this learning inefficiency by means of the 'globality degree' measure. To mitigate this, we propose a method called chain-of-sketch (CoS). Similar to the chain-of-thought and scratchpad techniques used in language models, CoS breaks the original task into intermediate visual steps to help learn a complex task. In addition, we show that not all CoS strategies perform equally well. Our key insight is to impose a Markovian structure on the CoS frames. This leads to the introduction of 'inductive CoS' which achieves better out-of-distribution generalization and performs well even with smaller models compared to non-inductive variants.

视觉推理思维链模型泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。