arXiv:2605.26520cs.CVcs.AI2026-05

让AI像人一样边看图边推理,还能自我修正。

InterSketch: An Interleaved Reasoning Model with Self-correcting Visual Sketch and Stepwise Reward

论文配图:InterSketch: An Interleaved Reasoning Model with Self-correcting Visual Sketch and Stepwise Reward
图 1 · 摘自论文原文
  • 用图文交替推理+自修正画草图,提升长程理解能力
  • 在多个视觉推理任务上超越Gemini-3-Pro等闭源模型
  • 适合需要深度视觉分析与多步推理的复杂场景

尽管视觉语言模型(VLMs)已具备多轮视觉推理能力,其推理过程仍较浅显,且以文本为中心,难以应对复杂视觉挑战。相比之下,人类思维通常具有长程、图文交织的推理链条(VT-CoT)。为此,我们提出InterSketch,一种通过自修正视觉草图和分步奖励机制增强图文交织推理能力的模型。该模型利用外部工具动态生成中间视觉草图,并将其与文本推理交错,实现对长程视觉理解任务的有效感知与逻辑推理。首先,在冷启动阶段,构建高质量的图文交织式推理数据集,并引入反思机制,使模型具备多轮交错推理与自纠错能力;随后在强化学习阶段,设计分步奖励机制,缓解长程推理中奖励信号稀疏的问题。在多个视觉推理基准上的实验证明,InterSketch效果显著,甚至优于Gemini-3-Pro等专有模型。

原文摘要 · Abstract (English)

While vision-language models (VLMs) have exhibited multi-turn visual reasoning capabilities, their reasoning trajectories remain relatively shallow and are dominated by a text-centric paradigm, limiting their applicability to complex visual challenges. In contrast, human-like thought typically involves long-horizon reasoning with an interleaved visual-textual chain-of-thought (VT-CoT). To bridge this gap, we introduce InterSketch, an interleaved reasoning model to enhance the VT-CoT capability via self-correcting and stepwise reward mechanisms. InterSketch dynamically generates intermediate visual sketches using external tools and interleaves them with textual reasoning, enabling effective perception and logical reasoning over long-horizon visual understanding tasks. Specifically, in the first cold-start stage, we propose a synthesized high-quality interleaved VT-CoT dataset and include a reflection mechanism to enable the model's capability in multi-turn interleaved reasoning and self-correction. In the subsequent reinforcement learning (RL) stage, we design a stepwise reward mechanism to mitigate the sparsity of reward signals inherent in end-only supervision over long-horizon reasoning. Extensive experiments on visual reasoning benchmarks demonstrate the effectiveness of InterSketch, even outperforming proprietary models such as Gemini-3-Pro.

视觉推理图文交织自修正强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。