arXiv:2509.25866cs.CV2025-09被引 11

让视觉模型像画画一样思考,直接在图像嵌入空间中生成视觉推理。

DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning

  • 在图像嵌入空间内原生生成视觉思维,无需调用外部工具
  • 构建3.1万条图文交织的推理轨迹数据集,标注准确率高
  • 适合研究视觉推理、多模态模型设计的学者和开发者

“以图思考”代表了视觉语言模型(VLMs)推理范式的重大转变,从以文本为主导的链式思考转向图像交互式推理。通过调用视觉工具或生成中间视觉表示,VLMs 可迭代关注细粒度区域,实现更深入的图像理解与更忠实的多模态推理。然而,这一新兴范式在数据构建精度、结构设计和应用场景拓展方面仍有较大探索空间。为此,我们提出 DeepSketcher,一个包含图文交错数据集和自包含模型的完整工具套件。该数据集包含 31,000 条链式思考(CoT)推理轨迹,涵盖多样化的工具调用与编辑后图像,覆盖多种数据类型与操作指令,具有高标注精度。基于此资源,我们设计了一种模型,可进行图文交织推理,并直接在视觉嵌入空间中原生生成“视觉思考”,而非调用外部工具并重复编码生成图像。该设计实现了无工具、更灵活的“以图思考”。在多模态推理基准上的大量实验表明其性能优异,验证了数据集的价值与模型设计的有效性。

原文摘要 · Abstract (English)

The "thinking with images" paradigm represents a pivotal shift in the reasoning of Vision Language Models (VLMs), moving from text-dominant chain-of-thought to image-interactive reasoning. By invoking visual tools or generating intermediate visual representations, VLMs can iteratively attend to fine-grained regions, enabling deeper image understanding and more faithful multimodal reasoning. As an emerging paradigm, however, it still leaves substantial room for exploration in data construction accuracy, structural design, and broader application scenarios, which offer rich opportunities for advancing multimodal reasoning. To further advance this line of work, we present DeepSketcher, a comprehensive suite comprising both an image-text interleaved dataset and a self-contained model. The dataset contains 31k chain-of-thought (CoT) reasoning trajectories with diverse tool calls and resulting edited images, covering a wide range of data types and manipulation instructions with high annotation accuracy. Building on this resource, we design a model that performs interleaved image-text reasoning and natively generates "visual thoughts" by operating directly in the visual embedding space, rather than invoking external tools and repeatedly re-encoding generated images. This design enables tool-free and more flexible "thinking with images". Extensive experiments on multimodal reasoning benchmarks demonstrate strong performance, validating both the utility of the dataset and the effectiveness of the model design.

多模态推理视觉思考图像生成VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。