arXiv:2505.22525cs.CVcs.AI2025-05被引 68

让AI通过生成图像进行视觉思维,实现更像人类的推理。

Thinking with Generated Images

  • 模型自主生成中间视觉步骤,融合图文思考过程。
  • 复杂多物体任务准确率提升50%(38%→57%)。
  • 适合科研、设计、分析等需要视觉想象的场景。

我们提出「基于生成图像的思维」,一种革新大型多模态模型视觉推理的新范式。该方法使模型能通过自发生成中间视觉思考步骤,在文本与视觉模态间原生地进行跨模态思考。现有方法受限于固定输入图像或纯文本链式思考。本方法通过两种机制实现:(1) 基于中间视觉子目标的视觉生成,将复杂任务分解为可逐步生成与整合的组件;(2) 带自我批判的视觉生成,模型先生成初始视觉假设,再通过文本推理分析其缺陷,并据此生成优化结果。在视觉生成基准测试中,我们的模型在处理复杂多物体场景时,相对基线实现高达50%的性能提升(从38%到57%)。该方法赋能生物学家探索新型蛋白质结构、建筑师迭代空间设计、法医重建案发现场及篮球运动员构想战术策略,使AI具备类人化的视觉想象力与迭代优化能力。开源代码已发布于https://github.com/GAIR-NLP/thinking-with-generated-images。

原文摘要 · Abstract (English)

We present Thinking with Generated Images, a novel paradigm that fundamentally transforms how large multimodal models (LMMs) engage with visual reasoning by enabling them to natively think across text and vision modalities through spontaneous generation of intermediate visual thinking steps. Current visual reasoning with LMMs is constrained to either processing fixed user-provided images or reasoning solely through text-based chain-of-thought (CoT). Thinking with Generated Images unlocks a new dimension of cognitive capability where models can actively construct intermediate visual thoughts, critique their own visual hypotheses, and refine them as integral components of their reasoning process. We demonstrate the effectiveness of our approach through two complementary mechanisms: (1) vision generation with intermediate visual subgoals, where models decompose complex visual tasks into manageable components that are generated and integrated progressively, and (2) vision generation with self-critique, where models generate an initial visual hypothesis, analyze its shortcomings through textual reasoning, and produce refined outputs based on their own critiques. Our experiments on vision generation benchmarks show substantial improvements over baseline approaches, with our models achieving up to 50% (from 38% to 57%) relative improvement in handling complex multi-object scenarios. From biochemists exploring novel protein structures, and architects iterating on spatial designs, to forensic analysts reconstructing crime scenes, and basketball players envisioning strategic plays, our approach enables AI models to engage in the kind of visual imagination and iterative refinement that characterizes human creative, analytical, and strategic thinking. We release our open-source suite at https://github.com/GAIR-NLP/thinking-with-generated-images.

视觉推理多模态生成模型AI思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。