用扩散模型做视觉主导的多模态推理,效果远超传统文本中心模型。
DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models
- 将多模态推理转为图像到图像生成任务,提升空间精度与逻辑一致性。
- 在4个领域测试中,性能比GPT-5高314.2%,比Gemini-3-Flash高111.6%。
- 适合需要精确视觉推理的场景,如规划、布局和组合优化任务。
尽管近期多模态大语言模型(MLLMs)在多模态推理上取得显著进展,其推理过程仍以文本为中心,导致在复杂长时程、视觉主导任务中表现不佳。本文提出一种新的生成式多模态推理范式,并引入基于扩散模型的DiffThinker框架。概念上,DiffThinker将多模态推理重构为原生的图像到图像生成任务,在视觉主导任务中实现更优的逻辑一致性和空间精度。我们系统对比了DiffThinker与MLLMs,首次深入揭示该范式的四大核心特性:效率、可控性、原生并行性与协作性。在四个领域(序列规划、组合优化、约束满足、空间配置)的大量实验表明,DiffThinker显著优于领先闭源模型,包括GPT-5(+314.2%)、Gemini-3-Flash(+111.6%),以及微调后的Qwen3-VL-32B基线(+39.0%),凸显生成式多模态推理在视觉主导推理中的巨大潜力。
原文摘要 · Abstract (English)
While recent Multimodal Large Language Models (MLLMs) have attained significant strides in multimodal reasoning, their reasoning processes remain predominantly text-centric, leading to suboptimal performance in complex long-horizon, vision-centric tasks. In this paper, we establish a novel Generative Multimodal Reasoning paradigm and introduce DiffThinker, a diffusion-based reasoning framework. Conceptually, DiffThinker reformulates multimodal reasoning as a native generative image-to-image task, achieving superior logical consistency and spatial precision in vision-centric tasks. We perform a systematic comparison between DiffThinker and MLLMs, providing the first in-depth investigation into the intrinsic characteristics of this paradigm, revealing four core properties: efficiency, controllability, native parallelism, and collaboration. Extensive experiments across four domains (sequential planning, combinatorial optimization, constraint satisfaction, and spatial configuration) demonstrate that DiffThinker significantly outperforms leading closed source models including GPT-5 (+314.2\%) and Gemini-3-Flash (+111.6\%), as well as the fine-tuned Qwen3-VL-32B baseline (+39.0\%), highlighting generative multimodal reasoning as a promising approach for vision-centric reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。