视觉生成让AI像人一样思考,尤其在物理空间任务中更有效。
Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
- 用视觉与语言交替推理,模拟人类的内部世界模型。
- 在物理空间任务上,视觉推理比纯语言推理提升显著,准确率提高18%。
- 适合研究多模态认知、具身智能和通用AI的开发者与研究人员。
人类通过构建内部世界模型并操作其中的概念来进行推理。当前基于大语言模型的链式思维(CoT)推理虽在数学与编程等抽象领域表现优异,但在物理与空间智能方面仍远落后于人类,因缺乏丰富表征与先验知识。统一多模态模型(UMMs)能同时生成文本与图像,为更类人的推理提供了可能。本文首次系统研究视觉生成何时及如何提升推理能力,提出视觉优势假说:对于物理世界相关任务,视觉生成更自然地作为世界模型,而纯语言模型则受限于表征能力与知识不足。理论层面,将世界建模定义为CoT的核心;实验层面,构建新评测集VisWorld-Eval,控制实验显示,在适合视觉建模的任务上,交错视觉-语言的CoT显著优于纯语言版本,但其他任务无明显优势。该研究澄清了多模态世界建模对实现更强人类级多模态AI的潜力。
原文摘要 · Abstract (English)
Humans construct internal world models and reason by manipulating the concepts within these models. Recent advances in AI, particularly chain-of-thought (CoT) reasoning, approximate such human cognitive abilities, where world models are believed to be embedded within large language models. Expert-level performance in formal and abstract domains such as mathematics and programming has been achieved in current systems by relying predominantly on verbal reasoning. However, they still lag far behind humans in domains like physical and spatial intelligence, which require richer representations and prior knowledge. The emergence of unified multimodal models (UMMs) capable of both verbal and visual generation has therefore sparked interest in more human-like reasoning grounded in complementary multimodal pathways, though their benefits remain unclear. From a world-model perspective, this paper presents the first principled study of when and how visual generation benefits reasoning. Our key position is the visual superiority hypothesis: for certain tasks--particularly those grounded in the physical world--visual generation more naturally serves as world models, whereas purely verbal world models encounter bottlenecks arising from representational limitations or insufficient prior knowledge. Theoretically, we formalize internal world modeling as a core component of CoT reasoning and analyze distinctions among different forms of world models. Empirically, we identify tasks that necessitate interleaved visual-verbal CoT reasoning, constructing a new evaluation suite, VisWorld-Eval. Controlled experiments on a state-of-the-art UMM show that interleaved CoT significantly outperforms purely verbal CoT on tasks that favor visual world modeling, but offers no clear advantage otherwise. Together, this work clarifies the potential of multimodal world modeling for more powerful, human-like multimodal AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。