让AI像人一样用图像思考,突破语言局限。
Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
- 将视觉信息作为思维中间步骤,构建动态认知空间。
- 提出三阶段演进框架:工具使用→程序操控→内在想象。
- 系统梳理方法、评测与应用,指引未来研究方向。
多模态推理的进展主要依赖于文本链式思维(CoT),该范式将视觉视为静态初始上下文,导致感知数据与符号思维间存在根本性语义鸿沟。人类认知常超越语言,以视觉为动态心智画板。这一趋势正重塑AI,推动模型从‘思考图像’转向‘用图像思考’。新范式使视觉成为思维过程中的可操作中间步骤,实现从被动输入到主动认知工作区的转变。本文系统梳理了智能演进的三个关键阶段:外部工具探索、程序化操控、内在想象。提出四项贡献:(1)建立‘用图像思考’的理论基础与三阶段框架;(2)全面回顾各阶段核心技术方法;(3)分析评估基准与典型应用;(4)识别核心挑战并展望未来方向。旨在为更强大、更契合人类认知的多模态智能提供清晰路线图。
原文摘要 · Abstract (English)
Recent progress in multimodal reasoning has been significantly advanced by textual Chain-of-Thought (CoT), a paradigm where models conduct reasoning within language. This text-centric approach, however, treats vision as a static, initial context, creating a fundamental "semantic gap" between rich perceptual data and discrete symbolic thought. Human cognition often transcends language, utilizing vision as a dynamic mental sketchpad. A similar evolution is now unfolding in AI, marking a fundamental paradigm shift from models that merely think about images to those that can truly think with images. This emerging paradigm is characterized by models leveraging visual information as intermediate steps in their thought process, transforming vision from a passive input into a dynamic, manipulable cognitive workspace. In this survey, we chart this evolution of intelligence along a trajectory of increasing cognitive autonomy, which unfolds across three key stages: from external tool exploration, through programmatic manipulation, to intrinsic imagination. To structure this rapidly evolving field, our survey makes four key contributions. (1) We establish the foundational principles of the think with image paradigm and its three-stage framework. (2) We provide a comprehensive review of the core methods that characterize each stage of this roadmap. (3) We analyze the critical landscape of evaluation benchmarks and transformative applications. (4) We identify significant challenges and outline promising future directions. By providing this structured overview, we aim to offer a clear roadmap for future research towards more powerful and human-aligned multimodal AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。