首份系统综述揭示多模态大模型视觉语言感知的演进历程
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

- 将视觉语言感知视为统一能力,类比人类先天感知机制
- 提出五阶段演化框架,梳理关键模型与里程碑方法
- 揭示开放挑战,为通用多模态智能提供研究路线图
多模态大语言模型(MLLMs)近期在统合视觉-语言理解与推理方面取得显著进展,尤其以OpenAI的O系列和DeepSeek的R系列模型为代表,推动感知中心智能范式变革。然而,现有综述多碎片化,分别关注视觉或语言,难以捕捉跨模态感知的整合演进。为此,本文首次系统性地探讨了MLLM中统一的视觉-语言感知。具体贡献包括:(1) 将MLLM感知形式化为类同人类先天感知的内在统一能力;(2) 提出五阶段演化分类法,梳理各阶段代表性方法与关键里程碑;(3) 识别开放挑战,展望迈向真正通用、统一的多模态智能的前沿方向。本研究旨在为人工通用智能(AGI)的发展提供基础认知与可行动路线。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially following the introduction of models such as OpenAI's O-series and DeepSeek's R-series, which have driven a paradigm shift toward perception-centric intelligence. However, there remains a lack of systematic surveys that examine perception from a truly unified vision-language perspective -- one that treats vision and language as an inseparable modality. Existing reviews are often fragmented, focusing separately on either vision or language, and thus rarely capture the cross-modal evolution of perception as an integrated capability. To bridge this gap, we present the first systematic survey of unified vision-language perception in MLLMs. Specifically, we (1) formalize MLLM perception as an intrinsic, unified vision-language capability analogous to human innate perception, (2) introduce a five-stage taxonomy tracing the paradigm evolution of MLLM perception and survey representative methods and milestones at each phase, and (3) identify open challenges and outline promising research directions toward truly general, unified multimodal intelligence. We hope our study will provide both a foundational understanding and an actionable roadmap to foster further innovation on the path toward artificial general intelligence (AGI).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。