剖析视觉语言模型从感知到认知的断层,提出统一分析框架
From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- 构建感知与认知双层框架,解析多模态理解断层
- 指出当前模型易幻觉,源于感知与推理脱节
- 适合研究多模态推理与下一代AI的学者参考
多模态大语言模型(MLLMs)追求对物理世界的深度、类人理解与交互,但常在信息获取(感知)与推理(认知)间表现出浅层且不连贯的整合,导致一系列推理失败,其中幻觉最为突出。这暴露了根本性挑战:能处理像素不等于能构建连贯可信的内部世界模型。为此,本文提出全新统一分析框架「从感知到认知」,将视觉-语言交互理解拆解为两个相互依赖的层级:感知层——准确提取视觉信息并实现细粒度文本指令对齐;认知层——基于感知基础,进行主动、多步、目标导向的高级推理,核心是动态的观察-思考-验证循环。该框架系统分析当前MLLMs在两层中的关键瓶颈,综述前沿方法,涵盖从增强低层视觉表征到改进高层推理范式的技术。同时,回顾关键评测基准并展望未来方向。本综述旨在为研究社区提供清晰结构化视角,理解当前模型内在局限,并指引构建具备深度推理与真实世界理解能力的下一代模型之路。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) strive to achieve a profound, human-like understanding of and interaction with the physical world, but often exhibit a shallow and incoherent integration when acquiring information (Perception) and conducting reasoning (Cognition). This disconnect leads to a spectrum of reasoning failures, with hallucination being the most prominent. Collectively, these issues expose a fundamental challenge: the ability to process pixels does not yet confer the ability to construct a coherent, credible internal world model. To systematically dissect and address this challenge, this survey introduces a novel and unified analytical framework: ``From Perception to Cognition." We deconstruct the complex process of vision-language interactive understanding into two interdependent layers: Perception, the foundational ability to accurately extract visual information and achieve fine-grained alignment with textual instructions; and Cognition, the higher-order capability for proactive, multi-step, goal-oriented reasoning built upon this perceptual foundation, the core of which is the formation of a dynamic observe-think-verify reasoning loop. Guided by this framework, this paper systematically analyzes the key bottlenecks of current MLLMs at both layers. It surveys the landscape of cutting-edge methods designed to address these challenges, spanning from techniques that enhance low-level visual representations to those that improve high-level reasoning paradigms. Furthermore, we review critical benchmarks and delineate future research directions. This survey aims to provide the research community with a clear, structured perspective for understanding the intrinsic limitations of current MLLMs and to illuminate the path toward building next-generation models capable of deep reasoning and a genuine understanding of the world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。